Glossary · Models
RLHF
RLHF, or reinforcement learning from human feedback, is a training method in which humans rank model outputs, a reward model learns those preferences, and the language model is optimized to score well on it.
RLHF sits in the Models part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.
Also called Reinforcement learning from human feedback.