Glossary · Models
DPO
DPO, or direct preference optimization, is a method for aligning a language model with preference data by training directly on pairs of preferred and rejected responses, without a separate reward model.
DPO sits in the Models part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.
Also called Direct preference optimization, Preference optimization.