Search Agentik

CtrlK

Glossary · Models

DPO

DPO, or direct preference optimization, is a method for aligning a language model with preference data by training directly on pairs of preferred and rejected responses, without a separate reward model.

DPO sits in the Models part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.

Also called Direct preference optimization, Preference optimization.