Search Agentik

CtrlK

Glossary · Foundations

PPO

PPO, or proximal policy optimization, is a reinforcement learning algorithm that makes limited, stable updates to a policy at each step. It has been widely used in RLHF for language models.

PPO sits in the Foundations part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.

Also called Proximal policy optimization.