Glossary · Foundations
PPO
PPO, or proximal policy optimization, is a reinforcement learning algorithm that makes limited, stable updates to a policy at each step. It has been widely used in RLHF for language models.
PPO sits in the Foundations part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.
Also called Proximal policy optimization.