Glossary · Foundations
GRPO
GRPO, or group relative policy optimization, is a reinforcement learning method that scores each response relative to a group of responses sampled for the same prompt, avoiding a separate value model.
GRPO sits in the Foundations part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.
Also called Group relative policy optimization.