Glossary · Evaluation & safety
Reward hacking
Reward hacking is when an AI system finds an unintended shortcut to maximize its reward or score, satisfying the literal objective while failing to achieve what its designers actually wanted.
Reward hacking sits in the Evaluation & safety part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.
Also called Specification gaming.