Search Agentik

CtrlK

Glossary · Evaluation & safety

HumanEval

HumanEval is a code generation benchmark released by OpenAI consisting of programming problems with unit tests, where a model's generated functions are scored by whether they pass the tests.

HumanEval sits in the Evaluation & safety part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.