Glossary · Evaluation & safety
HumanEval
HumanEval is a code generation benchmark released by OpenAI consisting of programming problems with unit tests, where a model's generated functions are scored by whether they pass the tests.
HumanEval sits in the Evaluation & safety part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.