Glossary · Evaluation & safety
MMLU
MMLU is a widely used benchmark of multiple-choice questions spanning dozens of subjects, such as law, medicine and mathematics, used to measure a language model's general knowledge and reasoning.
MMLU sits in the Evaluation & safety part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.
Also called Massive Multitask Language Understanding.