Glossary · Evaluation & safety
Content moderation
Content moderation is the process of detecting and handling harmful, illegal or policy-violating content in user inputs or AI outputs, using automated classifiers, rules, human reviewers or a combination of these.
Content moderation sits in the Evaluation & safety part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.