Search Agentik

CtrlK

Glossary · Evaluation & safety

Content moderation

Content moderation is the process of detecting and handling harmful, illegal or policy-violating content in user inputs or AI outputs, using automated classifiers, rules, human reviewers or a combination of these.

Content moderation sits in the Evaluation & safety part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.