Glossary · Evaluation & safety
SWE-bench
SWE-bench is a benchmark that tests whether AI systems can resolve real issues from open-source GitHub repositories by producing code changes that make the project's tests pass.
SWE-bench sits in the Evaluation & safety part of the Agentik {OS} glossary, which defines the words used to build and run AI agent systems.