SWE-bench¶
What it is¶
SWE-bench is a benchmark for evaluating LLMs on real-world software engineering tasks. It uses actual issues from GitHub and requires the model to generate a functional patch that passes existing tests. As of January 2027, it remains the industry standard for measuring the autonomous coding capabilities of frontier models like Claude 5.1, GPT-5.5 / GPT-5.6, Gemini 4.0 Pro / Ultra, and DeepSeek-V4.
With the launch of SWE-bench Multilingual, the benchmark has been expanded to support a wider array of programming languages (C, C++, Go, Java, JavaScript, TypeScript, PHP, Ruby, and Rust), making it a truly language-agnostic evaluator for autonomous engineering agents across diverse modern software stacks.
System Architecture¶
The following diagram illustrates SWE-bench's dockerized execution sequence, from task instance ingestion and FastMCP 3.1 agent tool interactions to patch application and automated test suite evaluation:
sequenceDiagram
autonumber
participant Harness as SWE-bench Evaluation Harness
participant Agent as Autonomous Agent (FastMCP 3.1)
participant Container as Isolated Docker Container
participant TestSuite as Repository PyTest / Cargo Test Suite
participant Verifier as Pydantic v2 Patch Verifier
Harness->>Container: Spawn repository workspace (base commit SHA)
Harness->>Agent: Send GitHub issue problem statement
loop Search & Edit Cycle
Agent->>Container: Execute CLI search / FastMCP file edits
Container-->>Agent: Return file content / terminal stdout
end
Agent->>Verifier: Output unified git diff patch proposal
Verifier->>Verifier: Validate diff syntax & schema via Pydantic v2
Verifier->>Container: Apply unified patch proposal
Container->>TestSuite: Run test suite (FAIL_TO_PASS & PASS_TO_PASS)
TestSuite-->>Harness: Return unit test pass/fail report
Harness->>Harness: Compute instance resolution score
What problem it solves¶
Measures whether LLMs can perform practical software engineering work—understanding codebases, diagnosing issues, and producing working fixes—rather than just solving isolated coding puzzles. It identifies "stalling" behaviors and evaluates the robustness of agentic loops in a terminal environment, often leveraging FastMCP 3.1 (Model Context Protocol) for dynamic tool discovery and execution.
Where it fits in the stack¶
Benchmarking / Eval. It is used as a reference benchmark for evaluating real-world software engineering capabilities of AI agents and coding assistants.
Typical use cases¶
- Evaluating AI coding agents on their ability to resolve real GitHub issues.
- Comparing models on practical software engineering tasks.
- Tracking progress of AI agents toward autonomous software development.
- Choosing whether an agent is ready for repository-maintenance work that requires reading tests, editing files, and producing a valid patch.
Strengths¶
- Authenticity: Based on real-world GitHub issues, providing high-fidelity evaluation.
- End-to-End Skill Assessment: Requires reading code, understanding issues, and writing functional patches.
- Objective Validation: Results are verified against existing test suites from the source repositories.
- Large Dataset: Over 2,000 tasks across multiple popular repositories and programming languages.
Limitations¶
- Computational Cost: Requires a Docker environment and significant compute to run full test suites.
- Language Bias: Historically focused on Python repositories in the standard dataset, although this is now resolved by the SWE-bench Multilingual dataset covering 9 major programming languages.
- Static Nature: While updated, older subsets can suffer from data contamination in newer model training sets.
- Limited Scope: Does not typically evaluate documentation-only changes or complex multi-repository dependencies.
When to use it¶
- When evaluating AI agents or LLMs on real-world software engineering capability.
- When comparing coding agents that claim to autonomously resolve issues.
- When performing high-signal regression testing on agentic coding frameworks.
When not to use it¶
- When evaluating basic code generation from simple specifications (use HumanEval instead).
- When you need quick, lightweight benchmarking without Docker infrastructure.
- For testing non-code reasoning (use MMLU or GPQA).
Getting started¶
SWE-bench requires a Docker environment to safely execute untrusted code and run test suites.
1. Installation¶
pip install swebench
2. Basic Inference¶
To get started, you can run an inference pass using a lightweight model or a specific subset:
python -m swebench.inference.run_api \
--dataset_name princeton-nlp/SWE-bench_Lite \
--model_name claude-5-1-sonnet-20261022 \
--output_dir ./predictions
CLI examples¶
The following commands demonstrate how to interact with the SWE-bench evaluation harness.
# Install the SWE-bench package from source for the latest updates
pip install git+https://github.com/princeton-nlp/SWE-bench.git
# Run predictions for the 'Verified' subset using a local model endpoint
python -m swebench.inference.run_api --dataset_name princeton-nlp/SWE-bench_Verified --output_dir ./eval_results
# Execute evaluation using the Docker harness to verify generated patches
docker run -v $(pwd)/predictions:/predictions swebench/swe-bench-eval --predictions /predictions/predictions.jsonl --output_dir /results
API examples¶
Programmatic Prediction Validation using Pydantic v2 & FastMCP Server Pattern¶
Below is a complete FastMCP 3.1 server implementation that validates SWE-bench patch proposals and evaluates task instances using Pydantic v2 schemas:
import json
from typing import Optional, List, Dict
from pydantic import BaseModel, Field, ValidationError, field_validator
from mcp.server.fastmcp import FastMCP
# Initialize FastMCP 3.1 Server for SWE-bench Evaluation
mcp = FastMCP("SWE-bench-Evaluation-Server", version="3.1")
class SWEBenchPrediction(BaseModel):
instance_id: str = Field(..., description="The unique SWE-bench task identifier (e.g. django__django-12345)")
model_name: str = Field(..., description="Name of the model generating the patch")
patch: str = Field(..., description="The generated git diff patch proposing the bugfix")
explanation: Optional[str] = Field(None, description="Optional reasoning chain leading to this fix")
tokens_used: Optional[int] = Field(None, ge=0, description="Inference token count consumed")
@field_validator("patch")
@classmethod
def validate_is_git_patch(cls, value: str) -> str:
if value and not value.startswith("diff --git"):
raise ValueError("Patch must be a valid unified diff starting with 'diff --git'")
return value
class SWEBenchResult(BaseModel):
instance_id: str = Field(..., description="Task ID evaluated")
resolved: bool = Field(..., description="Whether all FAIL_TO_PASS tests passed and PASS_TO_PASS tests remained green")
test_stdout: str = Field(..., description="Execution logs from PyTest or Cargo test suite")
@mcp.tool(name="verify_swe_bench_patch", description="Validates a proposed patch diff against SWE-bench schema guidelines.")
def verify_swe_bench_patch(raw_prediction_json: str) -> str:
"""Parses raw model output JSON and verifies patch formatting."""
try:
data = json.loads(raw_prediction_json)
pred = SWEBenchPrediction.model_validate(data)
return json.dumps({
"status": "VALID",
"instance_id": pred.instance_id,
"patch_length": len(pred.patch)
}, indent=2)
except json.JSONDecodeError:
return json.dumps({"error": "Invalid JSON string"})
except ValidationError as e:
return json.dumps({"error": "Schema validation failed", "details": e.errors()})
if __name__ == "__main__":
sample_prediction = """
{
"instance_id": "django__django-12345",
"model_name": "claude-5-1-sonnet-20261022",
"patch": "diff --git a/django/db/models/fields/__init__.py b/django/db/models/fields/__init__.py\\n--- a/django/db/models/fields/__init__.py\\n+++ b/django/db/models/fields/__init__.py\\n@@ -1,1 +1,2 @@\\n",
"tokens_used": 14205
}
"""
print(verify_swe_bench_patch(sample_prediction))
Related tools / concepts¶
- HumanEval — Basic code generation benchmark.
- LongCLI-Bench — Long-horizon CLI task evaluation.
- DREAM: Deep Research Evaluation with Agentic Metrics — Agentic research evaluation.
- Aider — Terminal-based AI coding assistant.
- Claude Code — Anthropic's official coding agent.
- OpenHands — Open-source agentic platform.
- Terminal-Bench — Evaluating tool use in CLI environments.
- Benchmarking — Overview of AI evaluation frameworks.
- Model Context Protocol — Standard for agentic tool discovery.
Sources / references¶
- Official Website
- GitHub Repository
- SWE-bench Paper (arXiv:2310.06770)
- SWE-bench Verified Announcement
- SWE-bench Multilingual Release Discussion - Reddit
- SWE-bench Multilingual Official Documentation
Contribution Metadata¶
- Last reviewed: 2027-01-07
- Confidence: high