Skip to content

SWE-bench

What it is

SWE-bench is a benchmark for evaluating LLMs on real-world software engineering tasks. It uses actual issues from GitHub and requires the model to generate a functional patch that passes existing tests. As of January 2027, it remains the industry standard for measuring the autonomous coding capabilities of frontier models like Claude 5.1, GPT-5.5 / GPT-5.6, Gemini 4.0 Pro / Ultra, and DeepSeek-V4.

With the launch of SWE-bench Multilingual, the benchmark has been expanded to support a wider array of programming languages (C, C++, Go, Java, JavaScript, TypeScript, PHP, Ruby, and Rust), making it a truly language-agnostic evaluator for autonomous engineering agents across diverse modern software stacks.

System Architecture

The following diagram illustrates SWE-bench's dockerized execution sequence, from task instance ingestion and FastMCP 3.1 agent tool interactions to patch application and automated test suite evaluation:

sequenceDiagram
    autonumber
    participant Harness as SWE-bench Evaluation Harness
    participant Agent as Autonomous Agent (FastMCP 3.1)
    participant Container as Isolated Docker Container
    participant TestSuite as Repository PyTest / Cargo Test Suite
    participant Verifier as Pydantic v2 Patch Verifier

    Harness->>Container: Spawn repository workspace (base commit SHA)
    Harness->>Agent: Send GitHub issue problem statement
    loop Search & Edit Cycle
        Agent->>Container: Execute CLI search / FastMCP file edits
        Container-->>Agent: Return file content / terminal stdout
    end
    Agent->>Verifier: Output unified git diff patch proposal
    Verifier->>Verifier: Validate diff syntax & schema via Pydantic v2
    Verifier->>Container: Apply unified patch proposal
    Container->>TestSuite: Run test suite (FAIL_TO_PASS & PASS_TO_PASS)
    TestSuite-->>Harness: Return unit test pass/fail report
    Harness->>Harness: Compute instance resolution score

What problem it solves

Measures whether LLMs can perform practical software engineering work—understanding codebases, diagnosing issues, and producing working fixes—rather than just solving isolated coding puzzles. It identifies "stalling" behaviors and evaluates the robustness of agentic loops in a terminal environment, often leveraging FastMCP 3.1 (Model Context Protocol) for dynamic tool discovery and execution.

Where it fits in the stack

Benchmarking / Eval. It is used as a reference benchmark for evaluating real-world software engineering capabilities of AI agents and coding assistants.

Typical use cases

  • Evaluating AI coding agents on their ability to resolve real GitHub issues.
  • Comparing models on practical software engineering tasks.
  • Tracking progress of AI agents toward autonomous software development.
  • Choosing whether an agent is ready for repository-maintenance work that requires reading tests, editing files, and producing a valid patch.

Strengths

  • Authenticity: Based on real-world GitHub issues, providing high-fidelity evaluation.
  • End-to-End Skill Assessment: Requires reading code, understanding issues, and writing functional patches.
  • Objective Validation: Results are verified against existing test suites from the source repositories.
  • Large Dataset: Over 2,000 tasks across multiple popular repositories and programming languages.

Limitations

  • Computational Cost: Requires a Docker environment and significant compute to run full test suites.
  • Language Bias: Historically focused on Python repositories in the standard dataset, although this is now resolved by the SWE-bench Multilingual dataset covering 9 major programming languages.
  • Static Nature: While updated, older subsets can suffer from data contamination in newer model training sets.
  • Limited Scope: Does not typically evaluate documentation-only changes or complex multi-repository dependencies.

When to use it

  • When evaluating AI agents or LLMs on real-world software engineering capability.
  • When comparing coding agents that claim to autonomously resolve issues.
  • When performing high-signal regression testing on agentic coding frameworks.

When not to use it

  • When evaluating basic code generation from simple specifications (use HumanEval instead).
  • When you need quick, lightweight benchmarking without Docker infrastructure.
  • For testing non-code reasoning (use MMLU or GPQA).

Getting started

SWE-bench requires a Docker environment to safely execute untrusted code and run test suites.

1. Installation

pip install swebench

2. Basic Inference

To get started, you can run an inference pass using a lightweight model or a specific subset:

python -m swebench.inference.run_api \
    --dataset_name princeton-nlp/SWE-bench_Lite \
    --model_name claude-5-1-sonnet-20261022 \
    --output_dir ./predictions

CLI examples

The following commands demonstrate how to interact with the SWE-bench evaluation harness.

# Install the SWE-bench package from source for the latest updates
pip install git+https://github.com/princeton-nlp/SWE-bench.git

# Run predictions for the 'Verified' subset using a local model endpoint
python -m swebench.inference.run_api --dataset_name princeton-nlp/SWE-bench_Verified --output_dir ./eval_results

# Execute evaluation using the Docker harness to verify generated patches
docker run -v $(pwd)/predictions:/predictions swebench/swe-bench-eval --predictions /predictions/predictions.jsonl --output_dir /results

API examples

Programmatic Prediction Validation using Pydantic v2 & FastMCP Server Pattern

Below is a complete FastMCP 3.1 server implementation that validates SWE-bench patch proposals and evaluates task instances using Pydantic v2 schemas:

import json
from typing import Optional, List, Dict
from pydantic import BaseModel, Field, ValidationError, field_validator
from mcp.server.fastmcp import FastMCP

# Initialize FastMCP 3.1 Server for SWE-bench Evaluation
mcp = FastMCP("SWE-bench-Evaluation-Server", version="3.1")

class SWEBenchPrediction(BaseModel):
    instance_id: str = Field(..., description="The unique SWE-bench task identifier (e.g. django__django-12345)")
    model_name: str = Field(..., description="Name of the model generating the patch")
    patch: str = Field(..., description="The generated git diff patch proposing the bugfix")
    explanation: Optional[str] = Field(None, description="Optional reasoning chain leading to this fix")
    tokens_used: Optional[int] = Field(None, ge=0, description="Inference token count consumed")

    @field_validator("patch")
    @classmethod
    def validate_is_git_patch(cls, value: str) -> str:
        if value and not value.startswith("diff --git"):
            raise ValueError("Patch must be a valid unified diff starting with 'diff --git'")
        return value

class SWEBenchResult(BaseModel):
    instance_id: str = Field(..., description="Task ID evaluated")
    resolved: bool = Field(..., description="Whether all FAIL_TO_PASS tests passed and PASS_TO_PASS tests remained green")
    test_stdout: str = Field(..., description="Execution logs from PyTest or Cargo test suite")

@mcp.tool(name="verify_swe_bench_patch", description="Validates a proposed patch diff against SWE-bench schema guidelines.")
def verify_swe_bench_patch(raw_prediction_json: str) -> str:
    """Parses raw model output JSON and verifies patch formatting."""
    try:
        data = json.loads(raw_prediction_json)
        pred = SWEBenchPrediction.model_validate(data)
        return json.dumps({
            "status": "VALID",
            "instance_id": pred.instance_id,
            "patch_length": len(pred.patch)
        }, indent=2)
    except json.JSONDecodeError:
        return json.dumps({"error": "Invalid JSON string"})
    except ValidationError as e:
        return json.dumps({"error": "Schema validation failed", "details": e.errors()})

if __name__ == "__main__":
    sample_prediction = """
    {
        "instance_id": "django__django-12345",
        "model_name": "claude-5-1-sonnet-20261022",
        "patch": "diff --git a/django/db/models/fields/__init__.py b/django/db/models/fields/__init__.py\\n--- a/django/db/models/fields/__init__.py\\n+++ b/django/db/models/fields/__init__.py\\n@@ -1,1 +1,2 @@\\n",
        "tokens_used": 14205
    }
    """
    print(verify_swe_bench_patch(sample_prediction))

Sources / references

Contribution Metadata

  • Last reviewed: 2027-01-07
  • Confidence: high