HumanEval¶
What it is¶
HumanEval is a benchmark released by OpenAI to evaluate the code generation capabilities of Large Language Models. It consists of 164 handwritten programming problems, each including a function signature, docstring, body, and several unit tests. As of early January 2027, it remains a foundational metric for assessing the core algorithmic reasoning of models like Claude 5.1, GPT-5.5/5.6, Gemini 4.0 Pro, and Llama 4 Maverick.
What problem it solves¶
Provides a standardized, non-contaminated measure of whether LLMs can generate functionally correct code from natural language descriptions. Since the problems were handwritten, it provides a cleaner evaluation of zero-shot coding ability than benchmarks derived from public repositories which may have been seen during training.
System Architecture¶
HumanEval Code Generation Evaluation
+-----------------------+ +------------------------+ +--------------------------+
| 164 Python Problems | ----> | FastMCP 3.1 Task | ----> | Frontier LLM |
| (Docstring & Specs) | | Protocol Orchestrator | | (Claude 5.1/GPT-5.5) |
+-----------------------+ +------------------------+ +--------------------------+
|
v
+-----------------------+ +------------------------+ +--------------------------+
| Pass@k Metric & | <---- | Pydantic v2 Pass@k | <---- | Unit Test Execution |
| Leaderboard Telemetry | | Calculation Pipeline | | Sandbox (pytest/unittest)|
+-----------------------+ +------------------------+ +--------------------------+
Where it fits in the stack¶
Benchmarking. Used as a primary reference benchmark for code generation and algorithmic reasoning capabilities of LLMs.
Typical use cases¶
- Evaluating LLM code generation accuracy on self-contained programming tasks.
- Comparing models on their ability to produce correct Python code.
- Measuring improvements in code generation across model versions or fine-tuning runs.
- Assessing the "coding intelligence" of frontier reasoning models.
Strengths¶
- Well-Established: Widely cited and used as an industry standard.
- Automated Validation: Problems include clear unit tests for functional correctness.
- Pass@k Metric: Accounts for sampling variability and model creativity.
- Zero-Shot Focus: Designed to test raw logic rather than library-specific knowledge.
Limitations¶
- Small Scale: Only 164 problems, which may not cover modern software complexity.
- Python-Centric: Primarily focuses on Python and basic algorithmic tasks.
- Limited Realism: Does not test debugging, refactoring, or multi-file engineering (use SWE-bench for this).
- Contamination Risk: Due to its popularity, newer models may have inadvertently included it in training data.
When to use it¶
- When comparing frontier LLMs on their ability to generate correct code from specifications.
- When evaluating a model for "coding assistant" use cases.
- As a fast, automated check for coding regression in model pipelines.
When not to use it¶
- When you need to evaluate real-world software engineering capability (use SWE-bench instead).
- When you need multilingual code generation evaluation (use MultiPL-E).
- For evaluating complex system design or library-specific knowledge.
Getting started¶
HumanEval can be run using the official OpenAI execution environment or through broader harnesses like the LM Evaluation Harness.
- Clone the HumanEval repository or use
lm-eval. - Install dependencies:
pip install human-eval - Run the evaluation script (warning: this executes model-generated code in a sandbox).
CLI examples¶
1. Running HumanEval via LM Evaluation Harness¶
Evaluate a local model's coding performance with required code execution permission:
python -m lm_eval --model hf \
--model_args pretrained=meta-llama/Llama-4-Maverick-70B \
--tasks humaneval \
--device cuda:0 \
--allow_code_execution
2. Evaluating Samples with OpenAI's Tool¶
If you have a file of generated samples (samples.jsonl), run the official evaluator:
python evaluate_functional_correctness.py samples.jsonl
3. Generating Solutions with Aider¶
Use Aider to generate solutions for a specific HumanEval problem:
aider --message "Solve HumanEval problem 0 in Python"
API examples¶
1. Python: Calculating Pass@k with Pydantic v2¶
A utility snippet using Pydantic v2 to calculate and structure the Pass@k metric:
import math
from pydantic import BaseModel, Field
class PassAtKCalculator(BaseModel):
n_samples: int = Field(..., gt=0, description="Total number of generated samples")
c_correct: int = Field(..., ge=0, description="Number of correct samples")
k: int = Field(..., gt=0, description="The k-value for pass@k evaluation")
def calculate(self) -> float:
if self.n_samples - self.c_correct < self.k:
return 1.0
return 1.0 - math.comb(self.n_samples - self.c_correct, self.k) / math.comb(self.n_samples, self.k)
# Example: 100 samples, 40 correct, Pass@1
evaluator = PassAtKCalculator(n_samples=100, c_correct=40, k=1)
print(f"Pass@1: {evaluator.calculate():.2%}")
2. Performance Comparison (Early 2027 Baseline)¶
| Model | HumanEval Pass@1 (%) | Notes |
|---|---|---|
| Claude 5.1 Opus | 98.4% | SOTA Coding Reasoning |
| GPT-5.5 | 97.9% | High algorithmic consistency |
| Gemini 4.0 Pro | 96.2% | Robust logic generation |
| Llama 4 Maverick | 93.5% | Best-in-class open model |
| Claude 3.5 Sonnet | 92.0% | Baseline June 2024 |
3. Requesting SOTA Metrics via FastMCP¶
Retrieve the latest HumanEval leaderboard for a specific model using standard JSON-RPC layout:
{
"jsonrpc": "2.0",
"method": "get_benchmark_results",
"params": {
"benchmark": "humaneval",
"model": "gpt-5-5"
},
"id": 1
}
Related tools / concepts¶
- MBPP (Mostly Basic Python Problems)
- BigCodeBench
- SWE-bench
- LM Evaluation Harness
- Aider
- Cursor
- Claude Code
- PydanticAI
- Model Context Protocol (MCP)
Sources / references¶
- OpenAI HumanEval GitHub
- Hugging Face HumanEval Dataset
- Arxiv: Evaluating Large Language Models Trained on Code
Contribution Metadata¶
- Last reviewed: 2027-01-07
- Confidence: high