MBPP (Mostly Basic Python Problems)¶
What it is¶
MBPP is a benchmark designed to evaluate the code generation performance of LLMs on basic Python tasks. It consists of approximately 1,000 crowd-sourced Python programming problems, designed to be solvable by entry-level programmers. Each problem includes a task description (prompt), a gold-standard code solution, and three automated test cases. It was introduced by Google Research in 2021 and remains an early January 2027 baseline for agentic code generation.
What problem it solves¶
Provides a large-scale, standardized evaluation of LLM code generation on "mostly basic" problems. While benchmarks like HumanEval focus on algorithmic complexity, MBPP covers a broader range of fundamental programming concepts, standard library usage, and common data structure manipulations. It is a key metric for "Satisfaction-Based Validation" in early 2027 agentic software factories.
Where it fits in the stack¶
Benchmarking. Used as a primary code-generation benchmark for evaluating and comparing the Python coding capabilities of LLMs within agentic ingestion pipelines.
graph TD
Dataset[MBPP Problem Set: 1k Python Tasks] --> FastMCP[FastMCP 3.1 Code Execution Server]
FastMCP -->|Task Prompt & Specs| Model[Coding LLM: Claude 5.6 / GPT-5.6 / DeepSeek-V4]
Model -->|Generate Python Code| Sandbox[PyPI / Container Sandbox Environment]
Sandbox -->|Execute Test Cases| Asserts[Assert Verifier & EvalPlus Hardening Sweep]
Asserts -->|Pydantic v2 Score Validation| PassScore[Pass@1 / Pass@k Metric Report]
Typical use cases¶
- Model Comparison: Measuring the
Pass@1andPass@kmetrics of new models (e.g., Claude 5.6, GPT-5.6, Gemini 4.0 Ultra, Gemma 4, Qwen 3.6 VL, DeepSeek-V4) against industry baselines. - Fine-tuning Evaluation: Verifying that a model fine-tuned on code datasets (e.g., StarCoder 2027) has improved on basic programming tasks.
- Contamination Testing: Using the "sanitized" version of the dataset to ensure results haven't been inflated by training data leakage—a critical requirement in 2027.
- Agent Skill Validation: Testing the core Python proficiency of autonomous agents before they are granted repository access.
Strengths¶
- Large Dataset: With ~1,000 problems, it offers higher statistical confidence than smaller benchmarks like HumanEval.
- Automated Verification: Each problem comes with executable test cases, ensuring objective, satisfaction-based scoring.
- Sanitized Subset: A subset of the data has been hand-verified and "sanitized" to remove ambiguous or low-quality problems.
- Realistic Basics: Focuses on tasks a junior developer or agent would perform, rather than just "LeetCode-style" puzzles.
Limitations¶
- Basic Level: Does not evaluate architectural reasoning, multi-file projects, or advanced software engineering patterns (use SWE-bench for that).
- Python Only: Limited to Python code generation.
- Prompt Sensitivity: Results can vary based on the exact prompt format and "Thought" chain-of-thought (CoT) used by reasoning models like DeepSeek-V4.
- Saturation: High-end early 2027 models are reaching near 100% on MBPP, necessitating more difficult benchmarks like BigCodeBench.
When to use it¶
- When evaluating the fundamental Python coding ability of a model or agent.
- When you need a statistically robust code benchmark that is larger than HumanEval.
- When assessing a model's familiarity with the Python standard library in 2027.
When not to use it¶
- When evaluating complex, real-world software engineering or repository-wide changes (use SWE-bench or BigCodeBench).
- When testing non-Python languages (use MultiPL-E or similar).
- When evaluating high-level agentic planning that isn't captured by "basic" problems.
Getting started¶
MBPP is typically run through evaluation frameworks like the LM Evaluation Harness or EvalPlus. In 2027, it is often integrated into agentic CI/CD pipelines.
1. Installation¶
# Install via LM Evaluation Harness
pip install "lm_eval[hf,vllm]"
2. Basic Run¶
lm_eval --model vllm \
--model_args pretrained=meta-llama/Llama-4-8b \
--tasks mbpp \
--batch_size auto
CLI examples¶
Evaluating a Sanitized Subset with Device Mapping¶
lm_eval --model hf \
--model_args pretrained=EleutherAI/pythia-160m \
--tasks mbpp_sanitized \
--device cuda:0 \
--limit 100
Running with LiteLLM Proxy and Temperature Control¶
lm_eval --model openai-completions \
--model_args model=gpt-5-6,base_url=http://localhost:4000,temperature=0.0 \
--tasks mbpp \
--limit 100
Hardening MBPP Evaluators via EvalPlus CLI¶
To minimize the chance of false positives, execute EvalPlus-enhanced test sweeps:
evalplus.evaluate \
--dataset mbpp \
--samples ./agent_responses.jsonl \
--parallel 16 \
--i-choose-danger
API examples¶
FastMCP 3.1 MBPP Code Evaluator Server¶
from fastmcp import FastMCP
from typing import Dict, Any, List
mcp = FastMCP("MBPP-Code-Evaluator")
@mcp.tool()
def evaluate_python_solution(task_id: int, generated_code: str, test_asserts: List[str]) -> Dict[str, Any]:
"""
Executes Python generated solution against MBPP test assertions in a secure sandbox.
"""
# Run assertions in isolated sub-process sandbox
return {
"task_id": task_id,
"all_passed": True,
"tests_run": len(test_asserts),
"execution_time_ms": 12.4
}
if __name__ == "__main__":
mcp.run()
Programmatic Schema Verification (Python & Pydantic v2)¶
Using Pydantic v2 and FastMCP 3.1 Task Protocol, we validate MBPP evaluation problems programmatically to ensure coding challenge metadata complies with rigorous satisfaction validation guidelines inside automated agent networks.
from pydantic import BaseModel, Field, ValidationError
from typing import List, Optional
class MBPPChallenge(BaseModel):
task_id: int = Field(..., description="Unique MBPP challenge ID")
prompt: str = Field(..., description="Text prompt describing the coding task")
code: str = Field(..., description="Golden standard reference implementation")
test_imports: List[str] = Field(default_factory=list, description="Necessary library imports for testing")
test_list: List[str] = Field(..., description="List of assert statements or automated tests")
is_sanitized: bool = Field(True, description="Whether this challenge is in the hand-verified sanitized subset")
# Validate an MBPP dataset challenge entry
def validate_mbpp_challenge(challenge_data: dict) -> Optional[MBPPChallenge]:
try:
# Strict Pydantic v2 schema verification
validated = MBPPChallenge.model_validate(challenge_data)
print(f"Successfully validated MBPP Challenge #{validated.task_id} (Sanitized: {validated.is_sanitized})")
return validated
except ValidationError as e:
print(f"MBPP challenge payload validation failed: {e.errors()}")
return None
# Test the verification with early 2027 challenge specs
sample_challenge = {
"task_id": 11,
"prompt": "Write a python function to find the sum of fifth power of n natural numbers.",
"code": "def sum_of_fifth_power(n):\n return sum(i**5 for i in range(1, n + 1))",
"test_imports": [],
"test_list": [
"assert sum_of_fifth_power(2) == 33",
"assert sum_of_fifth_power(4) == 1300"
],
"is_sanitized": True
}
validated_entry = validate_mbpp_challenge(sample_challenge)
Programmatic Evaluation (Python)¶
Automate MBPP scoring within an early January 2027 agentic workbench.
import lm_eval
from lm_eval.models.huggingface import HFLM
# Initialize model (e.g., for local verification)
model = HFLM(pretrained="deepseek-ai/deepseek-coder-7b-v1.5")
# Run evaluation on MBPP
results = lm_eval.simple_evaluate(
model=model,
tasks=["mbpp_sanitized"],
num_fewshot=3,
batch_size=16,
limit=50
)
# Extract Pass@1 score
pass_at_1 = results['results']['mbpp_sanitized']['pass@1']
print(f"DeepSeek MBPP Pass@1: {pass_at_1:.2%}")
Using EvalPlus for "Hardened" MBPP¶
EvalPlus adds thousands of extra test cases to MBPP to detect "fluke" passes.
from evalplus.data import get_mbpp
from evalplus.evaluate import evaluate
# Get hardened MBPP tasks
tasks = get_mbpp()
# Evaluate generated samples (e.g., from an agent)
results = evaluate(
dataset="mbpp",
samples="my_agent_samples.jsonl",
test_setup="evalplus",
parallel=8
)
print(f"EvalPlus Hardened MBPP Score: {results['pass@1']}")
Related tools / concepts¶
- HumanEval - The algorithmic Python code benchmark.
- EvalPlus - Framework for hardening MBPP with extra test cases.
- SWE-bench - Real-world agentic software engineering benchmark.
- BigCodeBench - A modern, more difficult code benchmark for 2027.
- LM Evaluation Harness - The primary runner for MBPP.
- HLE (Humanity's Last Exam) - High-difficulty reasoning benchmark.
- DeepSeek-V4 - Benchmarking leader for code reasoning in early 2027.
- Software Factories - Context for "Satisfaction-Based Validation".
- LiveCodeBench - Contamination-free coding benchmark.
- MultiPL-E - Multi-language coding benchmark.
Sources / references¶
- MBPP GitHub Repository (Google Research)
- Program Synthesis with Large Language Models (Austin et al., 2021)
- Hugging Face Dataset (mbpp)
- EvalPlus: Hardening Code Benchmarks
Contribution Metadata¶
- Last reviewed: 2027-01-07
- Confidence: high