Humanity's Last Exam (HLE)¶
What it is¶
HLE is a benchmark designed to test the limits of LLMs on the most difficult human-level tasks. It consists of 3,000 highly complex, multi-disciplinary questions across over a hundred subjects (Mathematics, Physics, Biology, Humanities, etc.). Created by the Center for AI Safety (CAIS) and Scale AI, it represents a "frontier" benchmark where early January 2027 state-of-the-art models like Claude 5.6, GPT-5.6, Gemini 4.0 Ultra, Gemma 4, DeepSeek-V4, and Qwen 3.6 VL still perform poorly on the hardest subsets.
Evaluation Pipeline Architecture¶
graph TD
A[CAIS / Scale AI HLE Dataset] -->|Question Ingestion| B[FastMCP 3.1 Task Protocol Server]
B -->|Multimodal Ingestion ColQwen| C[Vision & Math Parser]
C -->|Prompts & Inputs| D[Frontier Model Under Test]
D -->|Candidate Response| E[LLM Judge / Scorer]
E -->|Automated Equivalence Verification| F[Pydantic v2 Validated Score]
F -->|Telemetry Logging| G[HLE Benchmark Leaderboard]
What problem it solves¶
Addresses the "saturation" of existing benchmarks like MMLU and GPQA. As frontier models reach or exceed human-level performance on older tests, those tests lose their utility as measurement tools. HLE provides a new ceiling for frontier reasoning research, ensuring that progress toward expert-level agentic intelligence remains measurable.
Where it fits in the stack¶
Benchmarking. Serves as a high-difficulty knowledge and reasoning benchmark for evaluating the upper limits of LLM and multi-modal model capabilities within agentic ingestion pipelines.
Typical use cases¶
- Frontier Model Evaluation: Comparing the reasoning capabilities of state-of-the-art models (Claude 5.6, GPT-5.6, Gemini 4.0 Ultra, Gemma 4, DeepSeek-V4, Qwen 3.6 VL).
- Multi-modal Assessment: Testing models on questions that require both textual reasoning and image understanding (14% of the dataset is multi-modal, evaluated using ColQwen).
- Calibration Testing: Measuring whether models accurately estimate their own confidence in their answers.
- Agentic Pre-training Validation: Verifying that new pre-training runs have significantly moved the needle on expert-level reasoning.
Strengths¶
- Extreme Difficulty: Designed to be the "last academic exam," remaining challenging even as frontier models rapidly evolve.
- Closed-ended & Verifiable: Answers are precise, allowing for automated, low-cost evaluation via agentic satisfaction loops and FastMCP 3.1 Task Protocol.
- Subject Diversity: Covers over 100 subjects with questions sourced from world-class experts.
- Private Set: Includes a held-out private set with regular rotation (v2 canary splits) to combat data contamination and benchmark hacking.
Limitations¶
- Not for Everyday Tasks: Does not measure "helpful assistant" capabilities or basic instruction following.
- Low Signal for Small Models: Smaller or older models often score near zero, making it difficult to distinguish between them.
- Requires LLM Judge: While answers are closed-ended, the variety of possible formats (decimals vs. fractions) often requires an LLM judge for automated scoring at scale.
When to use it¶
- When evaluating frontier models on the hardest available reasoning tasks in 2027.
- When existing benchmarks like MMLU or GPQA show signs of saturation (models scoring >90%).
- When testing a model's ability to handle world-class scientific or mathematical problems for agentic research.
When not to use it¶
- When evaluating models for general-purpose chat or basic RAG tasks.
- When you need a lightweight, fast-running benchmark for early-stage development.
- When you are optimizing for speed or low-cost inference rather than peak intelligence.
Getting started¶
HLE is typically executed via the UK Government's Inspect framework or the LM Evaluation Harness.
1. Environment Setup¶
# Install inspect and the evals package
pip install inspect-ai inspect_evals
# Configure API keys for frontier models
export ANTHROPIC_API_KEY="sk-ant-..."
export OPENAI_API_KEY="sk-proj-..."
2. Running the Benchmark¶
# Run HLE against a Claude 5.6 model
inspect eval inspect_evals/hle --model anthropic/claude-5-6-opus
CLI examples¶
Evaluation via Inspect CLI with concurrency controls¶
Evaluate a specific subject within HLE with adjusted concurrency:
inspect eval inspect_evals/hle \
--model openai/gpt-5-6 \
--limit 100 \
--subject "quantum_physics" \
--concurrency 10 \
--max-connections 5
Running with LM Evaluation Harness¶
HLE is also supported as a task in the standard harness:
lm_eval --model vllm \
--model_args pretrained=meta-llama/Llama-4-100b,tensor_parallel_size=4 \
--tasks hle \
--batch_size auto
API examples¶
FastMCP 3.1 HLE Ingestion & Evaluation Server Pattern¶
Exposing HLE benchmark runs over FastMCP 3.1 Task Protocol:
from mcp.server.fastmcp import FastMCP, Context
from pydantic import BaseModel, Field
from typing import Optional
mcp = FastMCP("HLE Frontier Evaluator Server")
class HLERunConfig(BaseModel):
model_identifier: str = Field(..., description="Frontier model to evaluate")
subject_filter: Optional[str] = Field(None, description="Specific academic subject or None for full suite")
sample_limit: int = Field(50, ge=1, le=3000)
class HLEResultSummary(BaseModel):
model_identifier: str
total_processed: int
accuracy_percentage: float
@mcp.tool()
async def execute_hle_eval(config: HLERunConfig, ctx: Context) -> HLEResultSummary:
"""Executes HLE benchmark subset over FastMCP 3.1 task protocol."""
ctx.info(f"Triggering HLE benchmark for {config.model_identifier} on subject {config.subject_filter}")
# Execution simulation
return HLEResultSummary(
model_identifier=config.model_identifier,
total_processed=config.sample_limit,
accuracy_percentage=24.5 # Early 2027 SOTA baseline on HLE
)
if __name__ == "__main__":
mcp.run()
Pydantic v2 Integration & Evaluator Verification¶
Below is a production-ready example of evaluating and validating HLE evaluation metadata programmatically with Pydantic v2 and FastMCP 3.1 Task Protocol before logging results into multi-agent databases.
from pydantic import BaseModel, Field, ValidationError
from typing import List, Optional
from datetime import datetime
class HLEQuestion(BaseModel):
question_id: str = Field(..., description="Unique ID of the HLE question")
subject: str = Field(..., description="Academic field (e.g., Quantum Physics, Topology)")
difficulty: int = Field(5, description="Expert difficulty tier from 1-5", ge=1, le=5)
has_image: bool = Field(False, description="Whether the question is multi-modal and requires a vision-native parser like ColQwen")
correct_answer: str = Field(..., description="Ground truth answer string")
class HLEEvaluationResult(BaseModel):
model_name: str = Field(..., description="Name of the model under evaluation")
evaluation_date: datetime = Field(default_factory=datetime.utcnow, description="When the evaluation was performed")
total_evaluated: int = Field(..., description="Total questions processed", ge=0)
accuracy: float = Field(..., description="Pass rate accuracy", ge=0.0, le=1.0)
validated_questions: List[HLEQuestion] = Field(default_factory=list, description="Validated subset of questions run")
# Execute a programmatic validation of an HLE score
def validate_and_log_hle_run(run_payload: dict) -> Optional[HLEEvaluationResult]:
try:
# Strict schema validation using Pydantic v2 model_validate
validated_eval = HLEEvaluationResult.model_validate(run_payload)
print(f"Validated HLE Run: {validated_eval.model_name} achieved {validated_eval.accuracy:.2%} accuracy.")
return validated_eval
except ValidationError as e:
print(f"HLE payload verification failed: {e.errors()}")
return None
# Test the function with simulated early 2027 run data
sample_eval = {
"model_name": "claude-5-6-opus",
"total_evaluated": 1,
"accuracy": 1.0,
"validated_questions": [
{
"question_id": "hle-math-40291",
"subject": "Algebraic Topology",
"difficulty": 5,
"has_image": False,
"correct_answer": "Z_2"
}
]
}
validated_run = validate_and_log_hle_run(sample_eval)
Python Integration (Inspect AI)¶
Automate HLE evaluation within a research pipeline:
from inspect_ai import eval, Epochs
from inspect_evals.hle import hle
# Run evaluation programmatically with custom epoch settings and model arguments
results = eval(
tasks=hle(),
model="anthropic/claude-5-6-sonnet",
limit=50,
epochs=Epochs(3, "at_least_once"),
model_args={
"temperature": 0.0,
"max_tokens": 4096
}
)
# Access scores
print(f"HLE Accuracy: {results[0].metrics['accuracy'].value}")
Custom Scorer Example¶
Using an LLM judge to verify HLE responses:
from inspect_ai.scorer import scorer, ScorerResult
from inspect_ai.solver import TaskState
@scorer(metrics=["accuracy"])
def hle_judge():
async def score(state: TaskState, target: str):
# Use a frontier model as a judge for complex formats
is_correct = await judge_response(state.output.completion, target)
return ScorerResult(
value="CORRECT" if is_correct else "INCORRECT",
answer=state.output.completion
)
return score
Related tools / concepts¶
- GPQA - Graduate-level Google-proof Q&A.
- MMLU - Massive Multitask Language Understanding.
- ARC (AI2 Reasoning Challenge) - Challenging questions for reasoning.
- GSM8K - Grade school math word problems.
- Chatbot Arena - Crowdsourced ELO ratings for LLMs.
- SWE-bench - Software engineering benchmark for agents.
- LM Evaluation Harness - Unified framework for running multiple benchmarks.
- ColQwen - Vision-native document parsing for multi-modal HLE tasks.
- DeepSeek-V4 - Reasoning benchmark leader in early 2027.
- Terminus 2 - Terminal-based reasoning benchmark.
Sources / references¶
- Humanity's Last Exam - Official Site
- Scale Labs Leaderboard
- Humanity's Last Exam - arXiv Paper (2025)
- Inspect AI Documentation
Contribution Metadata¶
- Last reviewed: 2027-01-07
- Confidence: high