DREAM: Deep Research Evaluation with Agentic Metrics¶
What it is¶
DREAM (Deep Research Evaluation with Agentic Metrics) is an agentic evaluation framework for deep research agents. It uses tool-calling agents to independently verify the factual correctness and temporal validity of research reports. As of early January 2027, it is the primary method for evaluating the "research depth" and multi-turn reasoning of frontier models like Claude 5.1, GPT-5.5 / 5.6, and Gemini 4.0 Pro.
What problem it solves¶
It addresses the "Mirage of Synthesis"—a defect in static LLM evaluation where fluent writing and plausible citations hide factual errors or reasoning flaws. Static judges cannot verify claims against real-world evidence; DREAM solves this by making the evaluator as capable (agentic) as the agent it is testing, utilizing FastMCP 3.1 for dynamic tool discovery and tool execution.
graph TD
AgentReport[Generated Research Report] -->|Extract Claims| ClaimExtractor[DREAM Claim Extractor]
ClaimExtractor -->|Unverified Claims| VerificationAgent[Agentic Evaluator Loop]
VerificationAgent -->|FastMCP 3.1 Protocol| SearchTools[Tavily / Web Search / FastMCP]
SearchTools -->|Live Web Evidence| VerificationAgent
VerificationAgent -->|Cross-Reference & Fact Check| VerdictEngine[Status & Decay Engine]
VerdictEngine -->|Output Verification Matrix| ResultReport[DREAM Evaluation Report]
Where it fits in the stack¶
Eval / Benchmarking: It is a framework for benchmarking and evaluating advanced LLM agentic performance, particularly for models when used in complex research loops. It bridges the gap between static benchmarks like GPQA and real-world utility.
Typical use cases¶
- Benchmarking Research Agents: Comparing how well different models or agent architectures (like OpenHands or custom research loops) generate accurate analyst-grade reports.
- Reasoning Probes: Systematically identifying reasoning defects in long-form generation.
- Fact-checking automation: Scaling the verification of AI-generated content against live web data using tools like Tavily.
- Temporal Verification: Checking if a report's data is still valid or has been superseded by newer events.
Strengths¶
- Parity-based Evaluation: Uses agents to evaluate agents, ensuring the evaluator has the tools necessary to verify modern information.
- Sensitivity to Decay: Significantly more sensitive to factual and temporal decay than static benchmarks like MMLU.
- Scalable and Reference-Free: Does not require a pre-defined ground truth for every query, allowing for flexible evaluation of open-ended research.
- Tool-Agnostic: Can be integrated with any tool-calling environment supporting FastMCP 3.1.
Limitations¶
- Operational Complexity: Requires a tool-calling environment for the evaluation agent, making it more complex to run than static Q&A.
- Cost: Agentic evaluation involves multiple LLM turns, increasing the cost of benchmarking.
- Latency: Evaluation can take several minutes per report, depending on the depth of the research required.
When to use it¶
- When evaluating agents that perform active research or use external tools.
- When static benchmarks are suspected of suffering from data contamination or lack of temporal awareness.
- To validate the factual integrity of long-form reports generated by Claude Code or similar tools.
When not to use it¶
- For evaluating simple base models on static knowledge where HumanEval is sufficient.
- When a fast, low-cost evaluation signal is needed for iterative model tuning.
- For non-research tasks like simple creative writing or translation.
Getting started¶
DREAM requires an environment where evaluation agents can execute search and browsing tools.
1. Installation¶
pip install dream-eval-framework
2. Environment Setup¶
Configure your LLM provider and search API keys. DREAM supports Claude 5.1 out of the box:
export ANTHROPIC_API_KEY="your-key"
export TAVILY_API_KEY="your-key"
3. Basic Verification¶
Run a verification check on a local report file:
dream-eval verify --report ./my_research_report.md
CLI examples¶
1. Running a Benchmark¶
Run the DREAM benchmark suite against a research agent endpoint:
dream-eval benchmark --agent-url "http://localhost:8080/chat" --tasks research_tasks_v1.json
2. Temporal Sensitivity Check¶
Check if a report is outdated by forcing the agent to prioritize recent news sources:
dream-eval verify --report report.md --temporal-weight 0.8
3. Exporting Results¶
Export the verification results to a structured JSON format for further analysis:
dream-eval verify --report report.md --output results.json
API examples¶
Basic Verification Loop (Python)¶
from dream_eval import DreamEvaluator
# Initialize the agentic evaluator with search tools via FastMCP 3.1
evaluator = DreamEvaluator(
model="claude-5-1-20261101",
mcp_servers=["https://search.api.tavily.com/mcp"]
)
report_content = "..." # The output from the research agent
# DREAM starts its independent verification
verification_results = evaluator.verify_claims(report_content)
for claim in verification_results.claims:
print(f"Claim: {claim.text}")
print(f"Status: {claim.verification_status}") # Verified | Refuted | Unverifiable
Programmatic Claim Validation using Pydantic v2¶
This Python script parses and validates agentic fact-checking results generated by the DREAM evaluation loop using Pydantic v2:
import json
from typing import List, Literal, Optional
from pydantic import BaseModel, Field, ValidationError, HttpUrl
class ClaimVerification(BaseModel):
claim_id: str = Field(..., description="Unique identifier for the parsed claim")
text: str = Field(..., description="The literal statement extracted from the research report")
status: Literal["Verified", "Refuted", "Unverifiable"] = Field(..., description="Verification status")
confidence: float = Field(..., ge=0.0, le=1.0, description="Confidence score of the verification agent")
sources: List[HttpUrl] = Field(default_factory=list, description="List of source URLs used to verify/refute the claim")
justification: str = Field(..., description="Reasoning or quotes from the sources justifying the decision")
class DreamEvaluationReport(BaseModel):
report_id: str = Field(..., description="ID of the research report being evaluated")
overall_accuracy_score: float = Field(..., ge=0.0, le=1.0, description="Fraction of claims verified")
temporal_decay_metric: float = Field(..., ge=0.0, le=1.0, description="Indicates temporal freshness penalty")
verified_claims: List[ClaimVerification] = Field(default_factory=list)
def validate_dream_report(raw_json: str) -> Optional[DreamEvaluationReport]:
try:
data = json.loads(raw_json)
# Validate using Pydantic v2
report = DreamEvaluationReport.model_validate(data)
return report
except json.JSONDecodeError:
print("Error: Invalid JSON.")
except ValidationError as e:
print(f"Validation failed: {e.errors()}")
return None
# Example usage:
if __name__ == "__main__":
sample_report = """
{
"report_id": "rep_10294",
"overall_accuracy_score": 0.95,
"temporal_decay_metric": 0.08,
"verified_claims": [
{
"claim_id": "c_1",
"text": "The Llama 4 architecture was officially open-sourced in late 2026.",
"status": "Verified",
"confidence": 0.99,
"sources": ["https://ai.meta.com/blog/llama-4-release/"],
"justification": "Verified against official Meta AI press release and tech specs."
}
]
}
"""
validated = validate_dream_report(sample_report)
if validated:
print("DREAM evaluation report successfully validated!")
print(validated.model_dump_json(indent=2))
Related tools / concepts¶
- Humanity's Last Exam (HLE) — High-difficulty static benchmark.
- SWE-bench — Evaluation for software engineering agents.
- GPQA — Graduate-level reasoning benchmark.
- LongCLI-Bench — Evaluating long-horizon terminal tasks.
- OpenHands — Open-source agentic platform often evaluated with DREAM.
- Model Context Protocol — The standard for tool discovery used by DREAM agents.
- Tavily — Search engine optimized for LLM agents.
Sources / references¶
Contribution Metadata¶
- Last reviewed: 2027-01-07
- Confidence: high