PA-bench¶
What it is¶
PA-bench is a comprehensive benchmark suite designed to evaluate the performance of Personal Assistant (PA) web agents on real-world workflows. It utilizes simulated environments (e.g., mock Gmail, mock Google Calendar) to provide a safe, reproducible, and cost-effective testbed for early January 2027 agentic orchestration.
Architecture & Agentic Simulation Flow¶
flowchart TD
subgraph Suite ["PA-bench Task Harness"]
TaskDef["Task Definition (Calendar / Email / Travel)"]
Orchestrator["Experiment Orchestrator"]
end
subgraph Simulation ["Simulated Backend Enclaves"]
DockerSim["Docker Container: PA Simulations"]
MockGmail["Mock Gmail API / Web UI"]
MockGCal["Mock Google Calendar API / Web UI"]
end
subgraph AgentLoop ["Agent Execution Loop"]
Agent["Web Agent (Claude 5.6 / GPT-5.6 / Gemini 4.0 Ultra)"]
Browser["Chromium Headless (v146 Side-Panel Hooks)"]
FastMCP["FastMCP 3.1 Task Server"]
end
subgraph Scoring ["Evaluation & Verification"]
Verify["Deterministic State Verifier"]
Scorecard["Pydantic v2 Validated Trajectory Report"]
end
TaskDef --> Orchestrator
Orchestrator --> DockerSim
DockerSim --> MockGmail & MockGCal
Orchestrator --> Agent
Agent --> Browser
Browser --> FastMCP
FastMCP --> MockGmail & MockGCal
MockGmail & MockGCal --> Verify
Verify --> Scorecard
What problem it solves¶
It addresses the lack of realistic evaluation frameworks for web-based agents by providing a set of complex, multi-step tasks that mirror actual user needs, such as booking travel, managing calendars, or conducting research across multiple websites. It is a critical tool for measuring "Agentic Session Orchestration" and risk mitigation in early 2027.
Where it fits in the stack¶
Eval. It provides the metrics and environment necessary to measure the effectiveness and reliability of autonomous web agents. It is the gold standard for evaluating "Agentic Hooks" and side-panel integration in Chrome v145+/146+.
Typical use cases¶
- Agent Comparison: Evaluating different agent architectures (e.g., Claude 5.6 vs. GPT-5.6 vs. Gemini 4.0 Ultra) on their ability to complete complex web tasks.
- Regression Testing: Ensuring that updates to an agent's reasoning or navigation logic don't break existing capabilities in the Ralph-loop.
- Research: Providing a standardized baseline for academic and industrial research into autonomous web navigation and "Computer Use" capabilities.
- Security Auditing: Testing agentic resilience against adversarial UI patterns using the SHARP (SharpAI Security Benchmark) methodology.
Strengths¶
- Real-world Focus: Tasks are based on actual personal assistant workflows rather than synthetic laboratory examples.
- End-to-End Evaluation: Measures the agent's ability to see a task through from start to finish, including handling unexpected UI states.
- Complexity: Includes tasks that require multi-site navigation, state management, and interaction with JMAP/Graph APIs via FastMCP 3.1 Task Protocol.
- Deterministic: Simulated backends ensure that benchmark runs are reproducible and not subject to real-world data drift.
Limitations¶
- Environment Stability: While simulations are more stable than the live web, maintaining them requires ongoing effort as real-world APIs evolve.
- Resource Intensive: Running full-scale web agent evaluations can be time and credit consuming, requiring significant inference budget in early 2027.
- Not for Code-Gen: Less effective for evaluating pure code generation or algorithmic reasoning (use MBPP or SWE-bench for those).
When to use it¶
- When developing or refining autonomous agents intended for web-based personal assistant tasks in 2027.
- When you need a high-signal metric for how well an agent handles real-world web complexity and "Computer Use".
- When validating "Agentic Calendar Orchestration" workflows.
When not to use it¶
- For evaluating models on pure reasoning or coding tasks without a web navigation component.
- If you lack the infrastructure to run autonomous browser-based agents or the LiteLLM inference plane.
- For evaluating low-level text completion or translation tasks.
Getting started¶
PA-bench is typically executed via its Python SDK, which manages simulated environments for email and calendar applications. It requires a local Docker environment for the simulation manager.
1. Installation¶
pip install pa-bench
# Ensure Docker is running
docker pull vibrantlabs/pa-simulations:latest
2. Basic Configuration¶
Configure your agent's API access via environment variables or a .env file:
export ANTHROPIC_API_KEY="sk-..."
export OPENAI_API_KEY="sk-proj-..."
export PA_BENCH_MODE="simulation"
CLI examples¶
Running a standard evaluation suite¶
Run the "Travel" category of tasks against an agent:
pa-bench run --suite travel --agent my_custom_agent --max_steps 50
Running with visual diagnostics¶
Execute the benchmark while preserving visual trajectories, specifying a target output directory:
pa-bench run \
--suite calendar_sync \
--agent agent_claude_5_6 \
--max_steps 75 \
--visualize \
--screenshot_dir ./diagnostics/screenshots/
Listing available tasks¶
pa-bench list-tasks --category research
Visualizing Agent Trajectories¶
Generate a video report of the agent's interaction:
pa-bench report --run_id RUN_123 --format webm
API examples¶
Trajectory Schema Validation & FastMCP 3.1 Server Integration¶
Using Pydantic v2 and FastMCP 3.1 Task Protocol, we validate web agent trajectories generated during PA-bench runs before persisting them to the database or passing them to evaluation engines (using frontier models like Claude 5.6, GPT-5.6, Gemini 4.0 Ultra, Gemma 4, DeepSeek-V4, and Qwen 3.6 VL).
import json
from typing import List, Optional
from datetime import datetime
from pydantic import BaseModel, Field, ValidationError
from mcp.server.fastmcp import FastMCP
# Initialize FastMCP 3.1 Server for PA-bench Trajectory Telemetry
mcp = FastMCP("PA-Bench-Evaluator", version="3.1")
class TrajectoryStep(BaseModel):
step_num: int = Field(..., description="Chronological step number", ge=1)
action: str = Field(..., description="Web action performed (click, type, navigate, wait)")
url: Optional[str] = Field(None, description="URL where the action occurred")
screenshot_path: Optional[str] = Field(None, description="Local path to screenshot artifact")
class PAEvaluationRun(BaseModel):
run_id: str = Field(..., description="Unique run identifier")
task_name: str = Field(..., description="Name of the task from PA-bench suite")
started_at: datetime = Field(default_factory=datetime.utcnow, description="Evaluation run start time")
steps: List[TrajectoryStep] = Field(default_factory=list, description="Sequence of actions taken by agent")
is_success: bool = Field(False, description="Whether final verification check succeeded")
@mcp.tool(name="validate_pa_trajectory", description="Validates PA-bench trajectory execution data using strict Pydantic v2 schema.")
def validate_pa_bench_run(run_data: dict) -> str:
try:
validated = PAEvaluationRun.model_validate(run_data)
return validated.model_dump_json(indent=2)
except ValidationError as e:
return json.dumps({"error": "Trajectory verification failed", "details": e.errors()})
if __name__ == "__main__":
sample_run = {
"run_id": "run-pa-9912",
"task_name": "calendar_sync_2027",
"is_success": True,
"steps": [
{
"step_num": 1,
"action": "navigate",
"url": "http://gcal.mock-env.local",
"screenshot_path": "./diagnostics/screenshots/step_01.png"
}
]
}
print(validate_pa_bench_run(sample_run))
Orchestrating an Evaluation¶
Integrate PA-bench into a CI/CD pipeline for agentic software factories.
from pa_bench import SimulationManager, ExperimentOrchestrator
from my_agent import CustomWebAgent
# Initialize simulations (Gmail, GCal, etc.)
sim_manager = SimulationManager()
sim_manager.spawn_instances(apps=["gmail", "google_calendar"])
# Configure orchestrator with early January 2027 settings
orchestrator = ExperimentOrchestrator(
agent=CustomWebAgent(model="claude-5-6-sonnet"),
max_steps=75,
resolution=(1280, 960),
mcp_enabled=True,
chrome_args=[
"--enable-extension-hooks",
"--side-panel-integration"
]
)
# Run benchmark suite
results = orchestrator.run_suite(tasks="calendar_sync_2027")
print(f"Success Rate: {results.success_rate}")
print(f"Mean Steps: {results.mean_steps_to_completion}")
# Cleanup
sim_manager.shutdown()
Defining a Custom Web Task¶
from pa_bench import TaskDefinition
task = TaskDefinition(
id="multi_calendar_sync_01",
goal="Sync my Fastmail and GCal events for next Tuesday.",
apps=["gmail", "google_calendar"],
eval_script="verify_sync.py"
)
Related tools / concepts¶
- Web Agents - Core architectural patterns.
- Agentic Workflows - Orchestration strategies.
- OpenHands - Open-source agentic development platform.
- SWE-bench - Software engineering benchmark.
- Terminal-bench - Shell interaction benchmark.
- GAIA (General AI Assistants) - Real-world assistant tasks.
- AssistantBench - Web-search and navigation benchmark.
- OSWorld - Operating system-wide agent evaluation.
- Skills in Chrome - Browser-native agentic hooks in early 2027.
- FastMCP 3.1 Task Protocol - The protocol for agentic tool use and tasks.
Sources / references¶
- PA-bench: Evaluating web agents on real world personal assistant workflows
- Vibrant Labs GitHub Repository
- Agentic Session Orchestration 2027 Whitepaper
Contribution Metadata¶
- Last reviewed: 2027-01-07
- Confidence: high