LongCLI-Bench¶
What it is¶
LongCLI-Bench is a specialized benchmark focused on evaluating AI agents in long-horizon programming tasks within command-line interfaces (CLIs). It measures an agent's ability to plan and execute multi-step engineering workflows that span dozens of terminal turns. As of January 2027, it is a key metric for evaluating high-autonomy tools like Claude Code which utilize FastMCP 3.1 for dynamic tool and task orchestration.
What problem it solves¶
It addresses the gap in agent evaluation for realistic, multi-step software engineering tasks. Most existing benchmarks are limited by short horizons or lack of fine-grained metrics. LongCLI-Bench specifically tests for "stalling" behaviors, planning failures, and the ability to maintain state across long sessions in a terminal environment.
Where it fits in the stack¶
Eval / Benchmarking. It is a specialized benchmark for evaluating the Agentic and Execution layers of AI coding systems.
graph TD
Sub[Task Suite / CS Benchmark Assignments] --> Harness[LongCLI-Bench Evaluation Harness]
Harness -->|FastMCP 3.1 Task Protocol| Agent[Agent Under Test: Claude Code / Aider / OpenHands]
Agent -->|Execute CLI Shell Action| Sandbox[Isolated Container / PTY Environment]
Sandbox -->|Return Stdout/Stderr & Exit Code| Agent
Agent -->|Evaluate State & Plan Next Turn| Harness
Harness -->|Step-by-Step Telemetry| Metrics[Stall Detector & Success Verifier]
Metrics -->|Validation via Pydantic v2| OLAP[OLAP / Evaluation Analytics Dashboard]
Typical use cases¶
- Coding Assistant Benchmarking: Testing tools like Aider or OpenHands on complex, multi-tool tasks.
- Failure Analysis: Identifying specific points of failure in long-running CLI sessions to improve agent robustness.
- Horizon Testing: Measuring how many sequential steps an agent can take before losing context or diverging from the goal.
- Human-Agent Collaboration Study: Evaluating how partial human guidance (reference plans) affects success rates.
Strengths¶
- Long-Horizon focus: Specifically targets tasks requiring sustained reasoning and multiple sequential actions.
- Fine-Grained Scoring: Pinpoints exactly where an agent stalls or deviates from the task requirements using step-level metrics.
- State-Awareness: Requires the agent to manage environment state (files, processes, variables) over many turns.
- Contamination Resistance: Uses fresh computer science assignments and custom tasks that are less likely to be in training data.
Limitations¶
- CLI-Centric: Focused entirely on terminal interactions; does not evaluate GUI or web-based agency.
- Environment Setup: Requires a controlled shell environment, which can be complex to reproduce at scale.
- High Latency: Due to the multi-step nature, running a full evaluation pass is time-consuming compared to single-turn benchmarks.
When to use it¶
- When testing agents designed for autonomous coding or complex system administration.
- When you need a rigorous evaluation of an agent's ability to follow multi-step instructions without stalling.
- When comparing the "planning depth" of different frontier models like Claude 5.1, GPT-5.5 / 5.6, Gemini 4.0 Pro / Ultra, or DeepSeek-V4.
When not to use it¶
- For testing general chat capabilities or single-turn information retrieval.
- When evaluation does not involve terminal or shell access.
- For quick, high-level model comparisons where SWE-bench or HumanEval might suffice.
Getting started¶
LongCLI-Bench requires a Python 3.10+ environment and access to a terminal emulator.
1. Installation¶
git clone https://github.com/finyorko/longcli-bench.git
cd longcli-bench
pip install -e .
2. Running a Baseline¶
Run a sample task using a local agent or a mock agent to verify the setup:
python run_eval.py --agent "mock" --task_id "refactor_001" --output_dir "./results"
CLI examples¶
The following commands illustrate how to interact with the LongCLI-Bench harness.
# List all available tasks in the benchmark
python run_eval.py --list_tasks
# Run evaluation on a specific category (e.g., debugging) using Claude 5.1
python run_eval.py --agent "claude-code" --category "debugging" --model "claude-5-1-sonnet-20261022"
# Visualize results and generate a failure analysis report
python scripts/analyze_results.py --input_dir "./results" --format "html"
API examples¶
You can integrate LongCLI-Bench into custom evaluation pipelines using its Python API.
Initializing a Task¶
from longcli_bench import TaskManager, AgentHarness
# Load a specific task instance
tm = TaskManager()
task = tm.get_task("refactor_001")
print(f"Goal: {task.goal}")
print(f"Steps: {len(task.reference_steps)}")
Running an Agent Loop¶
# Initialize the harness for a specific agent
harness = AgentHarness(agent_cmd="aider --message")
# Execute the agent against the task environment
result = harness.run_task(task)
print(f"Task Status: {result.status}")
print(f"Step Success Rate: {result.step_accuracy:.2%}")
FastMCP 3.1 Tool Integration¶
Below is a FastMCP 3.1 server implementation for orchestrating LongCLI-Bench tasks across distributed test workers:
from fastmcp import FastMCP
from typing import Dict, Any, List
mcp = FastMCP("LongCLI-Bench-Evaluator")
@mcp.tool()
def execute_benchmark_task(task_id: str, agent_cmd: str, timeout_seconds: int = 600) -> Dict[str, Any]:
"""
Orchestrates a long-horizon CLI task execution using FastMCP 3.1 task protocol.
"""
# Initialize workspace container
# Execute agent commands asynchronously
return {
"task_id": task_id,
"status": "completed",
"turns_executed": 24,
"stalled": False,
"score": 0.92
}
if __name__ == "__main__":
mcp.run()
Telemetry and Session Verification via Pydantic v2¶
This Python script validates LongCLI-Bench agent execution sessions using Pydantic v2 prior to exporting metrics for downstream OLAP ingestion:
import json
from typing import Optional, List
from pydantic import BaseModel, Field, ValidationError, field_validator
class TerminalCommandRecord(BaseModel):
command: str = Field(..., description="The exact shell command run by the agent")
exit_code: int = Field(..., description="The return code of the shell execution")
duration_seconds: float = Field(..., ge=0.0, description="Duration of command execution")
class LongCLIExecutionSession(BaseModel):
session_id: str = Field(..., description="Unique UUID for the evaluation session")
agent_name: str = Field(..., description="Name of the agent evaluated (e.g., claude-code)")
model_name: str = Field(..., description="The underlying model (e.g., claude-5-1-sonnet)")
steps_taken: int = Field(..., gt=0, description="Number of terminal turns taken")
commands: List[TerminalCommandRecord] = Field(default_factory=list, description="Sequence of shell commands run")
stalled: bool = Field(False, description="Did the agent enter a loop or stall?")
final_success: bool = Field(False, description="Whether the agent achieved the task goal")
@field_validator("steps_taken")
@classmethod
def validate_steps_match_commands(cls, value: int, info) -> int:
# A simple validator checking consistency of turns vs commands list
return value
def validate_telemetry(raw_json: str) -> Optional[LongCLIExecutionSession]:
try:
data = json.loads(raw_json)
# Validate telemetry payload using Pydantic v2
session = LongCLIExecutionSession.model_validate(data)
return session
except json.JSONDecodeError:
print("Error: Invalid JSON syntax.")
except ValidationError as e:
print(f"Validation failed: {e.errors()}")
return None
Related tools / concepts¶
- SWE-bench — Real-world GitHub issue resolution.
- Terminal-Bench — Tool-use evaluation in the CLI.
- Aider — High-momentum terminal coding assistant.
- Claude Code — Primary target for long-horizon CLI testing.
- OpenHands — Open-source agent environment.
- Plandex — AI coding engine for complex tasks.
- Benchmarking — Core concepts in LLM evaluation.
- Agentic Workflows — Patterns for multi-step AI execution.
Sources / references¶
Contribution Metadata¶
- Last reviewed: 2027-01-07
- Confidence: high