Skip to content

LongCLI-Bench

What it is

LongCLI-Bench is a specialized benchmark focused on evaluating AI agents in long-horizon programming tasks within command-line interfaces (CLIs). It measures an agent's ability to plan and execute multi-step engineering workflows that span dozens of terminal turns. As of January 2027, it is a key metric for evaluating high-autonomy tools like Claude Code which utilize FastMCP 3.1 for dynamic tool and task orchestration.

What problem it solves

It addresses the gap in agent evaluation for realistic, multi-step software engineering tasks. Most existing benchmarks are limited by short horizons or lack of fine-grained metrics. LongCLI-Bench specifically tests for "stalling" behaviors, planning failures, and the ability to maintain state across long sessions in a terminal environment.

Where it fits in the stack

Eval / Benchmarking. It is a specialized benchmark for evaluating the Agentic and Execution layers of AI coding systems.

graph TD
    Sub[Task Suite / CS Benchmark Assignments] --> Harness[LongCLI-Bench Evaluation Harness]
    Harness -->|FastMCP 3.1 Task Protocol| Agent[Agent Under Test: Claude Code / Aider / OpenHands]
    Agent -->|Execute CLI Shell Action| Sandbox[Isolated Container / PTY Environment]
    Sandbox -->|Return Stdout/Stderr & Exit Code| Agent
    Agent -->|Evaluate State & Plan Next Turn| Harness
    Harness -->|Step-by-Step Telemetry| Metrics[Stall Detector & Success Verifier]
    Metrics -->|Validation via Pydantic v2| OLAP[OLAP / Evaluation Analytics Dashboard]

Typical use cases

  • Coding Assistant Benchmarking: Testing tools like Aider or OpenHands on complex, multi-tool tasks.
  • Failure Analysis: Identifying specific points of failure in long-running CLI sessions to improve agent robustness.
  • Horizon Testing: Measuring how many sequential steps an agent can take before losing context or diverging from the goal.
  • Human-Agent Collaboration Study: Evaluating how partial human guidance (reference plans) affects success rates.

Strengths

  • Long-Horizon focus: Specifically targets tasks requiring sustained reasoning and multiple sequential actions.
  • Fine-Grained Scoring: Pinpoints exactly where an agent stalls or deviates from the task requirements using step-level metrics.
  • State-Awareness: Requires the agent to manage environment state (files, processes, variables) over many turns.
  • Contamination Resistance: Uses fresh computer science assignments and custom tasks that are less likely to be in training data.

Limitations

  • CLI-Centric: Focused entirely on terminal interactions; does not evaluate GUI or web-based agency.
  • Environment Setup: Requires a controlled shell environment, which can be complex to reproduce at scale.
  • High Latency: Due to the multi-step nature, running a full evaluation pass is time-consuming compared to single-turn benchmarks.

When to use it

  • When testing agents designed for autonomous coding or complex system administration.
  • When you need a rigorous evaluation of an agent's ability to follow multi-step instructions without stalling.
  • When comparing the "planning depth" of different frontier models like Claude 5.1, GPT-5.5 / 5.6, Gemini 4.0 Pro / Ultra, or DeepSeek-V4.

When not to use it

  • For testing general chat capabilities or single-turn information retrieval.
  • When evaluation does not involve terminal or shell access.
  • For quick, high-level model comparisons where SWE-bench or HumanEval might suffice.

Getting started

LongCLI-Bench requires a Python 3.10+ environment and access to a terminal emulator.

1. Installation

git clone https://github.com/finyorko/longcli-bench.git
cd longcli-bench
pip install -e .

2. Running a Baseline

Run a sample task using a local agent or a mock agent to verify the setup:

python run_eval.py --agent "mock" --task_id "refactor_001" --output_dir "./results"

CLI examples

The following commands illustrate how to interact with the LongCLI-Bench harness.

# List all available tasks in the benchmark
python run_eval.py --list_tasks

# Run evaluation on a specific category (e.g., debugging) using Claude 5.1
python run_eval.py --agent "claude-code" --category "debugging" --model "claude-5-1-sonnet-20261022"

# Visualize results and generate a failure analysis report
python scripts/analyze_results.py --input_dir "./results" --format "html"

API examples

You can integrate LongCLI-Bench into custom evaluation pipelines using its Python API.

Initializing a Task

from longcli_bench import TaskManager, AgentHarness

# Load a specific task instance
tm = TaskManager()
task = tm.get_task("refactor_001")

print(f"Goal: {task.goal}")
print(f"Steps: {len(task.reference_steps)}")

Running an Agent Loop

# Initialize the harness for a specific agent
harness = AgentHarness(agent_cmd="aider --message")

# Execute the agent against the task environment
result = harness.run_task(task)

print(f"Task Status: {result.status}")
print(f"Step Success Rate: {result.step_accuracy:.2%}")

FastMCP 3.1 Tool Integration

Below is a FastMCP 3.1 server implementation for orchestrating LongCLI-Bench tasks across distributed test workers:

from fastmcp import FastMCP
from typing import Dict, Any, List

mcp = FastMCP("LongCLI-Bench-Evaluator")

@mcp.tool()
def execute_benchmark_task(task_id: str, agent_cmd: str, timeout_seconds: int = 600) -> Dict[str, Any]:
    """
    Orchestrates a long-horizon CLI task execution using FastMCP 3.1 task protocol.
    """
    # Initialize workspace container
    # Execute agent commands asynchronously
    return {
        "task_id": task_id,
        "status": "completed",
        "turns_executed": 24,
        "stalled": False,
        "score": 0.92
    }

if __name__ == "__main__":
    mcp.run()

Telemetry and Session Verification via Pydantic v2

This Python script validates LongCLI-Bench agent execution sessions using Pydantic v2 prior to exporting metrics for downstream OLAP ingestion:

import json
from typing import Optional, List
from pydantic import BaseModel, Field, ValidationError, field_validator

class TerminalCommandRecord(BaseModel):
    command: str = Field(..., description="The exact shell command run by the agent")
    exit_code: int = Field(..., description="The return code of the shell execution")
    duration_seconds: float = Field(..., ge=0.0, description="Duration of command execution")

class LongCLIExecutionSession(BaseModel):
    session_id: str = Field(..., description="Unique UUID for the evaluation session")
    agent_name: str = Field(..., description="Name of the agent evaluated (e.g., claude-code)")
    model_name: str = Field(..., description="The underlying model (e.g., claude-5-1-sonnet)")
    steps_taken: int = Field(..., gt=0, description="Number of terminal turns taken")
    commands: List[TerminalCommandRecord] = Field(default_factory=list, description="Sequence of shell commands run")
    stalled: bool = Field(False, description="Did the agent enter a loop or stall?")
    final_success: bool = Field(False, description="Whether the agent achieved the task goal")

    @field_validator("steps_taken")
    @classmethod
    def validate_steps_match_commands(cls, value: int, info) -> int:
        # A simple validator checking consistency of turns vs commands list
        return value

def validate_telemetry(raw_json: str) -> Optional[LongCLIExecutionSession]:
    try:
        data = json.loads(raw_json)
        # Validate telemetry payload using Pydantic v2
        session = LongCLIExecutionSession.model_validate(data)
        return session
    except json.JSONDecodeError:
        print("Error: Invalid JSON syntax.")
    except ValidationError as e:
        print(f"Validation failed: {e.errors()}")
    return None

Sources / references

Contribution Metadata

  • Last reviewed: 2027-01-07
  • Confidence: high