Skip to content

Ollama Benchmark CLI

What it is

Ollama Benchmark CLI is a specialized tool for measuring the local inference performance of Large Language Models (LLMs) served via Ollama. It provides rigorous, low-level metrics for tokens-per-second (TPS), latency, and processing times, enabling developers to objectively compare model performance on their specific local hardware (such as Apple Silicon, multi-GPU rigs, and custom ARM64 nodes). In early January 2027, it serves as the standard for validating local "Agentic Latency"—the precise execution time of multi-step, local reasoning and FastMCP 3.1 Task Protocol tool calls.

Architecture & Benchmark Execution Loop

sequenceDiagram
    autonumber
    participant Client as Benchmark Harness
    participant MCP as FastMCP 3.1 Task Server
    participant Ollama as Ollama REST API
    participant Engine as Local Inference Engine (GGUF)

    Client->>MCP: Request Latency & Throughput Benchmark
    MCP->>Ollama: POST /api/generate (Model, Prompt, Context Window)
    activate Ollama
    Ollama->>Engine: Load Model Weights into VRAM / RAM
    Engine-->>Ollama: Prefill Complete (Prompt Processing)
    Ollama-->>MCP: Stream First Token (TTFT Recorded)
    Engine-->>Ollama: Token Generation Loop (Decoding)
    Ollama-->>MCP: Generation Complete (Prompt Eval & Eval Durations)
    deactivate Ollama
    MCP->>MCP: Calculate Prefill TPS, Decoding TPS & P99 Latency
    MCP-->>Client: Return Pydantic v2 Validated Benchmark Report

What problem it solves

Hardware configurations for local LLMs are highly diverse and unpredictable. Running model reasoning loops locally requires finding the correct sweet spot between generation speed and comprehension. Ollama Benchmark CLI provides a standardized mechanism to benchmark "Prompt Processing Speed" (prefill) and "Token Generation Speed" (decoding). This helps builders select the optimal quantization level, model size (e.g., gemma4:9b vs deepseek-v4:32b), and context limits to support smooth, real-time agent execution without hitting memory bottlenecks.

Where it fits in the stack

Benchmarking. Used for local infrastructure performance audits, specifically for models running on top of local services. It sits alongside LLMPerf but specializes in isolated, serverless, or on-premises environments.

Typical use cases

  • Quantization Optimization: Benchmarking performance differences across different GGUF precision levels (e.g., q4_K_M vs q8_0) to maximize local TPS.
  • Hardware Integration Tests: Measuring performance gains from thermal cooling upgrades, multi-GPU configurations, or CUDA/ROCm updates.
  • Agentic Latency Auditing: Testing TTFT (Time To First Token) for complex agentic system prompts on local models (such as Gemma 4, DeepSeek-V4, and Qwen 3.6 VL).
  • Stress-Testing and Thermal Throttling: Running high-load loops for extended periods to measure hardware degradation or performance throttling under sustained compute demands.

Strengths

  • Native Integration: Directly targets the Ollama REST API endpoints, requiring no complex driver wrappers.
  • Granular Latency Parsing: Explicitly separates prefill (prompt loading) from generation phase metrics.
  • Batch Comparisons: Supports evaluating multiple local models in a single, automated execution run with structured comparison tables.

Limitations

  • Ollama Specific: Cannot evaluate raw engines like vLLM, Aphrodite, or llama.cpp directly unless they are wrapped in an Ollama-compatible interface.
  • No Qualitative Evaluation: Only measures processing speed; it does not check if the model's response is accurate or contextually sound (use HLE or LM Evaluation Harness for quality evaluations).
  • Environment Dependency: Results are tightly coupled with the host hardware state (e.g., CPU load, GPU temperature) and cannot be compared across systems without strict environment controls.

When to use it

  • When provisioning or tuning local homelab nodes in early January 2027 to run low-latency local agents.
  • When validating the execution throughput of local model pools hosting FastMCP 3.1 Task Protocol tool-calling frameworks.
  • When testing hardware efficiency during model quantization swaps.

When not to use it

Getting started

Installation is straightforward via standard Python package managers. Ensure the local Ollama service is active before running evaluations.

pip install git+https://github.com/LarHope/ollama-benchmark.git

For multi-GPU local systems, configure Ollama with GPU-specific parameters before benchmarking:

# Example CUDA configuration for parallel GPUs in 2027
export CUDA_VISIBLE_DEVICES=0,1
export OLLAMA_NUM_PARALLEL=4

CLI examples

Benchmarking specific models with comparison table

ollama-benchmark --models gemma4:9b deepseek-v4:32b --table_output

Benchmarking with context limits and thread controls

Specify custom context window limits and thread counts to mirror the active 2027 agent environments:

ollama-benchmark \
    --models gemma4:9b \
    --num_ctx 16384 \
    --num_thread 8 \
    --output-json ./metrics/gemma4_stats.json

Benchmarking with custom prompt sequences

ollama-benchmark --models qwen3.6-vl-instruct --prompts "Explain quantum computing" "Write a fast Fibonacci in Python"

API examples

Parsing and Validating Ollama Benchmarks with FastMCP 3.1 & Strict Pydantic v2

This Python script demonstrates how to integrate an Ollama benchmark suite into a FastMCP 3.1 Task Protocol tool server and parse execution metrics using strict Pydantic v2 validation (BaseModel, Field, model_validate, ValidationError).

import sys
import requests
from typing import Dict, Any
from pydantic import BaseModel, Field, ValidationError
from mcp.server.fastmcp import FastMCP

# Initialize FastMCP 3.1 Server for Ollama Benchmarking
mcp = FastMCP("Ollama-Benchmark-Server", version="3.1")

class BenchmarkOptions(BaseModel):
    num_ctx: int = Field(default=8192, description="Context window size used for test")
    temperature: float = Field(default=0.0, description="Temperature parameter")
    num_predict: int = Field(default=512, description="Max tokens to predict")

class BenchmarkResult(BaseModel):
    model: str = Field(..., description="The name of the benchmarked model")
    prompt_tokens: int = Field(..., alias="prompt_eval_count", description="Number of tokens in prompt")
    prefill_duration_ns: int = Field(..., alias="prompt_eval_duration", description="Time spent in prefill (ns)")
    generation_tokens: int = Field(..., alias="eval_count", description="Number of tokens generated")
    generation_duration_ns: int = Field(..., alias="eval_duration", description="Time spent in token generation (ns)")
    total_duration_ns: int = Field(..., alias="total_duration", description="Total API response duration in ns")

    @property
    def prefill_tps(self) -> float:
        if self.prefill_duration_ns > 0:
            return self.prompt_tokens / (self.prefill_duration_ns / 1e9)
        return 0.0

    @property
    def generation_tps(self) -> float:
        if self.generation_duration_ns > 0:
            return self.generation_tokens / (self.generation_duration_ns / 1e9)
        return 0.0

@mcp.tool(name="run_ollama_benchmark", description="Executes local inference latency and throughput benchmark against Ollama.")
def run_local_benchmark(model_name: str, prompt: str, options_dict: Dict[str, Any]) -> str:
    try:
        options = BenchmarkOptions.model_validate(options_dict)
        payload = {
            "model": model_name,
            "prompt": prompt,
            "stream": False,
            "options": options.model_dump()
        }

        response = requests.post("http://localhost:11434/api/generate", json=payload, timeout=120)
        response.raise_for_status()
        raw_data = response.json()

        metrics = BenchmarkResult.model_validate(raw_data)
        return metrics.model_dump_json(indent=2)

    except ValidationError as ve:
        return f"Metrics validation error: {ve}"
    except requests.RequestException as re:
        return f"HTTP request failed: {re}"

if __name__ == "__main__":
    # Mock offline validation check for local server testing
    mock_payload = {
        "model": "gemma4:9b",
        "prompt_eval_count": 120,
        "prompt_eval_duration": 480000000,   # 0.48s (250 tps)
        "eval_count": 300,
        "eval_duration": 4000000000,         # 4.0s (75 tps)
        "total_duration": 4500000000
    }

    try:
        validated_metrics = BenchmarkResult.model_validate(mock_payload)
        print(f"Offline validation check successful for model: {validated_metrics.model}")
        print(f"  Prefill TPS: {validated_metrics.prefill_tps:.2f}")
        print(f"  Generation TPS: {validated_metrics.generation_tps:.2f}")
    except ValidationError as e:
        print(f"Offline validation check failed: {e}", file=sys.stderr)

Sources / References

Contribution Metadata

  • Last reviewed: 2027-01-07
  • Confidence: high