Ollama Benchmark CLI¶
What it is¶
Ollama Benchmark CLI is a specialized tool for measuring the local inference performance of Large Language Models (LLMs) served via Ollama. It provides rigorous, low-level metrics for tokens-per-second (TPS), latency, and processing times, enabling developers to objectively compare model performance on their specific local hardware (such as Apple Silicon, multi-GPU rigs, and custom ARM64 nodes). In early January 2027, it serves as the standard for validating local "Agentic Latency"—the precise execution time of multi-step, local reasoning and FastMCP 3.1 Task Protocol tool calls.
Architecture & Benchmark Execution Loop¶
sequenceDiagram
autonumber
participant Client as Benchmark Harness
participant MCP as FastMCP 3.1 Task Server
participant Ollama as Ollama REST API
participant Engine as Local Inference Engine (GGUF)
Client->>MCP: Request Latency & Throughput Benchmark
MCP->>Ollama: POST /api/generate (Model, Prompt, Context Window)
activate Ollama
Ollama->>Engine: Load Model Weights into VRAM / RAM
Engine-->>Ollama: Prefill Complete (Prompt Processing)
Ollama-->>MCP: Stream First Token (TTFT Recorded)
Engine-->>Ollama: Token Generation Loop (Decoding)
Ollama-->>MCP: Generation Complete (Prompt Eval & Eval Durations)
deactivate Ollama
MCP->>MCP: Calculate Prefill TPS, Decoding TPS & P99 Latency
MCP-->>Client: Return Pydantic v2 Validated Benchmark Report
What problem it solves¶
Hardware configurations for local LLMs are highly diverse and unpredictable. Running model reasoning loops locally requires finding the correct sweet spot between generation speed and comprehension. Ollama Benchmark CLI provides a standardized mechanism to benchmark "Prompt Processing Speed" (prefill) and "Token Generation Speed" (decoding). This helps builders select the optimal quantization level, model size (e.g., gemma4:9b vs deepseek-v4:32b), and context limits to support smooth, real-time agent execution without hitting memory bottlenecks.
Where it fits in the stack¶
Benchmarking. Used for local infrastructure performance audits, specifically for models running on top of local services. It sits alongside LLMPerf but specializes in isolated, serverless, or on-premises environments.
Typical use cases¶
- Quantization Optimization: Benchmarking performance differences across different GGUF precision levels (e.g.,
q4_K_Mvsq8_0) to maximize local TPS. - Hardware Integration Tests: Measuring performance gains from thermal cooling upgrades, multi-GPU configurations, or CUDA/ROCm updates.
- Agentic Latency Auditing: Testing TTFT (Time To First Token) for complex agentic system prompts on local models (such as Gemma 4, DeepSeek-V4, and Qwen 3.6 VL).
- Stress-Testing and Thermal Throttling: Running high-load loops for extended periods to measure hardware degradation or performance throttling under sustained compute demands.
Strengths¶
- Native Integration: Directly targets the Ollama REST API endpoints, requiring no complex driver wrappers.
- Granular Latency Parsing: Explicitly separates prefill (prompt loading) from generation phase metrics.
- Batch Comparisons: Supports evaluating multiple local models in a single, automated execution run with structured comparison tables.
Limitations¶
- Ollama Specific: Cannot evaluate raw engines like vLLM, Aphrodite, or llama.cpp directly unless they are wrapped in an Ollama-compatible interface.
- No Qualitative Evaluation: Only measures processing speed; it does not check if the model's response is accurate or contextually sound (use HLE or LM Evaluation Harness for quality evaluations).
- Environment Dependency: Results are tightly coupled with the host hardware state (e.g., CPU load, GPU temperature) and cannot be compared across systems without strict environment controls.
When to use it¶
- When provisioning or tuning local homelab nodes in early January 2027 to run low-latency local agents.
- When validating the execution throughput of local model pools hosting FastMCP 3.1 Task Protocol tool-calling frameworks.
- When testing hardware efficiency during model quantization swaps.
When not to use it¶
- For cloud-based model providers (use LLMPerf).
- For qualitative reasoning or capability validation (use LM Evaluation Harness).
Getting started¶
Installation is straightforward via standard Python package managers. Ensure the local Ollama service is active before running evaluations.
pip install git+https://github.com/LarHope/ollama-benchmark.git
For multi-GPU local systems, configure Ollama with GPU-specific parameters before benchmarking:
# Example CUDA configuration for parallel GPUs in 2027
export CUDA_VISIBLE_DEVICES=0,1
export OLLAMA_NUM_PARALLEL=4
CLI examples¶
Benchmarking specific models with comparison table¶
ollama-benchmark --models gemma4:9b deepseek-v4:32b --table_output
Benchmarking with context limits and thread controls¶
Specify custom context window limits and thread counts to mirror the active 2027 agent environments:
ollama-benchmark \
--models gemma4:9b \
--num_ctx 16384 \
--num_thread 8 \
--output-json ./metrics/gemma4_stats.json
Benchmarking with custom prompt sequences¶
ollama-benchmark --models qwen3.6-vl-instruct --prompts "Explain quantum computing" "Write a fast Fibonacci in Python"
API examples¶
Parsing and Validating Ollama Benchmarks with FastMCP 3.1 & Strict Pydantic v2¶
This Python script demonstrates how to integrate an Ollama benchmark suite into a FastMCP 3.1 Task Protocol tool server and parse execution metrics using strict Pydantic v2 validation (BaseModel, Field, model_validate, ValidationError).
import sys
import requests
from typing import Dict, Any
from pydantic import BaseModel, Field, ValidationError
from mcp.server.fastmcp import FastMCP
# Initialize FastMCP 3.1 Server for Ollama Benchmarking
mcp = FastMCP("Ollama-Benchmark-Server", version="3.1")
class BenchmarkOptions(BaseModel):
num_ctx: int = Field(default=8192, description="Context window size used for test")
temperature: float = Field(default=0.0, description="Temperature parameter")
num_predict: int = Field(default=512, description="Max tokens to predict")
class BenchmarkResult(BaseModel):
model: str = Field(..., description="The name of the benchmarked model")
prompt_tokens: int = Field(..., alias="prompt_eval_count", description="Number of tokens in prompt")
prefill_duration_ns: int = Field(..., alias="prompt_eval_duration", description="Time spent in prefill (ns)")
generation_tokens: int = Field(..., alias="eval_count", description="Number of tokens generated")
generation_duration_ns: int = Field(..., alias="eval_duration", description="Time spent in token generation (ns)")
total_duration_ns: int = Field(..., alias="total_duration", description="Total API response duration in ns")
@property
def prefill_tps(self) -> float:
if self.prefill_duration_ns > 0:
return self.prompt_tokens / (self.prefill_duration_ns / 1e9)
return 0.0
@property
def generation_tps(self) -> float:
if self.generation_duration_ns > 0:
return self.generation_tokens / (self.generation_duration_ns / 1e9)
return 0.0
@mcp.tool(name="run_ollama_benchmark", description="Executes local inference latency and throughput benchmark against Ollama.")
def run_local_benchmark(model_name: str, prompt: str, options_dict: Dict[str, Any]) -> str:
try:
options = BenchmarkOptions.model_validate(options_dict)
payload = {
"model": model_name,
"prompt": prompt,
"stream": False,
"options": options.model_dump()
}
response = requests.post("http://localhost:11434/api/generate", json=payload, timeout=120)
response.raise_for_status()
raw_data = response.json()
metrics = BenchmarkResult.model_validate(raw_data)
return metrics.model_dump_json(indent=2)
except ValidationError as ve:
return f"Metrics validation error: {ve}"
except requests.RequestException as re:
return f"HTTP request failed: {re}"
if __name__ == "__main__":
# Mock offline validation check for local server testing
mock_payload = {
"model": "gemma4:9b",
"prompt_eval_count": 120,
"prompt_eval_duration": 480000000, # 0.48s (250 tps)
"eval_count": 300,
"eval_duration": 4000000000, # 4.0s (75 tps)
"total_duration": 4500000000
}
try:
validated_metrics = BenchmarkResult.model_validate(mock_payload)
print(f"Offline validation check successful for model: {validated_metrics.model}")
print(f" Prefill TPS: {validated_metrics.prefill_tps:.2f}")
print(f" Generation TPS: {validated_metrics.generation_tps:.2f}")
except ValidationError as e:
print(f"Offline validation check failed: {e}", file=sys.stderr)
Related tools / concepts¶
- Ollama Service - The underlying model server.
- LLMPerf - Benchmarking API-based LLM performance.
- LM Evaluation Harness - Benchmarking model quality/accuracy.
- HLE (Humanity's Last Exam) - Frontier reasoning benchmark.
- MBPP - Code generation benchmark for Python.
- vLLM - High-performance inference server.
- Aphrodite Engine - High-throughput local inference engine.
- Terminus 2 - Benchmarking terminal-based agent interactions.
Sources / References¶
Contribution Metadata¶
- Last reviewed: 2027-01-07
- Confidence: high