MT-Bench¶
What it is¶
MT-Bench is a benchmark designed to evaluate the multi-turn conversational capabilities of Large Language Models (LLMs). It consists of 80 high-quality, multi-turn questions across eight categories: writing, roleplay, extraction, reasoning, math, coding, knowledge I (STEM), and knowledge II (humanities/social science). In the July 2026 landscape, it is increasingly integrated with the MCP 3.0 Task Protocol to automate complex, stateful evaluation loops across frontier models like Gemma 3, Claude 4.8 Opus, and GPT-5.5.
What problem it solves¶
Many traditional benchmarks only evaluate single-turn responses, failing to capture a model's ability to maintain context, follow instructions across multiple exchanges, and handle the dynamic nature of real-world conversations. MT-Bench specifically tests the "follow-up" capability of models, addressing the "goldfish memory" problem and verifying instruction adherence in deep, multi-turn dialogues.
Where it fits in the stack¶
Benchmarking. It is a core component of the LMSYS FastChat evaluation framework and is often used alongside the MCP 3.0 Task Protocol to benchmark autonomous agent persistence and context management.
Typical use cases¶
- Conversational AI Evaluation: Assessing how well a chatbot handles follow-up questions and maintains context.
- Model Comparison: Ranking chat-tuned models (e.g., Gemma 3 vs. Claude 4.8) based on their ability to handle multi-step instructions.
- LLM-as-a-Judge Validation: MT-Bench uses strong models as judges to provide automated, scalable scoring, now enhanced by the MCP 3.0 standardized task representations.
- Agentic Workflow Stress-Testing: Verifying that agents can maintain state across long-running tasks.
Strengths¶
- Multi-turn Focus: Specifically designed to test conversation depth and instruction adherence over multiple turns.
- Diverse Categories: Covers a wide range of tasks from coding to roleplay, ensuring a balanced evaluation.
- Strong Human Correlation: Scoring on MT-Bench shows high agreement (over 80%) with human expert preferences.
- MCP 3.0 Integration: Modern implementations leverage the Task Protocol for standardized, reproducible evaluation runs.
Limitations¶
- Judge Bias: If using an LLM as a judge, it may inherit the biases of that judge (e.g., preference for verbosity or certain styles).
- Small Sample Size: With only 80 questions, the results can have higher variance than larger benchmarks like AlpacaEval.
- Static Nature: Like all fixed benchmarks, it risks data contamination if questions are leaked into training sets.
When to use it¶
- When evaluating chat-tuned models where multi-turn interaction is a primary use case.
- When you need an automated conversational benchmark that aligns closely with human preference.
- For internal testing of "System 2" reasoning models in conversational contexts.
When not to use it¶
- For evaluating base (non-chat-tuned) models that are not designed for dialogue.
- When you only need to measure narrow technical capabilities like raw code execution (use BigCodeBench instead).
- When high-stakes safety evaluation is the primary goal.
Getting started¶
1. Installation¶
MT-Bench is part of the fastchat repository.
git clone https://github.com/lm-sys/FastChat.git
cd FastChat
pip install -e ".[model_worker,llm_judge]"
2. Generating Model Answers¶
Generate answers for the 80 questions using your local model or API.
# Example using a local Gemma 3 model via MCP 3.0
python fastchat/llm_judge/gen_model_answer.py \
--model-path google/gemma-3-27b-it \
--model-id gemma-3-27b-it
3. Grading with LLM-as-a-Judge¶
Use a strong model (like GPT-5.5) to grade the responses.
export OPENAI_API_KEY="your_api_key"
python fastchat/llm_judge/gen_judgment.py \
--model-list gemma-3-27b-it \
--parallel 4
CLI examples¶
The FastChat evaluation suite provides several CLI tools for MT-Bench.
# Show results summary
python fastchat/llm_judge/show_result.py
# Generate answers for a specific category
python fastchat/llm_judge/gen_model_answer.py --category reasoning --model-id my-custom-model
# Run judgments in parallel to save time
python fastchat/llm_judge/gen_judgment.py --model-list model1 model2 --parallel 8
# Export judgments to a JSON file (standardized for MCP 3.0)
python fastchat/llm_judge/gen_judgment.py --model-list model1 --output-file results.json
API examples¶
While MT-Bench is primarily a CLI-driven benchmark, it can be integrated into Python evaluation pipelines.
import json
from fastchat.llm_judge.common import load_questions, load_model_answers
# Load MT-Bench questions
questions = load_questions("fastchat/llm_judge/data/mt_bench/question.jsonl")
# Access a specific multi-turn question (Standardized via MCP 3.0 Task Protocol)
first_question = questions[0]
print(f"Turn 1: {first_question['turns'][0]}")
print(f"Turn 2: {first_question['turns'][1]}")
# Custom logic to process model answers
answers = load_model_answers("fastchat/llm_judge/data/mt_bench/model_answer/gemma-3-27b-it.jsonl")
Related tools / concepts¶
- Chatbot Arena - The primary leaderboard for human preferences.
- AlpacaEval - Simulator-based evaluator for instruction following.
- GSM8K - Basic math reasoning benchmark.
- MATH Benchmark - Advanced mathematical competition problems.
- HumanEval - Core coding benchmark.
- LM Evaluation Harness - Standard framework for single-turn benchmarks.
- OpenCompass - Comprehensive evaluation platform.
- BigCodeBench - Realistic code generation benchmark.
- MCP 3.0 - Protocol used for automated benchmarking tasks.
Sources / references¶
- FastChat GitHub (LLM Judge)
- MT-Bench Paper: "Judging LLM-as-a-judge" (Zheng et al., 2023)
- LMSYS Leaderboard
- Gemma 3 Technical Report
Contribution Metadata¶
- Last reviewed: 2026-07-21
- Confidence: high