Text Generation Inference (TGI)¶
What it is¶
Text Generation Inference (TGI) is a specialized toolkit for deploying and serving Large Language Models (LLMs). Developed by Hugging Face, it is designed for high-performance text generation in production environments. It is written in Rust and Python, offering a robust solution for serving the most popular open-weight models.
What problem it solves¶
TGI addresses the engineering challenges of serving LLMs at scale. It implements advanced optimizations like tensor parallelism for multi-GPU inference, dynamic batching to maximize throughput, and custom Rust kernels for faster generation. As of January 2027, it serves as a high-performance alternative to NVIDIA NIM, fully optimized for NVIDIA Blackwell and Rubin GPU architectures, and provides a critical backend for developers benchmarking self-hosted models against frontier services like Claude 5.6 (supporting advanced FastMCP 3.1 tooling and pipelines), GPT-5.6, Gemini 4.0 Ultra, DeepSeek-V4, and Llama 4.
Where it fits in the stack¶
Infra. It provides the high-performance serving layer for Hugging Face models, bridging the gap between raw weights and a production-ready API.
Typical use cases¶
- Enterprise-grade LLM APIs: Powering internal or external model services with high reliability.
- Multi-GPU Deployment: Serving very large models (e.g., Llama-4-70B, Qwen-3.8-72B, Gemma-3) that require tensor parallelism.
- Real-time Chat: Production backends for applications like Hugging Chat that require streaming responses.
- Agentic Workflows: Providing a high-speed completion endpoint for autonomous agents running in Claude Code and those communicating via FastMCP 3.1 protocols.
Strengths¶
- Production-Hardened: Battle-tested at Hugging Face for their own Inference API.
- Advanced Optimizations: Includes FlashAttention-3, PagedAttention, speculative decoding, and optimized custom Rust/Triton kernels.
- Flexible Serving: Supports a wide range of Hugging Face models out of the box, including Llama 4, Qwen 3.8, and Gemma 3.
- Enterprise Features: Robust monitoring via Prometheus, streaming support, and production-ready logging.
- Multi-LoRA: Efficiently serve multiple fine-tuned adapters on a single base model.
Limitations¶
- Licensing: Uses the Hugging Face Optimized Inference License (HFOIL), which has restrictions on commercial redistribution as a service.
- Setup Complexity: Docker is the primary and recommended way to run it, which may be a barrier for environments without container support.
- Hardware Specificity: Highly optimized for NVIDIA GPUs (specifically Ampere, Ada Lovelace, Blackwell, and Rubin), though support for other accelerators is evolving.
When to use it¶
- When you need a highly optimized, production-ready server for LLMs in the Hugging Face ecosystem.
- When you need to scale models across multiple GPUs efficiently using tensor parallelism.
- When serving multiple LoRA adapters simultaneously is a requirement for multi-tenant applications.
When not to use it¶
- For local development on consumer hardware where simpler tools like Ollama or llama.cpp suffice.
- If your commercial use case conflicts with the HFOIL license terms.
- When running on Apple Silicon (use MLX instead).
Getting started¶
Installation (Docker)¶
TGI is best run via Docker to ensure all Rust dependencies and CUDA kernels are correctly configured.
# Pull the latest TGI image
docker pull ghcr.io/huggingface/text-generation-inference:latest
Hello World¶
Launch a small model to verify the setup:
model=google/gemma-3-4b-it
volume=$PWD/data
docker run --gpus all --shm-size 1g -p 8080:80 \
-v $volume:/data \
ghcr.io/huggingface/text-generation-inference:latest \
--model-id $model
CLI examples¶
1. Launch with Quantization¶
Reduce VRAM requirements using bitsandbytes or 4-bit quantization.
docker run --gpus all --shm-size 1g -p 8080:80 \
ghcr.io/huggingface/text-generation-inference:latest \
--model-id meta-llama/Llama-4-8B-Instruct \
--quantize bitsandbytes-nf4
2. Multi-GPU Tensor Parallelism¶
Serve a large model across 4 GPUs.
docker run --gpus all --shm-size 1g -p 8080:80 \
ghcr.io/huggingface/text-generation-inference:latest \
--model-id Qwen/Qwen3.8-72B-Instruct \
--num-shard 4
3. Serving with LoRA Adapters¶
Enable LoRA support and specify adapter paths.
docker run --gpus all --shm-size 1g -p 8080:80 \
ghcr.io/huggingface/text-generation-inference:latest \
--model-id meta-llama/Llama-4-8B-Instruct \
--lora-adapters "adapter_1=path/to/lora1,adapter_2=path/to/lora2"
API examples¶
Basic Generation¶
curl 127.0.0.1:8080/generate \
-X POST \
-d '{
"inputs":"The future of AI is",
"parameters":{
"max_new_tokens":20,
"stop": ["\n"]
}
}' \
-H 'Content-Type: application/json'
Streaming Response with FastMCP 3.1 Task Payload¶
curl 127.0.0.1:8080/generate_stream \
-X POST \
-d '{
"inputs": "Explain quantum computing",
"parameters": {
"max_new_tokens": 100,
"temperature": 0.2
}
}' \
-H 'Content-Type: application/json'
Python Programmatic SDK Integration¶
Python integration with Pydantic v2 strict schema validation:
import requests
from pydantic import BaseModel, Field, HttpUrl
from typing import List, Optional
class TGIParameters(BaseModel):
max_new_tokens: int = Field(default=256, ge=1, le=4096)
temperature: float = Field(default=0.1, ge=0.0, le=2.0)
top_p: float = Field(default=0.95, ge=0.0, le=1.0)
stop: List[str] = Field(default_factory=lambda: ["</s>", "[/INST]"])
class TGIRequest(BaseModel):
inputs: str = Field(..., description="Prompt formatted for the served model")
parameters: TGIParameters = Field(default_factory=TGIParameters)
class TGIResponse(BaseModel):
generated_text: str = Field(..., description="Generated text completion")
def query_tgi_endpoint(prompt: str, server_url: str = "http://localhost:8080") -> str:
req_data = TGIRequest(inputs=f"[INST] {prompt} [/INST]")
headers = {"Content-Type": "application/json"}
response = requests.post(f"{server_url}/generate", json=req_data.model_dump(), headers=headers)
if response.status_code == 200:
res_data = TGIResponse.model_validate(response.json())
return res_data.generated_text
else:
raise RuntimeError(f"TGI Request failed: {response.status_code} - {response.text}")
if __name__ == "__main__":
try:
completion = query_tgi_endpoint("List three primary benefits of FastMCP 3.1 task protocol integration.")
print(f"Completion: {completion}")
except Exception as e:
print(f"Error connecting to TGI: {e}")
Related tools / concepts¶
- NVIDIA NIM — Enterprise inference microservices.
- Aphrodite Engine — High-performance inference engine.
- vLLM — High-throughput alternative using PagedAttention.
- SGLang — Optimized for structured generation.
- llama.cpp — The standard for CPU and local inference.
- Ollama — Easy-to-use local model management.
- MLX — Apple Silicon native inference.
- Inference engines — Overview of the LLM serving ecosystem.
- Docker — Containerization platform for TGI.
- Prometheus — Monitoring system supported by TGI.
Sources / References¶
- Official Website
- GitHub Repository
- TGI Documentation: Multi-LoRA
- Hugging Face Optimized Inference License
Contribution Metadata¶
- Last reviewed: 2027-01-07
- Confidence: high