Fish Audio (Fish Speech)¶
What it is¶
Fish Audio (Fish Speech) is an advanced, open-source multilingual text-to-speech (TTS) platform powered by a revolutionary Dual-Autoregressive (Dual-AR) architecture. It is designed for high-fidelity, expressive voice synthesis and zero-shot voice cloning with extremely low latency, supporting the Model Context Protocol (FastMCP 3.1) for real-time agentic audio orchestration and voice feedback.
What problem it solves¶
It solves the dependency on proprietary, high-latency, and expensive TTS APIs like ElevenLabs. By adopting a transformer-based voice architecture isomorphic to large language models, Fish Speech enables fine-grained emotional control and achieves industry-leading Real-Time Factor (RTF) using modern inference acceleration frameworks (like SGLang, TensorRT-LLM, and NVIDIA NIM) running on NVIDIA Blackwell (B200) or Rubin GPU architectures. It serves as a high-fidelity, real-time voice feedback layer for autonomous agents powered by frontier models like Claude 5.1, GPT-5.5, Gemini 4.0 Pro, and Llama 4.
Where it fits in the stack¶
Category: AI Assistants & Knowledge / Audio Generation. It integrates with the homelab and agentic automation stack via Model Context Protocol (FastMCP 3.1) to provide real-time voice feedback and streaming conversational speech for autonomous agents.
Typical use cases¶
- Expressive Narrators: Generating audiobooks or podcast content with precise, context-dependent emotional cues.
- Conversational AI: Powering real-time agents that sound natural, responsive, and maintain ultra-low latency using FastMCP 3.1 audio streams.
- Voice Cloning: Creating high-fidelity digital twins from as little as 10-20 seconds of reference audio.
- Multilingual Content: Synthesizing speech in over 80 languages without phoneme-level preprocessing or external alignment tools.
Strengths¶
- Fine-Grained Emotion Control: Supports inline natural language tags (e.g.,
[whisper],[excited],[laughing]) to control prosody, pacing, and emotion at the sub-word level. - Innovative Dual-AR Architecture: Combines a 4B parameter "Slow AR" model for semantic prediction with a 400M parameter "Fast AR" model for acoustic detail reconstruction.
- Extreme Performance: Highly optimized for NVIDIA Blackwell and Rubin GPUs, achieving a Real-Time Factor (RTF) of ~0.12 and Time-to-First-Audio (TTFA) of ~75ms using SGLang.
- RL Alignment: Uses Group Relative Policy Optimization (GRPO) to align generated speech with human acoustic preferences.
Limitations¶
- Hardware Intensity: The flagship 4B model requires significant VRAM (ideally NVIDIA Blackwell B200, H200, or Rubin) for optimal throughput.
- Model Size: While optimized, the combined Dual-AR system is larger than lightweight on-device models like Kokoro TTS.
- Setup Complexity: Requires specialized CUDA environments, Docker, or NVIDIA NIM containers for maximum acceleration.
When to use it¶
- High-Fidelity Audio: When audio quality, natural intonation, and expressiveness are the top priorities.
- Rapid Voice Cloning: When you need to clone a voice from a very short sample (less than 30 seconds).
- GPU-Rich Environments: When you have access to high-end NVIDIA GPUs to leverage SGLang and TensorRT-LLM acceleration.
When not to use it¶
- CPU-Only Deployment: Performance is limited on consumer CPUs without dedicated GPU acceleration.
- Low-Latency Mobile Apps: The 4B model is too heavy for on-device mobile inference; use Kokoro TTS instead.
- Basic Voice Alerts: For simple notification sounds where expressiveness and complex intonations are unnecessary.
Getting started¶
Installation¶
# Clone the repository
git clone https://github.com/fishaudio/fish-speech.git
cd fish-speech
# Install dependencies using uv
uv sync --extra vllm
Hello-World¶
# Launch the Gradio-based interface
python -m tools.webui
CLI examples¶
Generate Speech with Voice Cloning¶
python -m tools.llama.generate \
--text "Hello, this is a test of Fish Audio S2 Pro using the updated FastMCP 3.1 Dual-AR pipeline." \
--prompt-text "Reference audio transcript" \
--prompt-tokens "path/to/reference.wav" \
--output "output.wav"
Batch Processing from JSON¶
python -m tools.llama.generate_batch --config batch_tasks.json --output_dir ./results/
Model Management¶
python -m tools.download_models --model-size 4b --lang all
API examples¶
Inference via FastAPI (Strict Pydantic v2 Validation)¶
import requests
from pydantic import BaseModel, Field, field_validator
class TTSRequest(BaseModel):
text: str = Field(..., description="The text to synthesize.")
reference_id: str = Field("target_voice_01", description="Reference speaker ID.")
format: str = Field("wav", description="Output audio format.")
@field_validator('text')
@classmethod
def validate_text_length(cls, v: str) -> str:
if not v.strip():
raise ValueError("Text payload cannot be empty")
return v
# Send TTS request to running local instance
req = TTSRequest(
text="The quick brown fox jumps over the lazy dog [laughing].",
reference_id="target_voice_01",
format="wav"
)
response = requests.post(
"http://localhost:8080/v1/tts",
json=req.model_dump()
)
with open("output.wav", "wb") as f:
f.write(response.content)
Using the Python SDK with FastMCP 3.1¶
from fish_speech import FishTTS
from pydantic import BaseModel, Field
class SpeechGenerationConfig(BaseModel):
model_path: str = Field("weights/fish-speech-v1.5")
reference_audio: str = Field("ref.wav")
output_path: str = Field("out.wav")
config = SpeechGenerationConfig()
tts = FishTTS(model_path=config.model_path)
tts.synthesize(
text="Synthesizing with emotional tags and FastMCP 3.1 streaming [excited].",
reference_audio=config.reference_audio,
output_path=config.output_path
)
Related tools / concepts¶
- KokoClone — Lightweight local alternative for TTS.
- Whisper — SOTA audio transcription.
- SGLang — The inference framework powering Fish Audio's speed.
- ElevenLabs — Proprietary industry standard for TTS.
- NVIDIA — Provider of Blackwell/Rubin GPU architectures and NIM microservices.
- Model Context Protocol (FastMCP 3.1) — Standard for agentic tool and resource connection.
- Audiobookshelf — Target service for Fish Audio content.
- Jellyfin — Media server for hosting synthesized audio.
Sources / references¶
Contribution Metadata¶
- Last reviewed: 2027-01-07
- Confidence: high