Skip to content

Distilabel

What it is

Distilabel is an open-source framework designed for scalable, high-fidelity synthetic data generation and structured AI feedback (RLHF / RLAIF). As of early 2027, Distilabel has progressed to v2.5.0+, establishing itself as an industry standard for preparing training and fine-tuning datasets for frontier models like Claude 5.6, GPT-5.6, Gemini 4.0 Ultra, and DeepSeek-V4. It enables developers to construct complex multi-step pipelines that orchestrate LLMs to generate, mutate, score, and filter datasets using advanced "LLM-as-a-judge" patterns.

What problem it solves

Creating high-quality instruction and preference datasets for model training is a major bottleneck in AI development. Manual labeling is expensive and slow, while raw synthetic generation without filtering is prone to redundancy and noise. Distilabel solves these problems by: - Declarative Pipelines: Providing a clear, standard way to define data generation and feedback steps as Python pipelines. - Robust Scale: Native support for parallel execution across model APIs (Anthropic, OpenAI) and high-throughput local backends like vLLM or Ollama. - High Data Quality: Built-in scoring, ranking, and deduplication modules that filter out low-quality data. - Dynamic Tool Calling: Native Model Context Protocol (FastMCP 3.1) support to supply synthetic agents with live tools during generation.

Where it fits in the stack

Distilabel sits in the Frameworks / Data-Generation layer. It serves as the primary data engineering and preparation pipeline that feeds model-training frameworks like Unsloth, Axolotl, and LLaMA Factory.

Typical use cases

  • Evol-Instruct Pipelines: Taking simple prompt seeds and evolving them into highly complex multi-turn instructions using frontier models.
  • Preference Dataset Creation (RLHF/DPO): Generating multiple responses to a prompt and using Claude 5.6 as a judge to score and output structured pairwise preferences.
  • VLM/RAG Data Enrichment: Synthesizing high-quality question-answering pairs from document repositories or image databases.
  • Agent Trajectory Synthesis: Simulating multi-step tool-use conversations using MCP 3.1 servers to train specialized action models.

Strengths

  • Provider Agnostic: Switch easily between local backends (vLLM, Ollama) and commercial APIs (Anthropic, OpenAI, Gemini).
  • Enterprise Reliability: Handles API rate limits, connection retries, state caching, and step-by-step pipeline recovery.
  • Rich Library: Pre-built components for standard datasets (e.g., UltraFeedback, DEITA, self-instruct).
  • Hugging Face Hub Native: Directly loads from and pushes to the Hugging Face Hub.

Limitations

  • Cost Accumulation: Running large-scale data generation using commercial frontier model APIs can result in very high token costs.
  • Prompt Sensitivity: Quality is deeply tied to system prompt engineering; subtle model updates can alter generation distributions.

When to use it

  • When you need to scale fine-tuning data from hundreds of seeds to tens of thousands of highly varied instruction-response pairs.
  • To set up automated, reproducible "LLM-as-a-judge" data filtering and scoring systems.
  • When generating structured preference pairs (chosen vs. rejected) for DPO/RLHF alignment.

When not to use it

  • For basic data loading or simple filtering that can be accomplished with standard pandas or Hugging Face datasets scripts.
  • If you lack access to capable generator models (either local GPUs or commercial APIs).

Getting started

Installation

pip install distilabel[vllm,anthropic,openai]

Minimal Python Example

from distilabel.pipeline import Pipeline
from distilabel.steps import LoadDataFromHub
from distilabel.llms import AnthropicLLM

with Pipeline(name="hello-world") as pipeline:
    loader = LoadDataFromHub(repo_id="instruction-dataset")
    llm = AnthropicLLM(model="claude-5-1-sonnet")
    # ... define pipeline steps ...

CLI examples

# Run a declarative pipeline from a YAML configuration file
distilabel pipeline run --config my_pipeline.yaml

# Check the status of active pipelines
distilabel pipeline status

# List installed distilabel pipeline templates
distilabel templates list

API examples

Generating Evol-Instructions

Using Claude 5.1 to evolve instruction complexity over multiple iterations:

from distilabel.pipeline import Pipeline
from distilabel.steps import LoadDataFromHub
from distilabel.steps.tasks import EvolInstruction
from distilabel.llms import AnthropicLLM

with Pipeline(name="evol-instruct-pipeline") as pipeline:
    loader = LoadDataFromHub(repo_id="HuggingFaceH4/instruction-dataset")
    llm = AnthropicLLM(model="claude-5-1-sonnet")

    evolve = EvolInstruction(
        llm=llm,
        num_evolutions=2,
    )

    loader >> evolve

if __name__ == "__main__":
    pipeline.run()

Python (Preference Dataset Record Validation with Pydantic v2)

To ensure downstream training runs do not crash due to malformed JSON, synthetic records (such as UltraFeedback-style preference scores or DPO pairs) should be validated using Pydantic v2:

import json
from typing import List, Dict, Any, Optional, Literal
from pydantic import BaseModel, Field, field_validator

# 1. Define the validation schema for a Preference alignment pair
class PreferenceEvaluation(BaseModel):
    judge_model: str = Field(..., serialization_alias="judgeModel", validation_alias="judgeModel")
    score: float = Field(..., ge=1.0, le=10.0)
    critique: str

class PreferenceDatasetRecord(BaseModel):
    record_id: str = Field(..., serialization_alias="recordId", validation_alias="recordId")
    instruction: str
    chosen_response: str = Field(..., serialization_alias="chosenResponse", validation_alias="chosenResponse")
    rejected_response: str = Field(..., serialization_alias="rejectedResponse", validation_alias="rejectedResponse")
    evaluation: PreferenceEvaluation

    @field_validator("chosen_response", "rejected_response")
    @classmethod
    def validate_responses(cls, v: str) -> str:
        if len(v.strip()) < 10:
            raise ValueError("Response text is too short to be viable training data.")
        return v

# 2. Simulated Distilabel pipeline output payload for a single synthetic instruction
distilabel_record_payload = {
    "recordId": "rec-distilabel-552",
    "instruction": "Explain quantum superposition in simple words.",
    "chosenResponse": "Imagine a spinning coin. While spinning, it's both heads and tails at once. That's superposition.",
    "rejectedResponse": "It is a linear combination of all possible eigenstates in a Hilbert space prior to measurement.",
    "evaluation": {
        "judgeModel": "Claude 5.1",
        "score": 9.5,
        "critique": "The chosen response uses a great spinning coin analogy, making it highly accessible compared to the rejected jargon."
    }
}

# 3. Perform strict validation
try:
    record = PreferenceDatasetRecord(**distilabel_record_payload)
    print("Distilabel synthetic dataset record validated successfully via Pydantic v2!")
    print(f"Record ID: {record.record_id}")
    print(f"Instruction: {record.instruction}")
    print(f"Chosen (Score: {record.evaluation.score}): {record.chosen_response}")
    print(f"Critique: {record.evaluation.critique}")
except Exception as e:
    print(f"Record validation failed: {e}")
  • Fine-tuning Open Models — The primary training method utilizing generated data.
  • Unsloth — For ultra-fast single-GPU model fine-tuning.
  • Axolotl — For multi-GPU configuration-driven training.
  • vLLM — Highly optimized generation engine used as a backend.
  • Ollama — Simple local inference backend.
  • Glaive — Enterprise synthetic agentic data platform.
  • Model Context Protocol (MCP) — Used to power tool usage in synthetic agents.
  • Instructor — For structured output extraction.

Sources / references

Contribution Metadata

  • Last reviewed: 2027-01-07
  • Confidence: high