Skip to content

Tube Archivist

Tube Archivist is a self-hosted YouTube archive that allows you to index and download YouTube videos, metadata, and comments to your own server.

What it is

Tube Archivist is an open-source media management system designed specifically for preserving YouTube content. As of early January 2027, it features a robust integration with the Deno runtime for enhanced download reliability, handles advanced cookie-passing techniques for age-restricted content, and provides advanced tools for metadata persistence and Elasticsearch-based searching. It integrates with AI agents via MCP 3.1 / FastMCP 3.1 to allow natural language triggers for channel archival and indexing across multi-agent workflows.

What problem it solves

YouTube videos can be deleted, made private, or censored without notice. Tube Archivist provides a way to build a permanent, offline, and searchable library of your favorite content, ensuring long-term access to tutorials, documentaries, and educational material while eliminating dependency on third-party platform availability and advertising.

Where it fits in the stack

It serves as a content preservation layer within the media management stack. It sits alongside general-purpose media servers like Jellyfin or Plex, but provides deep specialization for YouTube-specific metadata (comments, descriptions, subtitles) and automated channel monitoring for agentic workflows powered by Claude 5.6, GPT-5.6, Gemini 4.0 Ultra, and DeepSeek-V4.

Typical use cases

  • Automatically monitoring and downloading new videos from subscribed channels or playlists.
  • Archiving high-quality educational series for offline reference and AI-assisted summarization.
  • Building a private, ad-free "YouTube" experience for family members.
  • Researching video trends and comments at scale using its integrated search engine.
  • Preserving a record of metadata even if the original video is removed from YouTube.

Strengths

  • Comprehensive Preservation: Captures thumbnails, descriptions, comments, subtitles, and high-quality video files.
  • Advanced Search: Integrated Elasticsearch/OpenSearch for rapid full-text search across the entire archive.
  • Native Automation: Built-in scheduling for periodic channel rescans and downloads.
  • Metadata Resilience: Supports embedding all indexed metadata directly into the media files for reconstruction from the library files themselves.
  • Agentic Ready: Robust REST API for integration with tools like Claude 5.6 or DeepSeek-V4 for automated content analysis.

Limitations

  • Storage Intensive: Storing high-resolution video archives can consume terabytes of storage rapidly.
  • Resource Usage: Requires secondary containers for Redis and Elasticsearch, which can be memory-intensive.
  • Maintenance: Ongoing site layout changes on YouTube require frequent yt-dlp updates within the container.

When to use it

  • When you want to ensure permanent, offline access to specific YouTube content.
  • When you need to search across video descriptions and comments at scale for research or knowledge management.
  • To provide a private, curated media experience without tracking or platform bias.
  • For building a local knowledge base from video content via AI ingestion.

When not to use it

  • For occasional, one-off video downloads (use a simple CLI tool like yt-dlp).
  • If server resources (RAM/CPU/Storage) are extremely limited.
  • If you prefer a single-binary solution without the complexity of multiple service containers.

Getting started

Installation (Docker Compose)

Tube Archivist requires Redis and Elasticsearch/OpenSearch.

services:
  tubearchivist:
    container_name: tubearchivist
    image: bbilly1/tubearchivist:latest
    ports:
      - 8000:8000
    volumes:
      - /path/to/media:/youtube
      - /path/to/cache:/cache
    environment:
      - ES_URL=http://archivist-es:9200
      - REDIS_HOST=archivist-redis
      - HOST_UID=1000
      - HOST_GID=1000
      - TA_USERNAME=admin
      - TA_PASSWORD=password
    depends_on:
      - archivist-es
      - archivist-redis

  archivist-redis:
    image: redis/redis-stack-server:latest

  archivist-es:
    image: bbilly1/tubearchivist-es:latest
    environment:
      - "ELASTIC_PASSWORD=verysecret"
      - "ES_JAVA_OPTS=-Xms512m -Xmx512m"
      - "discovery.type=single-node"
    volumes:
      - /path/to/es:/usr/share/elasticsearch/data

Hello World

  1. Start the stack and navigate to http://localhost:8000.
  2. Log in with your configured credentials.
  3. Go to Downloads, paste a YouTube URL, and click Index and Download.

CLI examples

Tube Archivist is primarily managed via the UI, but the container provides tools for maintenance:

# Force an immediate rescan of the media directory to pick up external changes
docker exec tubearchivist python manage.py rescan

# Manually update yt-dlp to the latest version within the container
docker exec tubearchivist pip install -U yt-dlp

# Trigger a manual scan of channel tabs (shorts, streams, videos)
docker exec tubearchivist python manage.py ta_index_channel_tabs

API examples

Integrate Tube Archivist metadata parsing and download triggers into Python scripts or FastMCP 3.1 servers.

Python: FastMCP 3.1 Server for Automated Download and Video Validation

This example showcases a production-ready FastMCP 3.1 tool utilizing Pydantic v2 schemas to trigger video ingestion and validate download responses. It allows frontier models like Claude 5.6, GPT-5.6, and Gemini 4.0 Ultra to dynamically archive requested YouTube tutorials and extract descriptions.

import requests
from pydantic import BaseModel, Field, HttpUrl
from mcp.server.fastmcp import FastMCP

# Initialize FastMCP Server
mcp = FastMCP("TubeArchivistManager")

TA_URL = "http://localhost:8000/api"
HEADERS = {"Authorization": "Token YOUR_API_TOKEN"}

class ArchivalRequest(BaseModel):
    url: HttpUrl = Field(description="The valid YouTube video or channel URL to archive")
    bypass_cache: bool = Field(default=False, description="Whether to bypass local cache and force download")

class ArchivalResponse(BaseModel):
    success: bool = Field(description="Whether the archival task was successfully queued")
    message: str = Field(description="Response message from the Tube Archivist API")
    task_id: str = Field(default="", description="The unique ID of the triggered download task")

@mcp.tool()
def trigger_youtube_download(request_data: ArchivalRequest) -> str:
    """
    Submits a download request to the local Tube Archivist instance, validates input,
    and returns a Pydantic v2 validated status object.
    """
    payload = {
        "url": str(request_data.url),
        "bypass_cache": request_data.bypass_cache
    }

    try:
        response = requests.post(f"{TA_URL}/download/", headers=HEADERS, json=payload, timeout=10)

        if response.status_code == 201:
            data = response.json()
            result = ArchivalResponse(
                success=True,
                message="Video successfully queued for download",
                task_id=data.get("task_id", "N/A")
            )
        else:
            result = ArchivalResponse(
                success=False,
                message=f"Failed to queue video. API responded with status {response.status_code}: {response.text}"
            )

        return result.model_dump_json(indent=2)
    except requests.RequestException as e:
        return ArchivalResponse(
            success=False,
            message=f"Network exception when connecting to Tube Archivist API: {str(e)}"
        ).model_dump_json(indent=2)

if __name__ == "__main__":
    mcp.run()
  • Plex — For streaming your YouTube archive to smart TVs and mobile devices.
  • Jellyfin — The recommended open-source alternative for media streaming.
  • Audiobookshelf — For managing audio-only YouTube archives or podcasts.
  • Changedetection.io — To monitor YouTube channels for visual or metadata changes.
  • n8n — For advanced automation (e.g., notifying you in Element when a video is archived).
  • SearXNG — For private searching before adding videos to the archive.
  • Home Assistant — For dashboard integration and download notifications.
  • Tailscale — For secure remote access to your video library.
  • PO Token Management — For mitigating 403 errors during download.

Sources / References

Contribution Metadata

  • Last reviewed: 2027-01-07
  • Confidence: high