| Local Embedding Models |
https://nomic.ai/embeddings |
tool, infrastructure, rag |
integrated |
tools/infrastructure/local-embeddings.md |
Offline local embedding models for RAG |
| LanceDB |
https://github.com/lancedb/lancedb |
tool, infrastructure, vector-db |
integrated |
tools/infrastructure/lancedb.md |
Embedded serverless columnar vector store |
| Model Context Protocol Servers |
https://modelcontextprotocol.io/servers |
tool, automation, mcp |
integrated |
tools/automation_orchestration/mcp-servers.md |
Self-hostable MCP server ecosystem |
| PrivateGPT |
https://github.com/zylon-ai/private-gpt |
tool, ai_knowledge, rag |
integrated |
tools/ai_knowledge/privategpt.md |
Fully local air-gapped RAG |
| Ramalama |
https://github.com/containers/ramalama |
tool, infrastructure, containers |
integrated |
tools/infrastructure/ramalama.md |
Container-native local model runner |
| llama-swap |
https://github.com/sammcj/llama-swap |
tool, infrastructure, model-routing |
integrated |
tools/infrastructure/llama-swap.md |
Model proxy for hot-swapping local GGUF models on demand |
| text-generation-webui |
https://github.com/oobabooga/text-generation-webui |
tool, infrastructure, web-ui |
integrated |
tools/infrastructure/text-generation-webui.md |
Local web UI and multi-backend inference server |
| Otaku |
https://www.reddit.com/r/LocalLLaMA/comments/1w85blf/otaku_an_llm_frontend/ |
tool |
integrated |
tools/ai_knowledge/otaku.md |
An LLM frontend interface. |
| nInfer |
https://www.reddit.com/r/LocalLLaMA/comments/1w8f8fa/ninfer_fork_555k_contextfp4_for_5090_with_yarn/ |
infrastructure |
integrated |
tools/infrastructure/ninfer.md |
Inference engine fork supporting high context and FP4 quantization. |
| Vortex |
https://www.infoq.com/presentations/vortex-columnar-file-format-gpu-streaming/ |
infrastructure |
integrated |
tools/infrastructure/vortex.md |
A columnar file format designed for GPU streaming. |