| ⭐ Ollama |
Go |
MIT |
Get up and running with Llama 3.3, Mistral, Qwen, and other large language models locally with a simple CLI and REST API. |
GitHub |
| vLLM |
Python / C++ |
Apache-2.0 |
A high-throughput and memory-efficient inference and serving engine for LLMs with PagedAttention. |
GitHub |
| YYLO |
TypeScript |
MIT |
Command-line orchestrator for AI coding agents: runs Claude Code, Codex CLI, and Gemini CLI in parallel git worktrees with Kanban task state, typed merge flows, and per-agent skills. |
GitHub |
| LocalAI |
Go / C++ |
MIT |
Local, OpenAI-compatible REST API for running LLMs, audio transcription, image generation, and embeddings on commodity CPU/GPU. |
GitHub |
| AnythingLLM |
JavaScript |
MIT |
The all-in-one Desktop & Docker AI application with full RAG pipeline, multi-agent workspaces, and complete privacy. |
GitHub |
| Jan |
TypeScript / C++ |
AGPL-3.0 |
Open-source desktop alternative to ChatGPT that runs 100% offline on local hardware with zero data leakage. |
GitHub |
| llama.cpp |
C / C++ |
MIT |
High-performance LLM inference in pure C/C++ with minimal setup and state-of-the-art quantization across Apple Silicon and GPUs. |
GitHub |
| LiteLLM |
Python |
MIT |
Call 100+ LLM APIs using the standardized OpenAI input/output format with load balancing, proxy, and spend tracking. |
GitHub |
| PrivateGPT |
Python |
Apache-2.0 |
100% private, local document question and answering engine using open LLMs without internet connection or data leaks. |
GitHub |
| GPT4All |
C++ / Python |
MIT |
Open-source ecosystem of consumer hardware-optimized LLMs running locally with desktop client and Python bindings. |
GitHub |
| Khoj |
Python |
GPL-3.0 |
Open-source, personal AI second brain that searches, reasons over, and answers questions from your personal docs and notes. |
GitHub |
| Chroma |
Python |
Apache-2.0 |
The AI-native open-source embedding database with simple Python and JavaScript client SDKs. |
GitHub |
| Qdrant |
Rust |
Apache-2.0 |
Vector similarity search engine with extended filtering support, payload indexing, and high-performance Rust core. |
GitHub |
| Milvus |
Go / C++ |
Apache-2.0 |
Cloud-native open-source vector database built for massive scale similarity search and AI vector retrieval. |
GitHub |
| Weaviate |
Go |
BSD-3-Clause |
Open-source vector database that stores both objects and vectors, allowing for seamless hybrid search. |
GitHub |
| Bark |
Python |
MIT |
Transformer-based text-to-audio model created by Suno that can generate highly realistic speech, music, and ambient noise. |
GitHub |
| Whisper.cpp |
C/C++ |
MIT |
High-performance inference of OpenAI's Whisper automatic speech recognition in pure C/C++ without external dependencies. |
GitHub |
| ComfyUI |
Python |
GPL-3.0 |
The most powerful and modular stable diffusion GUI and backend with graph/nodes-based workflow interface. |
GitHub |
| TGI (Text Generation Inference) |
Rust / Python |
Apache-2.0 |
Hugging Face's production-ready toolkit for deploying and serving Large Language Models at scale with Tensor Parallelism. |
GitHub |
| SGLang |
Python / C++ |
Apache-2.0 |
Fast serving engine for large language models and vision-language models with RadixAttention for multi-turn caching. |
GitHub |
| TensorRT-LLM |
C++ / Python |
Apache-2.0 |
NVIDIA's specialized library providing state-of-the-art acceleration and optimized execution for LLM inference. |
GitHub |
| ExLlamaV2 |
Python / C++ |
MIT |
Fast inference library for running local LLMs on modern NVIDIA GPUs with EXL2 quantization. |
GitHub |
| llamafile |
C++ |
Apache-2.0 |
Distribute and run LLMs with a single multi-GB binary that runs across Windows, macOS, Linux, and FreeBSD. |
GitHub |
| Aphrodite Engine |
Python |
Apache-2.0 |
High-throughput, memory-efficient LLM inference engine supporting many architectures and quantizations. |
GitHub |
| ⭐ Unsloth |
Python |
Apache-2.0 |
Finetune Llama 3.3, Mistral, and Qwen 2x-5x faster with 70% less memory using optimized hand-written GPU kernels. |
GitHub |
| Axolotl |
Python |
Apache-2.0 |
Post-training framework for fine-tuning hundreds of open-source language models with standard YAML configurations. |
GitHub |
| LLaMA-Factory |
Python |
Apache-2.0 |
Unified web UI and CLI for easy fine-tuning and evaluation of 100+ LLMs and multi-modal models. |
GitHub |
| AutoAWQ |
Python / C++ |
Apache-2.0 |
Activation-aware Weight Quantization package for 4-bit LLM quantization with 3x speedup. |
GitHub |
| AutoGPTQ |
Python / CUDA |
MIT |
Easy-to-use LLMs quantization package with user-friendly APIs, based on GPTQ algorithm. |
GitHub |
| bitsandbytes |
C++ / Python |
MIT |
Accessible 8-bit and 4-bit optimizers and quantization functions for deep learning on CUDA and ROCm. |
GitHub |
| PEFT |
Python |
Apache-2.0 |
Hugging Face's state-of-the-art Parameter-Efficient Fine-Tuning library (LoRA, QLoRA, Prefix Tuning). |
GitHub |
| TRL |
Python |
Apache-2.0 |
Transformer Reinforcement Learning library with Supervised Fine-Tuning (SFT), DPO, and PPO training loops. |
GitHub |
| Stable Diffusion WebUI |
Python |
AGPL-3.0 |
The ubiquitous browser interface for Stable Diffusion based on Gradio library with extensive extensions. |
GitHub |
| InvokeAI |
Python / TypeScript |
Apache-2.0 |
Professional generative AI engine and canvas for creative visual artists and digital designers. |
GitHub |
| Fooocus |
Python |
GPL-3.0 |
Focus on prompting and generating without complex parameter tuning, inspired by Midjourney simplicity. |
GitHub |
| LLaVA |
Python |
Apache-2.0 |
Visual instruction tuning toward large multimodal models: chat and reasoning over high-res images. |
GitHub |
| Segment Anything (SAM) |
Python |
Apache-2.0 |
Meta's foundation model for image segmentation, allowing promptable cutouts of any object with one click. |
GitHub |
| YOLOv8 |
Python |
AGPL-3.0 |
Ultralytics YOLOv8 for real-time object detection, image segmentation, pose estimation, and classification. |
GitHub |
| PaddleOCR |
Python |
Apache-2.0 |
Awesome multilingual OCR toolkits supporting 80+ languages, table recognition, and document structure analysis. |
GitHub |
| ⭐ Kokoro-TTS |
Python |
Apache-2.0 |
Lightweight, open-weight text-to-speech model producing studio-grade human voices with only 82M parameters. |
GitHub |
| Piper TTS |
C++ |
MIT |
Fast, local neural text to speech system that sounds great and is optimized for Raspberry Pi and low-power hardware. |
GitHub |
| ChatTTS |
Python |
CC-BY-NC-4.0 |
A generative text-to-speech model specifically designed for conversational scenarios and dialogue synthesis. |
GitHub |
| Faster-Whisper |
Python |
MIT |
Reimplementation of OpenAI's Whisper model using CTranslate2, delivering up to 4x faster transcription. |
GitHub |
| ⭐ pgvector |
C |
PostgreSQL |
Open-source vector similarity search extension for PostgreSQL, supporting exact and approximate nearest neighbor. |
GitHub |
| Faiss |
C++ / Python |
MIT |
Meta's library for efficient similarity search and clustering of dense vectors, indexing billions of vectors on GPUs. |
GitHub |
| LanceDB |
Rust / Python |
Apache-2.0 |
Developer-friendly, serverless open-source vector database for multi-modal AI with zero-copy storage. |
GitHub |
| Vespa |
C++ / Java |
Apache-2.0 |
The open big data serving engine: store, search, organize, and machine-learn over large vector and text datasets in real-time. |
GitHub |
| Annoy |
C++ / Python |
Apache-2.0 |
Spotify's C++ library with Python bindings to search for points in space that are close to a given query point. |
GitHub |
| AutoRound |
Python |
Apache-2.0 |
Intel's advanced weight-only quantization algorithm for low-bit LLM optimization without accuracy degradation. |
GitHub |
| FastChat |
Python |
Apache-2.0 |
An open platform for training, serving, and evaluating LLM-based chatbots, the core engine behind LMSYS Chatbot Arena. |
GitHub |
| GPTCache |
Python |
MIT |
Semantic cache for LLM queries to cut costs and speed up response time up to 10x using vector similarity. |
GitHub |