Local-first AI Knowledge Base — Fast · Private · BYOK
Rust + Tauri v2 desktop app for private RAG (Retrieval-Augmented Generation) with your own LLM API key.
Features · Quick Start · Architecture · Tech Stack · 📖 中文文档
English (current) · 简体中文
EchoMind is a desktop application that lets you chat with your local documents using any OpenAI-compatible LLM. Your files never leave your machine — parsing, chunking, embedding, and vector storage all happen locally. You bring your own API key (BYOK), so you stay in full control of costs and data.
Core value proposition: Rust speed · Privacy by design · MIT open source
| EchoMind | AnythingLLM | Open WebUI | Jan | |
|---|---|---|---|---|
| Runtime | Rust + Tauri (~15 MB) | Electron (~150 MB+) | Python + Docker | Tauri |
| RAM usage | Very low | High | Medium-high | Medium |
| RAG knowledge base | ✅ | ✅ | ✅ | ❌ |
| BYOK (own API key) | ✅ | ✅ | ✅ | Local models |
| Local embedding (ONNX) | ✅ | ✅ | ✅ | ❌ |
| Local LLM (GGUF) | ✅ | ❌ | ❌ | ✅ |
| Database encryption | ✅ SQLCipher | ❌ | ❌ | ❌ |
| License | MIT | MIT | MIT | MIT |
| Zero server cost | ✅ | ❌ | ❌ | ✅ |
- Multi-format support — Markdown, text, code files (Rust/TS/Python/Go), HTML, PDF, DOCX, PPTX, EPUB, XLSX/CSV
- 100% local processing — parsing, chunking, embedding, and vector storage all on-device
- Semantic chunking — paragraph → sentence → clause recursive splitting with code block preservation
- Section-aware splitting — Markdown heading hierarchy → section-boundary chunks
- ONNX embedding — all-MiniLM-L6-v2 (384-dim, ~30 MB) via fastembed; no external API
- Custom embedding models — upload custom ONNX models
- SQLite vector store — WAL mode, FTS5 full-text index, zero configuration
- HNSW index — approximate nearest neighbor for sub-linear search
- File deduplication — MD5 content hashing prevents duplicate imports
- Crash recovery — interrupted indexing tasks auto-recovered on restart
- Hybrid retrieval — vector search + BM25 keyword matching → RRF fusion
- Cross-Encoder reranking — bge-reranker-base for precision boost
- HyDE query rewriting — LLM generates hypothetical answer → embed → search
- Knowledge graph — entity extraction + relation mining → graph traversal retrieval
- Agentic RAG — ReAct multi-step reasoning with parallel tool execution
- Progressive context injection — start with top-2 chunks, expand if insufficient
- Speculative RAG — draft model generates, verify model confirms
- Retrieval memory — adaptive method selection based on query type
- Semantic cache — three-tier cache (exact / semantic / retrieval) for instant responses
- Context compaction — LLM-based history summarization replacing truncation
- Progress phases — preparing → retrieving → generating, no blank wait
- Cancellable generation — stop mid-response; partial content preserved
- Multi-turn conversation — full chat history with auto-extracted titles
- Branch tree — ChatGPT-style visual conversation branching
- GGUF inference — mistral.rs v0.9.0, pure Rust
- GPU acceleration — Metal (macOS) / CUDA (NVIDIA) / Accelerate (Apple BLAS)
- PagedAttention — efficient KV cache management for long conversations
- Sampling parameters — temperature, top-p, top-k, repetition penalty
- KV cache persistence — save/restore across sessions
- Custom GEMV kernel — self-developed quantization inference (Q4_0/Q4_K/Q8_0/Q8_K)
- Weight repacking — CPU cache-friendly Tile-Major layout
- Layer prefetch —
madvise(MADV_WILLNEED)streaming prefetch - RAM budget — LRU eviction + system memory awareness
- Model download manager — pause/resume/cancel + crash recovery
- SQLCipher encryption — AES-256 transparent database encryption
- Argon2id key derivation — memory-hard KDF (m=19456KB, t=2, p=1) + PBKDF2 fallback
- PII detection & redaction — 8 types (email, phone, ID card, bank card, IP, SSN, passport, intl phone)
- Audit hash-chain — SHA-256 linked audit logs with tamper detection
- Auto-lock — idle timeout → locked state
- Brute-force protection — 5 failed attempts → exponential backoff
- Clipboard auto-clear — sensitive data auto-cleared after timeout
- API key masking —
****+ last 4 chars, never plaintext - Security posture — Dangerous / Auto / Strict tiers with shadow screening
forbid(unsafe_code)in production crates — memory safety guaranteed by Rust
- Markdown with code syntax highlighting (highlight.js)
- Mermaid diagrams — flowcharts, sequence diagrams, Gantt charts
- KaTeX math — inline and block LaTeX equations
- Chart.js — interactive data visualizations
- Bidirectional wiki-links — Obsidian-style
[[wiki-link]]with backlinks - No CDN — all frontend libraries locally vendored
- AutoDream — background idle tidying: duplicate detection, contradiction discovery
- Persistent memory — three-tier (Wing/Hall/Room) with LLM consolidation
- Code symbol search — tree-sitter AST extraction (Rust/TS/Python/Go)
- Code execution sandbox — Python/Node with timeout/memory/network limits
- DAG workflow — visual workflow builder with template management
- Web search fusion — DuckDuckGo Instant Answer + RRF local fusion
- Knowledge graph visualization — D3.js force-directed graph with community detection
- PDF export —
window.print()zero-dependency export - Conversation export — Markdown format with source citations
- Folder sync — file watcher + incremental sync (add/update/delete)
- macOS (Apple Silicon + Intel)
- Windows x64
- Linux x64
- Built with Tauri v2 — native performance, not Electron
- Initial alpha release with core RAG functionality
- All features fully open, no restrictions
Hexagonal (ports & adapters) architecture with 8 crates. Dependencies flow strictly inward:
crates/models → crates/prompt → crates/core → crates/infra → crates/tauri-app
(contracts) (prompts) (ports+logic) (adapters) (assembly)
↑
crates/compact
crates/context
| Crate | Role |
|---|---|
crates/models |
Domain contracts (Document, Chunk, ChatMessage, Conversation, etc.) |
crates/prompt |
Prompt building: SegmentedPrompt, RAG/Agent prompt construction, Cache policy |
crates/compact |
Context compaction engine: LLM-based history summarization |
crates/context |
System context registry: epoch management, durable baseline |
crates/core |
Port traits + business logic; chat engine, import service, security |
crates/infra |
Adapters: SqliteStorage, LocalEmbedder, OpenAIProvider, HNSW, LocalLlmEngine, OCR, VLM |
crates/tauri-app |
Tauri shell, 190+ IPC commands, AppState |
Frontend: Single-file SPA (ui/index.html) — 50 ES modules bundled via esbuild. Tailwind CSS (local JIT), vanilla JavaScript. No CDN, no framework.
Document Import (100% local):
import_files → Loader.load() → MD5 dedup → Splitter.split()
→ Storage.add_document() + add_chunk() → Embedder.embed_batch()
→ Storage.add_embedding() → EntityExtractor → doc-status-changed event
RAG Query (BYOK):
chat → embed query (local ONNX) → hybrid search (vector + BM25 → RRF)
→ rerank (bge-reranker) → build RAG prompt → LLM chat_stream (SSE)
→ chat_token events → chat_done → persist exchange
- Rust ≥ 1.85 (Edition 2024) — install
- Node.js ≥ 18 (for E2E tests, optional)
# Clone
git clone https://github.com/lisering/EchoMind.git
cd EchoMind
# Build all crates
cargo build
# Run in dev mode (hot-reload)
cargo tauri dev
# Build with all features
cargo build --features proNote: First build takes 5–10 minutes due to ML dependencies (fastembed/ort/tokenizers) compiled at
opt-level = 3. Incremental builds are fast.
# Rust unit + integration tests (987 tests)
cargo test
# Lint (zero warnings policy)
cargo clippy --all-targets -- -D warnings
cargo fmt --check
# Supply chain security
cargo audit
cargo deny check
# Frontend type check
npx tsc --noEmit
# Frontend build
node scripts/build-ui.mjs- Launch EchoMind
- Configure — Settings → enter your LLM provider details (API key, base URL, model name)
- Import — Drag files into the window (all formats supported)
- Wait for indexing to complete (local ONNX embedding — watch the progress badge)
- Chat — Type your question and get streaming answers with source citations
Any OpenAI-compatible API endpoint works:
| Provider | Base URL | Notes |
|---|---|---|
| OpenAI | https://api.openai.com/v1 |
Default |
| Anthropic | https://api.anthropic.com/v1 |
Via OpenAI-compatible endpoint |
| DeepSeek | https://api.deepseek.com/v1 |
Popular in China |
| Ollama (local) | http://localhost:11434/v1 |
Empty API key |
| LM Studio | http://localhost:1234/v1 |
Local model runner |
| Local GGUF | — | Built-in mistral.rs engine, no external service |
| Any OpenAI-compatible | Custom base URL | If it speaks OpenAI API, it works |
| Layer | Technology | Details |
|---|---|---|
| Language | Rust (Edition 2024) | Native async fn in trait, no async-trait macro |
| Desktop framework | Tauri v2 | Smaller, faster, more secure than Electron |
| Embedding | fastembed (ONNX Runtime) | all-MiniLM-L6-v2, 384-dim, ~30 MB |
| Local LLM | mistral.rs v0.9.0 | GGUF, Metal/CUDA, PagedAttention |
| Vector store | SQLite (rusqlite + r2d2) | WAL mode, FTS5, SQLCipher AES-256 |
| LLM API | OpenAI-compatible | SSE streaming, 30s connection timeout |
| Frontend | Vanilla JS ES modules | esbuild IIFE bundle, no React/Vue/Svelte |
| Rendering | marked.js, DOMPurify, highlight.js | + Mermaid, KaTeX, Chart.js, D3.js |
- Clippy — zero warnings policy (
-D warnings) with deny lints forunwrap_used,expect_used,panic,unreachable,todo,unimplemented - TDD — test-first development, 987 tests, unit tests co-located with source
- Supply chain —
cargo audit+cargo deny checkon every CI run - Documentation — all public types have
///doc comments - No
unsafe—forbid(unsafe_code)in production crates
MIT License — see LICENSE.
- Tauri — for the amazing Rust desktop framework
- fastembed — for making ONNX embedding effortless
- SQLite — for the world's most reliable embedded database
- mistral.rs — for pure Rust LLM inference
- The Rust community — for building tools that make software fast and safe
Made with ❤️ and Rust