Local-first retrieval-augmented generation in a single Go binary. Hybrid BM25 + vector search over your own files, with context packing and cited answers - no vector database, no Python, no server to run.
- Hybrid retrieval - BM25 and vector embeddings fused with Reciprocal Rank Fusion, each arm independently searchable
- Answers with citations -
rag askretrieves and answers in one LLM call, every claim tagged[file:lines] - HyDE query expansion - searches with a hypothetical answer instead of the question, cached so repeats are free
- Runs anywhere - one static binary, a single BoltDB file, or in the browser via WebAssembly
- Measured, not guessed - a built-in ground-truth benchmark reports recall and MRR
- Compact index - packed binary vectors with optional float16 and int8 quantization
Embeddings run locally through Ollama; answer generation uses a hosted API (DeepSeek or OpenAI).
$ rag ask -d ./books -q "How did Ned Stark die" --fast --hydeNed Stark was executed by beheading on the orders of King Joffrey, carried out by Ser Ilyn Payne [4.txt:L21107-21118][2.txt:L3019-3024][2.txt:L17370-17375].
3,601 tokens (with RAG) vs ~2.8M tokens (without RAG)
$ rag ask -d ./books -q "How did Joffrey die" --fast --hydeJoffrey died from poison, which was made to appear as though he choked on food, but maesters found no obstruction in his throat after it was slit open [3.txt:21634-21642]. He was poisoned at his own wedding feast [3.txt:26024-26035].
3,419 tokens (with RAG) vs ~2.8M tokens (without RAG)
$ rag ask -d ./books -q "How did Robb Stark die at the Red Wedding" --fast --hydeRobb Stark was betrayed and murdered by the Freys and Boltons at the Twins during his uncle Edmure Tully's wedding to Roslin Frey, an event known as the Red Wedding [4.txt:21114-21127]. The Freys conspired with Roose Bolton after Robb broke his pledge to marry a Frey daughter by wedding Lady Jeyne Westerling instead [4.txt:20767-20784].
3,954 tokens (with RAG) vs ~2.8M tokens (without RAG)
Answers generated by DeepSeek over hybrid retrieval (BM25 + mxbai-embed-large), one LLM call each. Token counts reported by the API.
go build -o rag ./cmd/ragFirst, create a rag.yaml config in your content directory:
# books/rag.yaml
index:
includes:
- "**/*.txt"
chunk_tokens: 140 # ~540 chars; see Tuning vector quality
chunk_overlap: 30
stemming: true
retrieve:
top_k: 20
mmr_lambda: 0.7
hybrid_enabled: true
bm25_weight: 0.35
embedding:
enabled: true
provider: ollama
model: mxbai-embed-large
dimension: 1024
include_path: false # true for code, false for prose
llm:
provider: openai # OpenAI-compatible chat endpoint
model: deepseek-web
base_url: http://127.0.0.1:8787/v1
api_key_env: "" # empty = keyless gateway, no Authorization header
pack:
token_budget: 4000The default endpoint above needs no key. To point at a keyed provider instead, set
base_url (or drop it for the provider default) and name the env var holding the key:
llm:
provider: openai
model: gpt-4o-mini
api_key_env: OPENAI_API_KEYPut that key in a .env file next to the config (or anywhere up the directory tree) -
it is read automatically, no export needed:
OPENAI_API_KEY=sk-...
# Index a directory (also generates embeddings)
rag index /path/to/books
# Ask a question and get an answer with citations
rag ask -d /path/to/books -q "how does auth work" --fast --hyde
# Or just search, and use the results yourself
rag query -d /path/to/books -q "authentication handler"
# Pack context to paste into another LLM
rag pack -d /path/to/books -q "how does auth work" -b 4000 -o context.json
# Generate a prompt for manual LLM orchestration
rag runprompt --runtime --ctx context.json -q "Explain the auth flow"Index files in a directory for later retrieval. Creates a .rag/index.db file.
rag index . # Index current directory
rag index /path/to/project # Index specific directoryFlags:
-d, --dir- Root directory (default: current directory)--config- Path to config file (default:./rag.yaml)--force-embed- Re-embed every chunk instead of reusing existing vectors
Search indexed files using BM25 retrieval with MMR deduplication.
rag query -q "database connection"
rag query -q "error handling" --top-k 10 --json
rag query -q "how to handle errors" --semanticFlags:
-q, --query- Search query (required)-k, --top-k- Number of results (default from config)--json- Output as JSON--no-mmr- Disable MMR reranking--semantic- Use embedding-only search (no BM25); mutually exclusive with--lexical--lexical- Use BM25-only search (no embeddings)--explain- Print which retrieval arms ran and how many candidates each produced--hyde- Expand the query with one LLM-generated hypothetical answer (1 API call, cached)-c, --context- Expand results by N lines before/after
Retrieve context and have a hosted LLM answer the question, with citations. This is the full pipeline: hybrid retrieval plus generation.
rag ask -q "how does authentication work"
rag ask -q "how does authentication work" --fast # exactly one LLM call
rag ask -q "how does authentication work" --hyde # better retrieval, cached probeFlags:
-q, --query- Question (required)--fast- Search once and answer: exactly one LLM call--expand- Expand the query with the LLM first (+1 call)--hyde- Expand retrieval with a hypothetical answer (+1 call, cached)--max-iters- Maximum retrieve/evaluate rounds (default 2)--semantic- Vector search only, no BM25--lexical- BM25 only, no embeddings-k, --top-k,-b, --budget,--explain
Every run prints a stats block: retrieval mode, rounds used, LLM calls, and input/output/total tokens (reported by the API when the provider returns a usage field).
The API key is read from .env automatically - see the llm: section under Configuration.
Vectors dominate the index. Three encodings are available, switchable in place with
rag compact --encoding - no re-embedding required:
| encoding | index size | recall@10 | MRR | top-20 order |
|---|---|---|---|---|
float32 (default) |
193 MB | 6/14 | 0.329 | baseline |
float16 |
132 MB (-32%) | 6/14 | 0.329 | identical |
int8 |
108 MB (-44%) | 6/14 | 0.329 | identical |
Measured on a 20,340-vector index of ~7.5 MB of prose. Both lossy encodings produced byte-identical rankings; int8 score deviations were around 0.0007. Query latency was unchanged at 0.13s.
rag compact -d ./books --encoding int8One caveat: conversion is one-way. Going from int8 back to float32 restores the
format but not the discarded precision - you would need to re-index. Measure with
rageval on your own corpus before converting, since these numbers come from one prose
corpus and quantization error is corpus-dependent.
Set embedding.vector_encoding to build new indexes directly in a given encoding.
CLAUDE.md requires measuring changes to the embedding model, chunk size or embedded-text
format against a ground-truth set rather than guessing. cmd/rageval is that harness.
go run ./cmd/rageval -corpus /path/to/indexed -k 1014 questions, recall, MMR disabled
MODE RECALL MRR RANKS
lexical 6/14 0.201 [- 4 - - - - - - 5 9 1 - 4 1]
semantic 6/14 0.329 [- 1 - - - - - - 1 - 1 2 9 1]
hybrid 6/14 0.310 [- 1 - - - - - - 2 - 1 3 2 1]
Each question is paired with an anchor phrase that must appear in a correctly retrieved
passage. A - means the passage was not found within -k. Add -modes hyde to include
HyDE, -v to print per-question ranks.
The bundled set in cmd/rageval/questions.json covers A Song of Ice and Fire; replace it
with questions and anchors for your own corpus.
Shrink the index on disk. Rewrites legacy JSON vector records as packed binary and reclaims free pages. No re-embedding, no change to results.
rag compact -d /path/to/content
# Rewrote 20340 vector records as packed binary
# Index: 579.3 MB -> 192.7 MB (66.7% smaller)Indexes created after this change already use the binary format; run it once on older indexes, and any time deletions have left free pages behind.
Pack relevant chunks into compressed context that fits a token budget.
rag pack -q "authentication flow" -b 2000
rag pack -q "API endpoints" -o context.json
rag pack -q "session handling" --lexicalFlags:
-q, --query- Search query (required)-b, --budget- Token budget (default from config)-o, --output- Output file (default: stdout)-k, --top-k- Candidate pool size
Generate formatted prompts from templates for manual LLM orchestration.
# Runtime prompt for question answering
rag runprompt --runtime --ctx context.json -q "How does auth work?"
# Builder prompt for context compression
rag runprompt --builder --ctx context.jsonFlags:
--runtime- Use runtime (answering) prompt template--builder- Use builder (compression) prompt template--ctx- Path to packed context JSON file (required)-q, --query- Override query for runtime prompt
Create a rag.yaml file in your project root:
index:
includes:
- "**/*.go"
- "**/*.py"
- "**/*.js"
- "**/*.ts"
- "**/*.md"
excludes:
- "**/node_modules/**"
- "**/vendor/**"
- "**/.git/**"
stemming: true
chunk_tokens: 512
chunk_overlap: 50
k1: 1.2
b: 0.75
retrieve:
top_k: 20
mmr_lambda: 0.7
dedup_jaccard: 0.8
pack:
token_budget: 4000
output: json
logging:
level: info| Section | Option | Description | Default |
|---|---|---|---|
index |
includes |
Glob patterns for files to index | Common code extensions |
index |
excludes |
Glob patterns to exclude | node_modules, vendor, .git |
index |
stemming |
Enable Porter stemming | true |
index |
chunk_tokens |
Max tokens per chunk | 512 |
index |
chunk_overlap |
Token overlap between chunks | 50 |
index |
k1 |
BM25 k1 parameter | 1.2 |
index |
b |
BM25 b parameter | 0.75 |
retrieve |
top_k |
Default number of results | 20 |
retrieve |
mmr_lambda |
MMR relevance vs diversity (0-1) | 0.7 |
retrieve |
dedup_jaccard |
Jaccard threshold for dedup | 0.8 |
retrieve |
hybrid_enabled |
Run the vector arm alongside BM25 | false |
retrieve |
rrf_k |
RRF fusion constant | 60 |
retrieve |
bm25_weight |
BM25 vs vector balance (0-1) | 0.5 |
embedding |
enabled |
Generate and use embeddings | false |
embedding |
provider |
ollama, openai, jina, deepseek, mock |
openai |
embedding |
model |
Embedding model name | text-embedding-3-small |
embedding |
dimension |
Vector size; probed from the provider when possible | model default |
embedding |
include_path |
Prefix embedded chunks with path:lines |
true |
embedding |
vector_encoding |
float32, float16 or int8 (see Vector encoding) |
float32 |
llm |
provider |
openai or deepseek; picks the default endpoint |
openai |
llm |
base_url |
Override the endpoint (any OpenAI-compatible server) | provider default |
llm |
model |
Chat model name | deepseek-web |
llm |
api_key_env |
Env var holding the key; also read from .env. Empty sends no Authorization header, and then base_url is required |
"" |
pack |
token_budget |
Default token budget | 4000 |
To enable semantic search alongside BM25 keyword search, install Ollama and pull an embedding model:
# Install Ollama (macOS)
brew install ollama
# Start Ollama server
ollama serve
# Pull an embedding model
ollama pull mxbai-embed-largeEmbeddings may run locally; text generation may not - see the llm: section.
Then add embedding config to your rag.yaml:
embedding:
enabled: true
provider: ollama
model: mxbai-embed-large
dimension: 1024
include_path: false # true for code, false for prose
retrieve:
hybrid_enabled: true
rrf_k: 60 # RRF fusion parameter
bm25_weight: 0.35 # Balance between BM25 and vector (0-1)Re-index to generate embeddings:
rag index /path/to/contentHybrid search runs both arms independently and fuses their rankings with Reciprocal Rank Fusion:
score(c) = bm25_weight / (rrf_k + rank_bm25) + (1 - bm25_weight) / (rrf_k + rank_vector)
Because the arms run independently, a chunk that only the vector arm finds still reaches
the results — BM25 does not gate the candidate pool. Use --explain to see how many
candidates each arm produced:
rag query -q "how are sessions validated" --explain
# retrieval: hybrid (bm25 + vector, RRF) (model=mxbai-embed-large, vectors=20340)
# candidates: bm25=32 vector=80 fused=97If embeddings are unavailable (provider down, index not embedded, model changed), query and pack print a warning and fall back to BM25 rather than failing silently.
rag pack uses the same retrieval path as rag query, so hybrid search applies there too.
Retrieval quality depends far more on the embedding model and chunk size than on the fusion parameters. Measured on a ~7.5MB prose corpus (14 questions, recall@10):
| Setting | Vector recall@10 |
|---|---|
nomic-embed-text, ~2200-char chunks |
2/14 |
mxbai-embed-large, ~540-char chunks |
6/14 |
Raising k shows the answers are being found but ranked low - vector recall@200 is 10/14.
If you need them in the top 10, add a reranking stage over a deep candidate pool; tuning
bm25_weight does not help (recall@10 was flat at 6/14 across 0.2-0.65).
embedding.include_path prepends path:startLine-endLine to each embedded chunk. This
helps for code, where the path carries real signal, and hurts for prose - on the corpus
above it cost 0.12 MRR. It defaults to true; set it to false for prose.
Changing embedding.model, embedding.dimension, or embedding.include_path invalidates
stored vectors. The index records what it was embedded with and re-embeds automatically.
--hyde makes exactly one LLM call per query to write a hypothetical answer passage,
appends it to the query, and searches with both. The generated text is cached in the index,
so repeating a query costs zero API calls. If the LLM fails, the search silently falls back
to the plain query.
Generation goes through an OpenAI-compatible chat endpoint - the adapter deliberately has no local-model provider:
llm:
provider: openai
model: deepseek-web
base_url: http://127.0.0.1:8787/v1
api_key_env: ""
max_tokens: 400Note: RRF scores are much smaller than BM25 scores (typically
0.005-0.03). If you setretrieve.min_score_threshold, tune it for whichever mode you actually run — a threshold picked for BM25 scores will filter out every hybrid result.
Embeddings are incremental. Re-running rag index only embeds chunks that do not
already have a vector, and drops vectors for chunks that no longer exist. Use
rag index --force-embed to re-embed everything. Changing embedding.model triggers a
rebuild, discarding vectors from the old model.
Use --semantic flag to search using only vector embeddings (no BM25 keyword matching):
rag query -q "a noble man betrayed by those he trusted" --semanticSemantic search is useful for:
- Natural language questions (e.g., "how to handle errors gracefully")
- Conceptual queries where exact keywords may not appear
- Finding related content even when terminology differs
Requires embeddings to be enabled and indexed (see Hybrid Search section above).
- Walks directory with glob patterns
- Checks file modification times for incremental updates
- Splits files into line-based chunks with token awareness
- Tokenizes with optional Porter stemming
- Builds inverted index with term frequencies
- Stores in BoltDB (
.rag/index.db)
- Tokenizes and stems query
- Scores chunks using BM25:
score(q,c) = Σ IDF(t) × (tf × (k1+1)) / (tf + k1 × (1-b + b×|c|/avgDl)) - Applies MMR for diversity:
MMR(c) = λ × relevance(c) - (1-λ) × max_similarity(c, selected) - Returns ranked, deduplicated results
- Calculates utility = score / token_count
- Greedily selects chunks by utility until budget exhausted
- Merges adjacent chunks from same file
- Outputs JSON with citations (path, line range, relevance)
{
"query": "authentication",
"budget_tokens": 4000,
"used_tokens": 1250,
"snippets": [
{
"path": "/src/auth/handler.go",
"range": "L45-89",
"why": "BM25 score: 2.34",
"text": "func Authenticate(..."
}
]
}RAG can run entirely in the browser via WebAssembly (BM25 search only, no embeddings).
make build-wasm
# Or manually:
GOOS=js GOARCH=wasm go build -o examples/wasm/rag.wasm ./cmd/wasmcd examples/wasm
python3 -m http.server 8080
# Open http://localhost:8080// Index content
ragIndex("file.txt", "Your text content here...")
// Search (returns JSON string)
const results = JSON.parse(ragQuery("search term", 5))
// Clear index
ragClear()
// Get statistics
const stats = JSON.parse(ragStats())See examples/wasm/README.md for details.
cmd/rag/main.go # Entrypoint
cmd/wasm/main.go # WASM entrypoint
internal/
├── domain/ # Core entities (Document, Chunk, etc.)
├── port/ # Interfaces (IndexStore, Retriever, etc.)
├── usecase/ # Business logic # Business logic
│ ├── index.go # Indexing orchestration
│ ├── embed.go # Incremental embedding sync
│ ├── retrieve.go # Search with BM25 + MMR
│ ├── pack.go # Context packing
│ └── ask.go # Retrieve, judge, iterate, answer
└── adapter/
├── fs/ # File system walker
├── store/ # BoltDB: index, vectors, HyDE cache
├── analyzer/ # Tokenizer + Porter stemmer
├── chunker/ # Line-based and AST chunking
├── embedding/ # Embedding providers (ollama, openai, jina, mock)
├── llm/ # Hosted chat providers (deepseek, openai)
└── retriever/ # BM25, MMR, hybrid RRF, semantic, HyDE
MIT - see LICENSE.