Skip to content

Latest commit

 

History

53 Commits

Folders and files

Repository files navigation

RAG Context Compressor

Local-first retrieval-augmented generation in a single Go binary. Hybrid BM25 + vector search over your own files, with context packing and cited answers - no vector database, no Python, no server to run.

Release Go License

  • Hybrid retrieval - BM25 and vector embeddings fused with Reciprocal Rank Fusion, each arm independently searchable
  • Answers with citations - rag ask retrieves and answers in one LLM call, every claim tagged [file:lines]
  • HyDE query expansion - searches with a hypothetical answer instead of the question, cached so repeats are free
  • Runs anywhere - one static binary, a single BoltDB file, or in the browser via WebAssembly
  • Measured, not guessed - a built-in ground-truth benchmark reports recall and MRR
  • Compact index - packed binary vectors with optional float16 and int8 quantization

Embeddings run locally through Ollama; answer generation uses a hosted API (DeepSeek or OpenAI).

Example: Querying A Song of Ice and Fire (4 books, ~2.8M tokens)

$ rag ask -d ./books -q "How did Ned Stark die" --fast --hyde

Ned Stark was executed by beheading on the orders of King Joffrey, carried out by Ser Ilyn Payne [4.txt:L21107-21118][2.txt:L3019-3024][2.txt:L17370-17375].

3,601 tokens (with RAG) vs ~2.8M tokens (without RAG)

$ rag ask -d ./books -q "How did Joffrey die" --fast --hyde

Joffrey died from poison, which was made to appear as though he choked on food, but maesters found no obstruction in his throat after it was slit open [3.txt:21634-21642]. He was poisoned at his own wedding feast [3.txt:26024-26035].

3,419 tokens (with RAG) vs ~2.8M tokens (without RAG)

$ rag ask -d ./books -q "How did Robb Stark die at the Red Wedding" --fast --hyde

Robb Stark was betrayed and murdered by the Freys and Boltons at the Twins during his uncle Edmure Tully's wedding to Roslin Frey, an event known as the Red Wedding [4.txt:21114-21127]. The Freys conspired with Roose Bolton after Robb broke his pledge to marry a Frey daughter by wedding Lady Jeyne Westerling instead [4.txt:20767-20784].

3,954 tokens (with RAG) vs ~2.8M tokens (without RAG)

Answers generated by DeepSeek over hybrid retrieval (BM25 + mxbai-embed-large), one LLM call each. Token counts reported by the API.


Installation

go build -o rag ./cmd/rag

Quick Start

First, create a rag.yaml config in your content directory:

# books/rag.yaml
index:
  includes:
    - "**/*.txt"
  chunk_tokens: 140      # ~540 chars; see Tuning vector quality
  chunk_overlap: 30
  stemming: true

retrieve:
  top_k: 20
  mmr_lambda: 0.7
  hybrid_enabled: true
  bm25_weight: 0.35

embedding:
  enabled: true
  provider: ollama
  model: mxbai-embed-large
  dimension: 1024
  include_path: false    # true for code, false for prose

llm:
  provider: openai       # OpenAI-compatible chat endpoint
  model: deepseek-web
  base_url: http://127.0.0.1:8787/v1
  api_key_env: ""        # empty = keyless gateway, no Authorization header

pack:
  token_budget: 4000

The default endpoint above needs no key. To point at a keyed provider instead, set base_url (or drop it for the provider default) and name the env var holding the key:

llm:
  provider: openai
  model: gpt-4o-mini
  api_key_env: OPENAI_API_KEY

Put that key in a .env file next to the config (or anywhere up the directory tree) - it is read automatically, no export needed:

OPENAI_API_KEY=sk-...
# Index a directory (also generates embeddings)
rag index /path/to/books

# Ask a question and get an answer with citations
rag ask -d /path/to/books -q "how does auth work" --fast --hyde

# Or just search, and use the results yourself
rag query -d /path/to/books -q "authentication handler"

# Pack context to paste into another LLM
rag pack -d /path/to/books -q "how does auth work" -b 4000 -o context.json

# Generate a prompt for manual LLM orchestration
rag runprompt --runtime --ctx context.json -q "Explain the auth flow"

Commands

rag index <path>

Index files in a directory for later retrieval. Creates a .rag/index.db file.

rag index .                      # Index current directory
rag index /path/to/project       # Index specific directory

Flags:

  • -d, --dir - Root directory (default: current directory)
  • --config - Path to config file (default: ./rag.yaml)
  • --force-embed - Re-embed every chunk instead of reusing existing vectors

rag query -q "<question>"

Search indexed files using BM25 retrieval with MMR deduplication.

rag query -q "database connection"
rag query -q "error handling" --top-k 10 --json
rag query -q "how to handle errors" --semantic

Flags:

  • -q, --query - Search query (required)
  • -k, --top-k - Number of results (default from config)
  • --json - Output as JSON
  • --no-mmr - Disable MMR reranking
  • --semantic - Use embedding-only search (no BM25); mutually exclusive with --lexical
  • --lexical - Use BM25-only search (no embeddings)
  • --explain - Print which retrieval arms ran and how many candidates each produced
  • --hyde - Expand the query with one LLM-generated hypothetical answer (1 API call, cached)
  • -c, --context - Expand results by N lines before/after

rag ask -q "<question>"

Retrieve context and have a hosted LLM answer the question, with citations. This is the full pipeline: hybrid retrieval plus generation.

rag ask -q "how does authentication work"
rag ask -q "how does authentication work" --fast    # exactly one LLM call
rag ask -q "how does authentication work" --hyde    # better retrieval, cached probe

Flags:

  • -q, --query - Question (required)
  • --fast - Search once and answer: exactly one LLM call
  • --expand - Expand the query with the LLM first (+1 call)
  • --hyde - Expand retrieval with a hypothetical answer (+1 call, cached)
  • --max-iters - Maximum retrieve/evaluate rounds (default 2)
  • --semantic - Vector search only, no BM25
  • --lexical - BM25 only, no embeddings
  • -k, --top-k, -b, --budget, --explain

Every run prints a stats block: retrieval mode, rounds used, LLM calls, and input/output/total tokens (reported by the API when the provider returns a usage field).

The API key is read from .env automatically - see the llm: section under Configuration.

Vector encoding

Vectors dominate the index. Three encodings are available, switchable in place with rag compact --encoding - no re-embedding required:

encoding index size recall@10 MRR top-20 order
float32 (default) 193 MB 6/14 0.329 baseline
float16 132 MB (-32%) 6/14 0.329 identical
int8 108 MB (-44%) 6/14 0.329 identical

Measured on a 20,340-vector index of ~7.5 MB of prose. Both lossy encodings produced byte-identical rankings; int8 score deviations were around 0.0007. Query latency was unchanged at 0.13s.

rag compact -d ./books --encoding int8

One caveat: conversion is one-way. Going from int8 back to float32 restores the format but not the discarded precision - you would need to re-index. Measure with rageval on your own corpus before converting, since these numbers come from one prose corpus and quantization error is corpus-dependent.

Set embedding.vector_encoding to build new indexes directly in a given encoding.

rageval - measuring retrieval quality

CLAUDE.md requires measuring changes to the embedding model, chunk size or embedded-text format against a ground-truth set rather than guessing. cmd/rageval is that harness.

go run ./cmd/rageval -corpus /path/to/indexed -k 10
14 questions, recall, MMR disabled

MODE             RECALL      MRR   RANKS
lexical           6/14     0.201   [- 4 - - - - - - 5 9 1 - 4 1]
semantic          6/14     0.329   [- 1 - - - - - - 1 - 1 2 9 1]
hybrid            6/14     0.310   [- 1 - - - - - - 2 - 1 3 2 1]

Each question is paired with an anchor phrase that must appear in a correctly retrieved passage. A - means the passage was not found within -k. Add -modes hyde to include HyDE, -v to print per-question ranks.

The bundled set in cmd/rageval/questions.json covers A Song of Ice and Fire; replace it with questions and anchors for your own corpus.

rag compact

Shrink the index on disk. Rewrites legacy JSON vector records as packed binary and reclaims free pages. No re-embedding, no change to results.

rag compact -d /path/to/content
# Rewrote 20340 vector records as packed binary
# Index: 579.3 MB -> 192.7 MB  (66.7% smaller)

Indexes created after this change already use the binary format; run it once on older indexes, and any time deletions have left free pages behind.

rag pack -q "<question>"

Pack relevant chunks into compressed context that fits a token budget.

rag pack -q "authentication flow" -b 2000
rag pack -q "API endpoints" -o context.json
rag pack -q "session handling" --lexical

Flags:

  • -q, --query - Search query (required)
  • -b, --budget - Token budget (default from config)
  • -o, --output - Output file (default: stdout)
  • -k, --top-k - Candidate pool size

rag runprompt

Generate formatted prompts from templates for manual LLM orchestration.

# Runtime prompt for question answering
rag runprompt --runtime --ctx context.json -q "How does auth work?"

# Builder prompt for context compression
rag runprompt --builder --ctx context.json

Flags:

  • --runtime - Use runtime (answering) prompt template
  • --builder - Use builder (compression) prompt template
  • --ctx - Path to packed context JSON file (required)
  • -q, --query - Override query for runtime prompt

Configuration

Create a rag.yaml file in your project root:

index:
  includes:
    - "**/*.go"
    - "**/*.py"
    - "**/*.js"
    - "**/*.ts"
    - "**/*.md"
  excludes:
    - "**/node_modules/**"
    - "**/vendor/**"
    - "**/.git/**"
  stemming: true
  chunk_tokens: 512
  chunk_overlap: 50
  k1: 1.2
  b: 0.75

retrieve:
  top_k: 20
  mmr_lambda: 0.7
  dedup_jaccard: 0.8

pack:
  token_budget: 4000
  output: json

logging:
  level: info

Configuration Options

Section Option Description Default
index includes Glob patterns for files to index Common code extensions
index excludes Glob patterns to exclude node_modules, vendor, .git
index stemming Enable Porter stemming true
index chunk_tokens Max tokens per chunk 512
index chunk_overlap Token overlap between chunks 50
index k1 BM25 k1 parameter 1.2
index b BM25 b parameter 0.75
retrieve top_k Default number of results 20
retrieve mmr_lambda MMR relevance vs diversity (0-1) 0.7
retrieve dedup_jaccard Jaccard threshold for dedup 0.8
retrieve hybrid_enabled Run the vector arm alongside BM25 false
retrieve rrf_k RRF fusion constant 60
retrieve bm25_weight BM25 vs vector balance (0-1) 0.5
embedding enabled Generate and use embeddings false
embedding provider ollama, openai, jina, deepseek, mock openai
embedding model Embedding model name text-embedding-3-small
embedding dimension Vector size; probed from the provider when possible model default
embedding include_path Prefix embedded chunks with path:lines true
embedding vector_encoding float32, float16 or int8 (see Vector encoding) float32
llm provider openai or deepseek; picks the default endpoint openai
llm base_url Override the endpoint (any OpenAI-compatible server) provider default
llm model Chat model name deepseek-web
llm api_key_env Env var holding the key; also read from .env. Empty sends no Authorization header, and then base_url is required ""
pack token_budget Default token budget 4000

Hybrid Search (BM25 + Vector Embeddings)

To enable semantic search alongside BM25 keyword search, install Ollama and pull an embedding model:

# Install Ollama (macOS)
brew install ollama

# Start Ollama server
ollama serve

# Pull an embedding model
ollama pull mxbai-embed-large

Embeddings may run locally; text generation may not - see the llm: section.

Then add embedding config to your rag.yaml:

embedding:
  enabled: true
  provider: ollama
  model: mxbai-embed-large
  dimension: 1024
  include_path: false   # true for code, false for prose

retrieve:
  hybrid_enabled: true
  rrf_k: 60             # RRF fusion parameter
  bm25_weight: 0.35     # Balance between BM25 and vector (0-1)

Re-index to generate embeddings:

rag index /path/to/content

Hybrid search runs both arms independently and fuses their rankings with Reciprocal Rank Fusion:

score(c) = bm25_weight / (rrf_k + rank_bm25) + (1 - bm25_weight) / (rrf_k + rank_vector)

Because the arms run independently, a chunk that only the vector arm finds still reaches the results — BM25 does not gate the candidate pool. Use --explain to see how many candidates each arm produced:

rag query -q "how are sessions validated" --explain
# retrieval: hybrid (bm25 + vector, RRF) (model=mxbai-embed-large, vectors=20340)
# candidates: bm25=32 vector=80 fused=97

If embeddings are unavailable (provider down, index not embedded, model changed), query and pack print a warning and fall back to BM25 rather than failing silently.

rag pack uses the same retrieval path as rag query, so hybrid search applies there too.

Tuning vector quality

Retrieval quality depends far more on the embedding model and chunk size than on the fusion parameters. Measured on a ~7.5MB prose corpus (14 questions, recall@10):

Setting Vector recall@10
nomic-embed-text, ~2200-char chunks 2/14
mxbai-embed-large, ~540-char chunks 6/14

Raising k shows the answers are being found but ranked low - vector recall@200 is 10/14. If you need them in the top 10, add a reranking stage over a deep candidate pool; tuning bm25_weight does not help (recall@10 was flat at 6/14 across 0.2-0.65).

embedding.include_path prepends path:startLine-endLine to each embedded chunk. This helps for code, where the path carries real signal, and hurts for prose - on the corpus above it cost 0.12 MRR. It defaults to true; set it to false for prose.

Changing embedding.model, embedding.dimension, or embedding.include_path invalidates stored vectors. The index records what it was embedded with and re-embeds automatically.

HyDE query expansion

--hyde makes exactly one LLM call per query to write a hypothetical answer passage, appends it to the query, and searches with both. The generated text is cached in the index, so repeating a query costs zero API calls. If the LLM fails, the search silently falls back to the plain query.

Generation goes through an OpenAI-compatible chat endpoint - the adapter deliberately has no local-model provider:

llm:
  provider: openai
  model: deepseek-web
  base_url: http://127.0.0.1:8787/v1
  api_key_env: ""
  max_tokens: 400

Note: RRF scores are much smaller than BM25 scores (typically 0.005-0.03). If you set retrieve.min_score_threshold, tune it for whichever mode you actually run — a threshold picked for BM25 scores will filter out every hybrid result.

Embeddings are incremental. Re-running rag index only embeds chunks that do not already have a vector, and drops vectors for chunks that no longer exist. Use rag index --force-embed to re-embed everything. Changing embedding.model triggers a rebuild, discarding vectors from the old model.

Semantic-Only Search

Use --semantic flag to search using only vector embeddings (no BM25 keyword matching):

rag query -q "a noble man betrayed by those he trusted" --semantic

Semantic search is useful for:

  • Natural language questions (e.g., "how to handle errors gracefully")
  • Conceptual queries where exact keywords may not appear
  • Finding related content even when terminology differs

Requires embeddings to be enabled and indexed (see Hybrid Search section above).

How It Works

Indexing

  1. Walks directory with glob patterns
  2. Checks file modification times for incremental updates
  3. Splits files into line-based chunks with token awareness
  4. Tokenizes with optional Porter stemming
  5. Builds inverted index with term frequencies
  6. Stores in BoltDB (.rag/index.db)

Retrieval

  1. Tokenizes and stems query
  2. Scores chunks using BM25:
    score(q,c) = Σ IDF(t) × (tf × (k1+1)) / (tf + k1 × (1-b + b×|c|/avgDl))
    
  3. Applies MMR for diversity:
    MMR(c) = λ × relevance(c) - (1-λ) × max_similarity(c, selected)
    
  4. Returns ranked, deduplicated results

Packing

  1. Calculates utility = score / token_count
  2. Greedily selects chunks by utility until budget exhausted
  3. Merges adjacent chunks from same file
  4. Outputs JSON with citations (path, line range, relevance)

Output Format

Packed Context JSON

{
  "query": "authentication",
  "budget_tokens": 4000,
  "used_tokens": 1250,
  "snippets": [
    {
      "path": "/src/auth/handler.go",
      "range": "L45-89",
      "why": "BM25 score: 2.34",
      "text": "func Authenticate(..."
    }
  ]
}

WebAssembly (Browser)

RAG can run entirely in the browser via WebAssembly (BM25 search only, no embeddings).

Build WASM

make build-wasm
# Or manually:
GOOS=js GOARCH=wasm go build -o examples/wasm/rag.wasm ./cmd/wasm

Run Demo

cd examples/wasm
python3 -m http.server 8080
# Open http://localhost:8080

JavaScript API

// Index content
ragIndex("file.txt", "Your text content here...")

// Search (returns JSON string)
const results = JSON.parse(ragQuery("search term", 5))

// Clear index
ragClear()

// Get statistics
const stats = JSON.parse(ragStats())

See examples/wasm/README.md for details.


Architecture

cmd/rag/main.go          # Entrypoint
cmd/wasm/main.go         # WASM entrypoint
internal/
├── domain/              # Core entities (Document, Chunk, etc.)
├── port/                # Interfaces (IndexStore, Retriever, etc.)
├── usecase/             # Business logic             # Business logic
│   ├── index.go         # Indexing orchestration
│   ├── embed.go         # Incremental embedding sync
│   ├── retrieve.go      # Search with BM25 + MMR
│   ├── pack.go          # Context packing
│   └── ask.go           # Retrieve, judge, iterate, answer
└── adapter/
    ├── fs/              # File system walker
    ├── store/           # BoltDB: index, vectors, HyDE cache
    ├── analyzer/        # Tokenizer + Porter stemmer
    ├── chunker/         # Line-based and AST chunking
    ├── embedding/       # Embedding providers (ollama, openai, jina, mock)
    ├── llm/             # Hosted chat providers (deepseek, openai)
    └── retriever/       # BM25, MMR, hybrid RRF, semantic, HyDE

License

MIT - see LICENSE.

Releases

Packages

Contributors

Languages