Skip to content

Repository files navigation

Demistify

Turn a codebase into a map you can read, instead of a chat log you have to scroll.

Upload a ZIP, PDF, Markdown file or raw text. Ask a question in plain English. Demistify retrieves the relevant chunks, clusters them into concepts, lays them out in 2D, and draws two independent kinds of relationship over the same nodes — what is similar, and what is probably read first. Click any concept for a generated explanation and the exact source chunks it came from.


The problem

Ask a "chat with your repo" tool how a system works and you get a paragraph. It is often a good paragraph — and it is still a linear answer to a question about a structure. You cannot see how many distinct ideas the answer touched, which of them are related, or where to start reading.

Demistify answers the same question with a map: the concepts involved, how they relate, and a suggested order to read them in.

What it does

  1. You upload a codebase or document set and ask a question.
  2. It retrieves the semantically relevant chunks — not whole files.
  3. It clusters those chunks into concepts and labels each one.
  4. It lays the concepts out in 2D and connects them two different ways.
  5. You click a concept to get an explanation plus the raw source chunks behind it.

Features

  • Two edge types over one node set. Semantic similarity and source order, toggled live. This is the core idea — see The two edge types.
  • Query-scoped graphs. The map is computed for your question, not a static whole-repo diagram nobody reads.
  • Concept labels derived by class-based TF-IDF, so a label says what makes a cluster different from the others on screen.
  • Source transparency. Every concept opens to the exact chunks it was built from.
  • Explanations cached by content hash, so a repeated view is a local dictionary lookup rather than a network call.
  • A suggested reading path, from a topological sort over the source-order edges.
  • Local embeddings. The model runs on your machine via ONNX Runtime.
  • Multi-format ingestion — ZIP archives, PDF, Markdown, plain text.

Architecture

flowchart TD
    A[Upload: ZIP / PDF / MD / text] --> B[Parse and chunk]
    B --> C[Token-budget enforcement<br/>hard 512-token ceiling]
    C --> D[BGE-small embeddings<br/>local ONNX Runtime]
    D --> E[Normalized 384-dim vectors]
    E --> F[(FAISS IndexFlatIP<br/>cosine similarity)]

    Q[Natural-language query] --> R[Intent classification]
    R --> S[FAISS retrieval, top_k=15]
    F --> S
    S --> T[Intent-weighted re-rank]
    T --> U[Cluster into concepts<br/>HDBSCAN, agglomerative fallback]
    U --> V[Label via class-based TF-IDF]
    V --> W[PCA to 2D layout]
    W --> X[Semantic edges + source-order edges]
    X --> Y[Interactive concept map]
    Y --> Z[Concept and source inspection]
    Z --> AA[Explanation via OpenRouter]
    AA --> AB[(Content-hash explanation cache)]
    AB -.cache hit.-> Z
Loading

Services. A Next.js frontend on localhost:3000 (Clerk-gated routes) and a FastAPI backend on localhost:8000 that owns all indexing and retrieval. They share nothing but HTTP and one directory on disk.

main.py                 app construction, CORS allowlist, startup warm-up
routers/upload.py       /upload, /upload-text, /refresh-index
routers/query.py        /graph-data, /concept-summary, /file-summary, /keep-alive

core/
  models.py             the ONNX embedder singleton
  chunker.py            PDF / ZIP / Markdown / text -> chunks; archive safety
  token_budget.py       second-stage split enforcing the 512-token window
  embedding.py          embed -> FAISS; atomic index/metadata writes
  retriever.py          index + query caches, FAISS search, intent re-rank
  clustering.py         HDBSCAN (agglomerative fallback) + c-TF-IDF labels
  graph_builder.py      PCA layout, semantic edges, source-order edges
  classifier.py         intent: trained model if present, keyword fallback otherwise
  explainer.py          OpenRouter call + content-hashed, disk-persisted cache
  supabase.py           optional remote index sync (inert while unconfigured)
  train_classifier.py   offline: data/dataset.jsonl -> intent_classifier.pkl

frontend_app/src/
  app/dashboard/page.tsx   main app: upload, query, visualize
  components/              FileUpload · QueryInterface · ConceptMap · ConceptDetails · Sidebar

Full detail: docs/ARCHITECTURE.md.

Technical stack

Layer Choice
Backend Python 3.11+, FastAPI, Uvicorn
Embeddings BAAI/bge-small-en-v1.5, int8 ONNX via ONNX Runtime
Vector search FAISS IndexFlatIP over L2-normalized vectors
Clustering HDBSCAN, agglomerative fallback
Labeling scikit-learn, class-based TF-IDF
Layout PCA (deterministic)
Explanations OpenRouter
Frontend Next.js 16, React 19, Tailwind 4, TypeScript
Auth Clerk
Tests standard-library unittest, no test dependency

How the pipeline works

Ingestion. An upload is size-capped as it streams and parsed off the event loop in a thread pool — chunking a ZIP is tens of seconds of pure CPU. ZIP handling never calls extractall: members are read in memory, and absolute paths, .. traversal and symlinks are rejected, with ceilings on total uncompressed bytes and member count.

Chunking and the token budget. Every source produces the same chunk shape, then a second stage hard-enforces the model's 512-token window. It re-measures after overlap is applied, because tokenization is not additive across a join. Chunk positions are renumbered per source file so a split cannot create duplicate positions.

Embedding. BGE-small runs locally through ONNX Runtime, producing normalized 384-dimensional vectors. Because vectors are unit-length at storage time, FAISS IndexFlatIP is cosine similarity. Queries get the BGE search prefix; passages do not.

Retrieval and intent. A query is classified as code_to_doc or doc_to_code, which re-weights results by document type. FAISS returns extra candidates, low-similarity hits are dropped, and the survivors are re-ranked. Nothing on the request path re-embeds text already in the index — vectors come back via index.reconstruct(i).

On intent classification, precisely: the classifier uses a trained logistic-regression model when intent_classifier.pkl exists, and falls back to keyword matching when it does not. That artifact is gitignored, so a fresh checkout runs the keyword fallback unless you train one with python core/train_classifier.py or supply your own.

Clustering and labels. Retrieved chunks are clustered by HDBSCAN, with agglomerative clustering as a fallback. Labels come from class-based TF-IDF computed across the concepts in one graph, so a label distinguishes a cluster from its neighbours rather than describing it in isolation.

Layout. PCA to 2D. Deterministic, so the same query lays out the same way twice — which matters more in a walkthrough than manifold fidelity on twenty points.

The two edge types

Semantic edges Source-order edges
Meaning "these concepts are about the same thing" "this is probably read before that"
Computed from cosine between cluster centroids earliest chunk position in a shared source file
Weight scale cosine, [0, 1] position delta, unbounded ≥ 1

Semantic edges are thresholded relative to the graph, not absolutely. Every concept in a graph comes from one corpus, so centroids share vocabulary and all pairwise similarities land in a narrow high band. A fixed low cutoff excludes nothing and produces a complete graph — every node wired to every other, which tells you nothing. The threshold is instead this graph's own mean pairwise similarity, plus each concept's strongest edge so nothing is orphaned.

What source order means — and does not. It is the relative position of chunks within shared source files. It is not chronological authorship order, and it is not cross-file execution order. Two concepts are linked only when they draw on the same source file, and the direction reflects which appeared earlier in that file.

Explanations. Once a graph renders, the frontend begins prefetching concept explanations sequentially in the background, so a request may be in flight before you click anything. Clicking a concept uses the prefetched or cached explanation when one is available, and otherwise requests it. On a cache miss the backend sends up to the first 10 of that concept's selected chunk contents to OpenRouter as prompt context, then caches the returned explanation.

The cache is keyed by a hash of the chunk content — never by concept id, because ids are positional within one query's clustering and restart at zero each query, so id-keying would confidently serve one concept's explanation for an unrelated one.


Running it locally

Requires Python 3.11 or newer. requirements.txt pins scipy==1.17.0, which requires ≥ 3.11, so installation fails on 3.10 with a message about scipy rather than about your Python version. Name the interpreter explicitly.

Backend — from the repository root:

python3.11 -m venv venv && source venv/bin/activate
pip install -r requirements.txt
uvicorn main:app --reload --port 8000

Wait for Demistify backend ready. — that is the embedder loading and the example queries being pre-embedded. Querying earlier still works; you just pay the warm-up on your first query.

Frontend — from frontend_app/:

npm ci        # exact locked versions; use `npm install` only to change a dependency
npm run dev

Then open http://localhost:3000, sign in, and upload something.

On macOS, start-backend.command and start-frontend.command launch each service after verifying the environment. They deliberately never install anything — an install seconds before a demo needs the network, takes minutes, and can change a working environment.

Configuration

Both contracts are tracked as templates that carry no live credentials:

cp .env.example .env
cp frontend_app/.env.local.example frontend_app/.env.local

Backend (.env) — all optional:

Variable Effect if unset
OPENROUTER_API_KEY explanations cannot be generated; cached ones still serve
SUPABASE_URL, SUPABASE_SERVICE_ROLE_KEY remote index sync stays dormant; local disk only
ALLOWED_ORIGINS defaults to http://localhost:3000

Tuning and limit overrides (ONNX_THREADS, MAX_UPLOAD_MB, MAX_CHUNKS_PER_UPLOAD, MAX_ZIP_UNCOMPRESSED_MB, MAX_ZIP_MEMBERS, MAX_CHUNK_TOKENS, MAX_BATCH_TOKENS) all have working defaults and are documented where they are defined.

Frontend (frontend_app/.env.local) — NEXT_PUBLIC_BACKEND_URL plus two Clerk keys, which are genuinely required: the routes are Clerk-gated, and the signed-in user id names the index directory.

Local, but not guaranteed offline

Demistify is designed to run locally on localhost. That is not the same as running fully offline, and the difference is worth stating plainly:

  • Clerk is a hosted service and is required to sign in.
  • OpenRouter is contacted to generate an explanation that is not already cached.
  • HuggingFace Hub is a bootstrap dependency — contacted only when the model or tokenizer assets are missing from the local cache, typically on first run.

Embedding, indexing, retrieval, clustering and layout all run on your machine, and cached explanations are served from local disk.


Trying it with the bundled demo

The repository ships a curated public demo corpus at demo/, derived from Demistify's own application source. Upload demo/Demistify Demo Codebase.zip once, then try:

  • How does the app explain a concept?
  • How is the FAISS index created and searched?
  • What happens when a user uploads a file?

Indexing is append-oriented, so uploading the same corpus twice duplicates its chunks. See demo/README.md.

Tests

python -m unittest discover -s tests -v

Standard-library unittest — no test dependency, nothing to install beyond the app's own requirements. Coverage is over core/ and the startup scripts: chunking and ZIP safety (including zip-slip and archive-size guards), embedder configuration, explainer caching and model fallback, FAISS index/metadata integrity, retriever behaviour, telemetry, and the token-budget splitter.

What the suite does and does not reach. Application service calls are mocked where applicable — no Clerk account is touched, no OpenRouter request is made, and tests that write to disk use their own temp directories, never your real vectorstore/. Most tests are entirely local.

The exception is deliberate: the token-budget tests exercise the real BGE tokenizer, because that is the only way to prove no emitted chunk exceeds the model's 512-token window under real tokenization rather than a stub. The tokenizer is resolved cache-first, so on a machine with no cached model assets the first run may download them from HuggingFace. Once cached, those tests run locally without that bootstrap. So the suite is not guaranteed to be fully offline on a cold machine.

Optional:

python core/train_classifier.py    # train the intent classifier; keyword fallback if absent

Operating envelope

This is a bounded portfolio and demonstration prototype, not a production-scale repository indexer. It is honest about where its edges are:

Guardrail Default
Upload size ceiling 25 MB per upload
Chunks per upload 1,500 (sampled round-robin across files, not truncated)
ZIP uncompressed budget 100 MB
ZIP member count 5,000
Chunk token ceiling 512 (the model's window)
Retrieval breadth top 15 chunks per query

25 MB is the normal ceiling, not a guarantee that every file under it will succeed — a dense archive can still hit the chunk, member or uncompressed-byte limits first, and indexing time grows with content. Treat multi-gigabyte repositories as out of scope.

Full list: docs/LIMITATIONS.md.

Privacy and security scope

  • Stored locally by default. Uploaded content, the FAISS index and the explanation cache are written to local disk under vectorstore/, which is gitignored and never committed. Indexing, retrieval, clustering and layout all run on your machine.
  • Explanations reach OpenRouter. Generating an uncached explanation sends up to 10 of that concept's selected source-chunk contents to OpenRouter as prompt context — the actual code or document text of those chunks, not a summary of it. Explanation prefetching may begin automatically once a graph renders, so this can happen before you click a concept. A cache hit makes no OpenRouter request at all.
  • Optional Supabase sync is off unless you turn it on. Under the documented default configuration it is dormant: .env.example ships SUPABASE_URL and SUPABASE_SERVICE_ROLE_KEY blank, and sync is skipped entirely unless both are set. If you do configure them, Demistify also uploads index.faiss and meta.pkl to Supabase Storage after each indexing run. meta.pkl holds the chunk metadata including each chunk's content, so enabling Supabase means your indexed source text is no longer strictly local. The original uploaded file is not itself sent — the index pair is.
  • The backend trusts a client-supplied user id and does not verify the Clerk session token server-side. That is an accepted limitation for local single-user use and is the first thing that would need to change before any shared deployment.
  • No rate limiting.

Run it on your own machine, with content you are comfortable indexing.

Project status

Feature-complete for its intended scope and actively maintained as a portfolio project. It runs locally, end to end, with a prepared demo corpus. It is not deployed, not multi-tenant, and not hardened for public hosting — by design.


Author

Built by Shrey Bansal — sole creator, designer and developer.

Third-party libraries, models and datasets remain the property of their respective authors and retain their own licenses.

Copyright

© 2026 Shrey Bansal. All rights reserved.

This project is published for portfolio and evaluation purposes. No open-source license is granted at this time.

About

Local-first code intelligence that turns codebases and documentation into searchable, interactive concept maps.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages