Turn a codebase into a map you can read, instead of a chat log you have to scroll.
Upload a ZIP, PDF, Markdown file or raw text. Ask a question in plain English. Demistify retrieves the relevant chunks, clusters them into concepts, lays them out in 2D, and draws two independent kinds of relationship over the same nodes — what is similar, and what is probably read first. Click any concept for a generated explanation and the exact source chunks it came from.
Ask a "chat with your repo" tool how a system works and you get a paragraph. It is often a good paragraph — and it is still a linear answer to a question about a structure. You cannot see how many distinct ideas the answer touched, which of them are related, or where to start reading.
Demistify answers the same question with a map: the concepts involved, how they relate, and a suggested order to read them in.
- You upload a codebase or document set and ask a question.
- It retrieves the semantically relevant chunks — not whole files.
- It clusters those chunks into concepts and labels each one.
- It lays the concepts out in 2D and connects them two different ways.
- You click a concept to get an explanation plus the raw source chunks behind it.
- Two edge types over one node set. Semantic similarity and source order, toggled live. This is the core idea — see The two edge types.
- Query-scoped graphs. The map is computed for your question, not a static whole-repo diagram nobody reads.
- Concept labels derived by class-based TF-IDF, so a label says what makes a cluster different from the others on screen.
- Source transparency. Every concept opens to the exact chunks it was built from.
- Explanations cached by content hash, so a repeated view is a local dictionary lookup rather than a network call.
- A suggested reading path, from a topological sort over the source-order edges.
- Local embeddings. The model runs on your machine via ONNX Runtime.
- Multi-format ingestion — ZIP archives, PDF, Markdown, plain text.
flowchart TD
A[Upload: ZIP / PDF / MD / text] --> B[Parse and chunk]
B --> C[Token-budget enforcement<br/>hard 512-token ceiling]
C --> D[BGE-small embeddings<br/>local ONNX Runtime]
D --> E[Normalized 384-dim vectors]
E --> F[(FAISS IndexFlatIP<br/>cosine similarity)]
Q[Natural-language query] --> R[Intent classification]
R --> S[FAISS retrieval, top_k=15]
F --> S
S --> T[Intent-weighted re-rank]
T --> U[Cluster into concepts<br/>HDBSCAN, agglomerative fallback]
U --> V[Label via class-based TF-IDF]
V --> W[PCA to 2D layout]
W --> X[Semantic edges + source-order edges]
X --> Y[Interactive concept map]
Y --> Z[Concept and source inspection]
Z --> AA[Explanation via OpenRouter]
AA --> AB[(Content-hash explanation cache)]
AB -.cache hit.-> Z
Services. A Next.js frontend on localhost:3000 (Clerk-gated routes) and a FastAPI
backend on localhost:8000 that owns all indexing and retrieval. They share nothing but
HTTP and one directory on disk.
main.py app construction, CORS allowlist, startup warm-up
routers/upload.py /upload, /upload-text, /refresh-index
routers/query.py /graph-data, /concept-summary, /file-summary, /keep-alive
core/
models.py the ONNX embedder singleton
chunker.py PDF / ZIP / Markdown / text -> chunks; archive safety
token_budget.py second-stage split enforcing the 512-token window
embedding.py embed -> FAISS; atomic index/metadata writes
retriever.py index + query caches, FAISS search, intent re-rank
clustering.py HDBSCAN (agglomerative fallback) + c-TF-IDF labels
graph_builder.py PCA layout, semantic edges, source-order edges
classifier.py intent: trained model if present, keyword fallback otherwise
explainer.py OpenRouter call + content-hashed, disk-persisted cache
supabase.py optional remote index sync (inert while unconfigured)
train_classifier.py offline: data/dataset.jsonl -> intent_classifier.pkl
frontend_app/src/
app/dashboard/page.tsx main app: upload, query, visualize
components/ FileUpload · QueryInterface · ConceptMap · ConceptDetails · Sidebar
Full detail: docs/ARCHITECTURE.md.
| Layer | Choice |
|---|---|
| Backend | Python 3.11+, FastAPI, Uvicorn |
| Embeddings | BAAI/bge-small-en-v1.5, int8 ONNX via ONNX Runtime |
| Vector search | FAISS IndexFlatIP over L2-normalized vectors |
| Clustering | HDBSCAN, agglomerative fallback |
| Labeling | scikit-learn, class-based TF-IDF |
| Layout | PCA (deterministic) |
| Explanations | OpenRouter |
| Frontend | Next.js 16, React 19, Tailwind 4, TypeScript |
| Auth | Clerk |
| Tests | standard-library unittest, no test dependency |
Ingestion. An upload is size-capped as it streams and parsed off the event loop in a
thread pool — chunking a ZIP is tens of seconds of pure CPU. ZIP handling never calls
extractall: members are read in memory, and absolute paths, .. traversal and symlinks
are rejected, with ceilings on total uncompressed bytes and member count.
Chunking and the token budget. Every source produces the same chunk shape, then a second stage hard-enforces the model's 512-token window. It re-measures after overlap is applied, because tokenization is not additive across a join. Chunk positions are renumbered per source file so a split cannot create duplicate positions.
Embedding. BGE-small runs locally through ONNX Runtime, producing normalized
384-dimensional vectors. Because vectors are unit-length at storage time, FAISS
IndexFlatIP is cosine similarity. Queries get the BGE search prefix; passages do not.
Retrieval and intent. A query is classified as code_to_doc or doc_to_code, which
re-weights results by document type. FAISS returns extra candidates, low-similarity hits
are dropped, and the survivors are re-ranked. Nothing on the request path re-embeds text
already in the index — vectors come back via index.reconstruct(i).
On intent classification, precisely: the classifier uses a trained logistic-regression model when
intent_classifier.pklexists, and falls back to keyword matching when it does not. That artifact is gitignored, so a fresh checkout runs the keyword fallback unless you train one withpython core/train_classifier.pyor supply your own.
Clustering and labels. Retrieved chunks are clustered by HDBSCAN, with agglomerative clustering as a fallback. Labels come from class-based TF-IDF computed across the concepts in one graph, so a label distinguishes a cluster from its neighbours rather than describing it in isolation.
Layout. PCA to 2D. Deterministic, so the same query lays out the same way twice — which matters more in a walkthrough than manifold fidelity on twenty points.
| Semantic edges | Source-order edges | |
|---|---|---|
| Meaning | "these concepts are about the same thing" | "this is probably read before that" |
| Computed from | cosine between cluster centroids | earliest chunk position in a shared source file |
| Weight scale | cosine, [0, 1] | position delta, unbounded ≥ 1 |
Semantic edges are thresholded relative to the graph, not absolutely. Every concept in a graph comes from one corpus, so centroids share vocabulary and all pairwise similarities land in a narrow high band. A fixed low cutoff excludes nothing and produces a complete graph — every node wired to every other, which tells you nothing. The threshold is instead this graph's own mean pairwise similarity, plus each concept's strongest edge so nothing is orphaned.
What source order means — and does not. It is the relative position of chunks within shared source files. It is not chronological authorship order, and it is not cross-file execution order. Two concepts are linked only when they draw on the same source file, and the direction reflects which appeared earlier in that file.
Explanations. Once a graph renders, the frontend begins prefetching concept explanations sequentially in the background, so a request may be in flight before you click anything. Clicking a concept uses the prefetched or cached explanation when one is available, and otherwise requests it. On a cache miss the backend sends up to the first 10 of that concept's selected chunk contents to OpenRouter as prompt context, then caches the returned explanation.
The cache is keyed by a hash of the chunk content — never by concept id, because ids are positional within one query's clustering and restart at zero each query, so id-keying would confidently serve one concept's explanation for an unrelated one.
Requires Python 3.11 or newer. requirements.txt pins scipy==1.17.0, which requires
≥ 3.11, so installation fails on 3.10 with a message about scipy rather than about your
Python version. Name the interpreter explicitly.
Backend — from the repository root:
python3.11 -m venv venv && source venv/bin/activate
pip install -r requirements.txt
uvicorn main:app --reload --port 8000Wait for Demistify backend ready. — that is the embedder loading and the example
queries being pre-embedded. Querying earlier still works; you just pay the warm-up on
your first query.
Frontend — from frontend_app/:
npm ci # exact locked versions; use `npm install` only to change a dependency
npm run devThen open http://localhost:3000, sign in, and upload something.
On macOS, start-backend.command and start-frontend.command launch each service after
verifying the environment. They deliberately never install anything — an install seconds
before a demo needs the network, takes minutes, and can change a working environment.
Both contracts are tracked as templates that carry no live credentials:
cp .env.example .env
cp frontend_app/.env.local.example frontend_app/.env.localBackend (.env) — all optional:
| Variable | Effect if unset |
|---|---|
OPENROUTER_API_KEY |
explanations cannot be generated; cached ones still serve |
SUPABASE_URL, SUPABASE_SERVICE_ROLE_KEY |
remote index sync stays dormant; local disk only |
ALLOWED_ORIGINS |
defaults to http://localhost:3000 |
Tuning and limit overrides (ONNX_THREADS, MAX_UPLOAD_MB, MAX_CHUNKS_PER_UPLOAD,
MAX_ZIP_UNCOMPRESSED_MB, MAX_ZIP_MEMBERS, MAX_CHUNK_TOKENS, MAX_BATCH_TOKENS) all
have working defaults and are documented where they are defined.
Frontend (frontend_app/.env.local) — NEXT_PUBLIC_BACKEND_URL plus two Clerk keys,
which are genuinely required: the routes are Clerk-gated, and the signed-in user id names
the index directory.
Demistify is designed to run locally on localhost. That is not the same as running fully offline, and the difference is worth stating plainly:
- Clerk is a hosted service and is required to sign in.
- OpenRouter is contacted to generate an explanation that is not already cached.
- HuggingFace Hub is a bootstrap dependency — contacted only when the model or tokenizer assets are missing from the local cache, typically on first run.
Embedding, indexing, retrieval, clustering and layout all run on your machine, and cached explanations are served from local disk.
The repository ships a curated public demo corpus at demo/, derived
from Demistify's own application source. Upload
demo/Demistify Demo Codebase.zip once, then try:
- How does the app explain a concept?
- How is the FAISS index created and searched?
- What happens when a user uploads a file?
Indexing is append-oriented, so uploading the same corpus twice duplicates its chunks.
See demo/README.md.
python -m unittest discover -s tests -vStandard-library unittest — no test dependency, nothing to install beyond the app's own
requirements. Coverage is over core/ and the startup scripts: chunking and ZIP safety
(including zip-slip and archive-size guards), embedder configuration, explainer caching
and model fallback, FAISS index/metadata integrity, retriever behaviour, telemetry, and
the token-budget splitter.
What the suite does and does not reach. Application service calls are mocked where
applicable — no Clerk account is touched, no OpenRouter request is made, and tests that
write to disk use their own temp directories, never your real vectorstore/. Most tests
are entirely local.
The exception is deliberate: the token-budget tests exercise the real BGE tokenizer, because that is the only way to prove no emitted chunk exceeds the model's 512-token window under real tokenization rather than a stub. The tokenizer is resolved cache-first, so on a machine with no cached model assets the first run may download them from HuggingFace. Once cached, those tests run locally without that bootstrap. So the suite is not guaranteed to be fully offline on a cold machine.
Optional:
python core/train_classifier.py # train the intent classifier; keyword fallback if absentThis is a bounded portfolio and demonstration prototype, not a production-scale repository indexer. It is honest about where its edges are:
| Guardrail | Default |
|---|---|
| Upload size ceiling | 25 MB per upload |
| Chunks per upload | 1,500 (sampled round-robin across files, not truncated) |
| ZIP uncompressed budget | 100 MB |
| ZIP member count | 5,000 |
| Chunk token ceiling | 512 (the model's window) |
| Retrieval breadth | top 15 chunks per query |
25 MB is the normal ceiling, not a guarantee that every file under it will succeed — a dense archive can still hit the chunk, member or uncompressed-byte limits first, and indexing time grows with content. Treat multi-gigabyte repositories as out of scope.
Full list: docs/LIMITATIONS.md.
- Stored locally by default. Uploaded content, the FAISS index and the explanation
cache are written to local disk under
vectorstore/, which is gitignored and never committed. Indexing, retrieval, clustering and layout all run on your machine. - Explanations reach OpenRouter. Generating an uncached explanation sends up to 10 of that concept's selected source-chunk contents to OpenRouter as prompt context — the actual code or document text of those chunks, not a summary of it. Explanation prefetching may begin automatically once a graph renders, so this can happen before you click a concept. A cache hit makes no OpenRouter request at all.
- Optional Supabase sync is off unless you turn it on. Under the documented default
configuration it is dormant:
.env.exampleshipsSUPABASE_URLandSUPABASE_SERVICE_ROLE_KEYblank, and sync is skipped entirely unless both are set. If you do configure them, Demistify also uploadsindex.faissandmeta.pklto Supabase Storage after each indexing run.meta.pklholds the chunk metadata including each chunk'scontent, so enabling Supabase means your indexed source text is no longer strictly local. The original uploaded file is not itself sent — the index pair is. - The backend trusts a client-supplied user id and does not verify the Clerk session token server-side. That is an accepted limitation for local single-user use and is the first thing that would need to change before any shared deployment.
- No rate limiting.
Run it on your own machine, with content you are comfortable indexing.
Feature-complete for its intended scope and actively maintained as a portfolio project. It runs locally, end to end, with a prepared demo corpus. It is not deployed, not multi-tenant, and not hardened for public hosting — by design.
Built by Shrey Bansal — sole creator, designer and developer.
Third-party libraries, models and datasets remain the property of their respective authors and retain their own licenses.
© 2026 Shrey Bansal. All rights reserved.
This project is published for portfolio and evaluation purposes. No open-source license is granted at this time.