Skip to content
uvesarshadPublic

About

Mine a codebase's git history for why the code is the way it is — not just what it is.

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

sprawl

Mine a codebase's git history for why the code is the way it is — not just what it is.

git blame tells you who. This tells you why.

sprawl reads a repo's own commit history and turns it into cited, structured answers instead of a wall of git log: which files carry the most bug-fix and defensive-change churn ("scars"), why a piece of code exists, what to watch out for before touching it, and whether that recorded intent has since drifted out of sync with the code it describes.

  • sprawl excavate — walks git history for every file-level node in the graph. The deterministic half (--no-llm, no network calls) computes a recency-weighted churn_score and detects "scar" commits (fixes, reverts, defensive changes). Add an LLM backend — or --github to also mine PRs, review threads, and linked issues — and it writes structured why / gotchas / scars / deprecated prose; every claim is checked against the actual commit messages, diffs, or PR text before being kept, so an unsupported claim is downgraded or dropped, not published.
  • sprawl drift — re-checks each node's recorded "why" against its current code and flags it (drifted / drift: possible) when they've diverged, so stale rationale doesn't quietly rot next to code that moved on.
  • sprawl biography — turns the same mined history into a long-form narrative: the oldest surviving code, the most-scarred files with the real incidents behind them (each cited to a commit or PR), the graveyard of abandoned approaches, and the eras the project has moved through. See Codebase biography.
sprawl excavate                # why-mining pass on cwd: git-history churn/scars/summary + LLM intent (why/scars/gotchas/deprecated), written into graph.json
sprawl excavate --github       # also mine the GitHub remote: commits → PRs → review threads → linked issues
sprawl drift                   # which excavated intent no longer matches its code (git-verified)
sprawl biography .             # long-form narrative account of the codebase, cited to commits/PRs

sprawl mines that history onto a full knowledge graph of the codebase underneath it — the same graph excavate/drift/biography write intent onto is the one you can query directly.

Type /sprawl in your AI coding assistant and it maps your entire project (code, docs, PDFs, images, videos) into a knowledge graph you can query instead of grepping through files.

  • Code maps for free, fully local. Code is parsed with tree-sitter AST: deterministic, no LLM, nothing leaves your machine. (Docs, PDFs, images and video use your assistant's model, or a configured API key, for a semantic pass.)
  • Every edge is explained. Each connection is tagged EXTRACTED (explicit in the source) or INFERRED (resolved by sprawl), so you can tell what was read directly from what was inferred.
  • Not a vector index. No embeddings, no vector store: a real graph you traverse. Ask a question, trace the path between two things, or explain one concept.

🤖 Using an AI coding assistant? Copy this into its chat to get started — installs the skill for every project on this machine. See For AI agents for a this-project-only variant and the full explanation.

Install and set up "sprawl" (a code-to-knowledge-graph tool) so it's available in
every project on this machine: run `uv tool install sprawl-graph` (or `pipx install
sprawl-graph`), then `sprawl install` to register its skill with yourself, then
`/sprawl .` to map this project into sprawl-out/. Read sprawl-out/GRAPH_REPORT.md and
summarize what you find. From now on, prefer `sprawl query "<question>"` or
`sprawl ask "<question>"` over grepping raw files when I ask about this codebase.

Credits. sprawl is built on graphify, an open-source codebase-to-knowledge-graph tool, which remains the underlying extraction and knowledge-graph engine. graphify is dual-licensed under Apache-2.0 and MIT; sprawl complies with those terms and preserves upstream attribution as documented in NOTICE, LICENSE, and LICENSE-MIT. sprawl is an independent fork and is not affiliated with or endorsed by the graphify project or Graphify Labs.


sprawl's interactive graph.html showing the FastAPI codebase as a force-directed knowledge graph with a legend of detected communities

The FastAPI codebase mapped by sprawl. Every node is a concept, colors are detected communities, and the whole thing is clickable in graph.html.

Get started (30 seconds):

uv tool install sprawl-graph      # install the CLI (or: pipx install sprawl-graph)
sprawl install               # register the skill with your AI assistant

Then, in your AI assistant:

/sprawl .

That's it. You get three files:

sprawl-out/
├── graph.html       open in any browser — click nodes, filter, search
├── GRAPH_REPORT.md  the highlights: key concepts, surprising connections, suggested questions
└── graph.json       the full graph — query it anytime without re-reading your files

Works in Claude Code, Cursor, Codex, Gemini CLI, GitHub Copilot, and 15+ more — pick your platform.


For AI agents

If a user has pointed you (an AI coding assistant) at this repository, or pasted the block from the top of this file, and asked you to set sprawl up or use it, do this:

  1. Install the CLI. uv tool install sprawl-graph (or pipx install sprawl-graph) — see Install for platform-specific notes and troubleshooting.
  2. Register the skill with yourself. sprawl install registers it in your user profile — available in every project on this machine. Add --project to scope the install to just the current repository instead (writes into the repo, e.g. .claude/skills/sprawl/SKILL.md, so it can be committed and shared with teammates). Either way this makes /sprawl available as a slash command in Claude Code, Cursor, Codex, and the 15+ other assistants sprawl supports.
  3. Build the graph. Run /sprawl . in the user's project (or sprawl extract <path> if you don't support slash commands). This writes sprawl-out/graph.html, sprawl-out/GRAPH_REPORT.md, and sprawl-out/graph.json, plus bookkeeping files sprawl uses internally (manifest.json, .sprawl_analysis.json, sprawl-out/cache/) — nothing outside sprawl-out/ is touched. See Team setup for what's safe to .gitignore.
  4. Prefer the graph over grepping. For codebase questions, use sprawl query "<question>" (raw subgraph) or sprawl ask "<question>" (a cited, ≤250-word answer) instead of reading files one by one — see Using the graph directly and ask vs query.
  5. Re-run after changes. /sprawl . --update re-extracts only changed files — code re-extraction needs no LLM call.

Copy-paste prompts a human can hand you, which work even if you cannot fetch this repository yourself. Both do the same thing — turn this project's code into a queryable knowledge graph so you can answer codebase questions by querying it instead of grepping or reading files blind — they differ only in whether the skill is registered for this repo or for every project on the machine.

This project only:

Install and set up "sprawl" (a code-to-knowledge-graph tool) for this project only: run
`uv tool install sprawl-graph` (or `pipx install sprawl-graph`), then
`sprawl install --project` to register its skill scoped to this repository, then
`/sprawl .` to map this project into sprawl-out/. Read sprawl-out/GRAPH_REPORT.md and
summarize what you find. From now on, prefer `sprawl query "<question>"` or
`sprawl ask "<question>"` over grepping raw files when I ask about this codebase.

All projects on this machine:

Install and set up "sprawl" (a code-to-knowledge-graph tool) so it's available in
every project on this machine: run `uv tool install sprawl-graph` (or `pipx install
sprawl-graph`), then `sprawl install` (no flags) to register its skill with yourself.
Then, in this project, run `/sprawl .` to map it into sprawl-out/. Read
sprawl-out/GRAPH_REPORT.md and summarize what you find. From now on, prefer
`sprawl query "<question>"` or `sprawl ask "<question>"` over grepping raw files when
I ask about this codebase.

This is an open-source tool you run locally, not a hosted service: installing it and reading sprawl-out/graph.json never sends the user's code anywhere unless they configure an API key for the optional semantic pass over docs/PDFs/images — see Privacy.

Contributing to sprawl's own source instead? This section is for agents using sprawl on someone else's project. If you're an agent working on sprawl's own codebase, read docs/overview.md instead — that's a separate audience with separate instructions.


Documentation

Full project documentation lives in /docs and is written for both humans and AI coding agents.

Start here: docs/overview.md

The overview contains the project's tech stack, architecture decisions, directory map of all doc files, and a glossary of domain terms. All other doc files are linked from there.

To UPDATE the docs after a change, read docs/maintenance.md.

AI agents (Claude Code, Gemini CLI, Codex, Cursor, etc.) should read docs/overview.md before making any changes to this codebase.


See it in action

sprawl path query: a terminal asks for the shortest path between FastAPI and ModelField, and the answer lights up hop by hop across the knowledge graph

Once the graph is built you query it instead of reading files. Real output, sprawl run on the FastAPI codebase shown above:

$ sprawl explain "APIRouter"
Node: APIRouter
  Source:    routing.py L2210
  Community: 2
  Degree:    47

Connections (47):
  --> RequestValidationError [uses] [INFERRED]
  --> Dependant [uses] [INFERRED]
  --> .get() [method] [EXTRACTED]
  <-- __init__.py [imports] [EXTRACTED]
  ...

$ sprawl path "FastAPI" "ModelField"
Shortest path (3 hops):
  FastAPI --uses--> DefaultPlaceholder <--references-- get_request_handler() --references--> ModelField

Every edge carries a confidence tag (EXTRACTED = explicit in the source, INFERRED = derived by resolution), so you can tell what was read directly from what was inferred. sprawl query "<question>" returns a scoped subgraph for a plain-language question, and sprawl path A B traces how any two things connect.


What it does

What you get out of the box:

Capability What you get
God nodes The most-connected concepts, so you see what everything flows through
Communities The graph split into subsystems (Leiden), with LLM-free labels
Cross-file links calls / imports / inherits / mixes_in resolved across ~40 languages via tree-sitter AST
Query, path, explain Ask a question, trace the path between two things, or explain one concept, all against graph.json
Rationale + doc refs # NOTE: / # WHY: comments and ADR/RFC citations become first-class nodes linked to the code
Beyond code Docs, PDFs, images, and video/audio all map into the same graph
Local-first Code is parsed locally with tree-sitter (no LLM, nothing leaves your machine); only the semantic pass over docs/media calls a backend, and only if you configure one

Benchmarks

Benchmark Metric sprawl Field
LOCOMO (n=300) recall@10 0.497 mem0 0.048, supermemory 0.149
LOCOMO (n=300) QA accuracy 45.3% supermemory 49.7%, mem0 27.3%
LongMemEval-S (n=50) QA accuracy 76% tied with dense RAG
Graph build LLM credits 0 per-token for most systems

Every system ran on the same harness with the same model and budgets, scored by a judge blind-validated against a second judge (90.6% agreement, Cohen's kappa 0.81). Full per-system tables, the code-intelligence result, and reproduction commands: BENCHMARKS.md.


Prerequisites

Requirement Minimum Check Install
Python 3.10+ python --version python.org
uv (recommended) any uv --version curl -LsSf https://astral.sh/uv/install.sh | sh
pipx (alternative) any pipx --version pip install pipx

macOS quick install (Homebrew):

brew install python@3.12 uv

Windows quick install:

winget install astral-sh.uv

Ubuntu/Debian:

sudo apt install python3.12 python3-pip pipx
# or install uv:
curl -LsSf https://astral.sh/uv/install.sh | sh

Install

Official package: The PyPI distribution is sprawl-graph (the name sprawl was already taken on PyPI by an unrelated package). The CLI command it installs is still sprawl. Other sprawl*/sprawl-* packages on PyPI are not affiliated.

Step 1 — install the package:

# Recommended (isolated env; if 'sprawl' isn't found after, run: uv tool update-shell):
uv tool install sprawl-graph

# Alternatives:
pipx install sprawl-graph
pip install sprawl-graph  # may need PATH setup — see note below

Step 2 — register the skill with your AI assistant:

sprawl install

That's it. Open your AI assistant and type /sprawl .

To install the assistant skill into the current repository instead of your user profile, add --project:

sprawl install --project
sprawl install --project --platform codex

Project-scoped installs write under the current directory, for example .claude/skills/sprawl/SKILL.md or .agents/skills/sprawl/SKILL.md (plus a references/ sidecar the skill loads on demand), and print a git add hint for files that can be committed. Per-platform commands that support project-scoped installs accept the same flag, for example sprawl claude install --project or sprawl codex install --project.

PowerShell note: Use sprawl . not /sprawl . — the leading slash is a path separator in PowerShell.

sprawl: command not found? uv tool install / pipx install put the sprawl command in their tool bin dir (~/.local/bin). If your shell can't find it right after install — common on a fresh macOS + zsh setup — that dir isn't on your PATH yet: run uv tool update-shell (or pipx ensurepath), then open a new terminal. With plain pip, add ~/.local/bin (Linux) or ~/Library/Python/3.x/bin (Mac) to your PATH, or run python -m sprawl.

Running with uvx / uv tool run instead of installing? Name the package, not the command: uvx --from sprawl-graph sprawl install. Plain uvx sprawl … fails (No solution found … no versions of sprawl) because uv tool run reads the first word as a package, and the package is sprawl-graph — the sprawl command lives inside it.

Avoid pip install on Mac/Windows if possible. The skill resolves Python at runtime from sprawl-out/.sprawl_python; if that points to a different environment than where pip installed the package, you'll get ModuleNotFoundError: No module named 'sprawl'. uv tool install and pipx install isolate the package in their own env and avoid this entirely.

Git hooks and uv tool / pipx: sprawl hook install embeds the current interpreter path directly into the hook scripts at install time, so the post-commit hook fires correctly even in GUI git clients and CI runners where ~/.local/bin is not on PATH. If you reinstall or upgrade sprawl, re-run sprawl hook install to refresh the embedded path.

Strict mode (Claude Code): sprawl install --project --strict makes the assistant actually use the graph. The default install nudges it to run sprawl query before reading files; strict mode blocks the first raw source read of a session and redirects it to the graph, then reverts to the nudge (so it fires at most once per session and never gets stuck). Toggle at runtime with SPRAWL_HOOK_STRICT=1/0; the default install is unchanged (soft nudge).

Pick your platform (20+ assistants, click to expand)
Platform Install command
Claude Code (Linux/Mac) sprawl install
Claude Code (Windows) sprawl install (auto-detected) or sprawl install --platform windows
CodeBuddy sprawl install --platform codebuddy
Codex sprawl install --platform codex
OpenCode sprawl install --platform opencode
Kilo Code sprawl install --platform kilo
GitHub Copilot CLI sprawl install --platform copilot
VS Code Copilot Chat sprawl vscode install
Aider sprawl install --platform aider
OpenClaw sprawl install --platform claw
Factory Droid sprawl install --platform droid
Trae sprawl install --platform trae
Trae CN sprawl install --platform trae-cn
Gemini CLI sprawl install --platform gemini
Hermes sprawl install --platform hermes
Kimi Code sprawl install --platform kimi
Amp sprawl amp install
Agent Skills (cross-framework) sprawl install --platform agents (alias --platform skills)
Kiro IDE/CLI sprawl kiro install
Pi coding agent sprawl install --platform pi
Cursor sprawl cursor install
Devin CLI sprawl devin install
Google Antigravity sprawl antigravity install

Codex users also need multi_agent = true under [features] in ~/.codex/config.toml for parallel extraction. CodeBuddy uses the same Agent tool and PreToolUse hook mechanism as Claude Code. Factory Droid uses the Task tool for parallel subagent dispatch. OpenClaw and Aider use sequential extraction (parallel agent support is still early on those platforms). Trae uses the Agent tool for parallel subagent dispatch and does not support PreToolUse hooks, so AGENTS.md is the always-on mechanism.

--platform agents (alias --platform skills) targets the generic cross-framework Agent-Skills locations: the spec's user-global ~/.agents/skills/ (read by npx skills and spec-compliant frameworks) for a global install, and ./.agents/skills/ for a project (--project) install. The bare sprawl install stays single-platform (Claude Code) by design — use the named agents platform when you want the skill discoverable by any framework that reads .agents/skills.

Codex uses $sprawl instead of /sprawl.

Optional extras (install only what you need)
Extra What it adds Install
pdf PDF extraction uv tool install "sprawl-graph[pdf]"
office .docx and .xlsx support uv tool install "sprawl-graph[office]"
google Google Sheets rendering uv tool install "sprawl-graph[google]"
video Video/audio transcription (faster-whisper + yt-dlp) uv tool install "sprawl-graph[video]"
mcp MCP stdio server uv tool install "sprawl-graph[mcp]"
neo4j Neo4j push support uv tool install "sprawl-graph[neo4j]"
falkordb FalkorDB push support uv tool install "sprawl-graph[falkordb]"
svg SVG graph export uv tool install "sprawl-graph[svg]"
leiden Leiden community detection (Python < 3.13 only) uv tool install "sprawl-graph[leiden]"
ollama Ollama local inference uv tool install "sprawl-graph[ollama]"
openai OpenAI / OpenAI-compatible APIs uv tool install "sprawl-graph[openai]"
gemini Google Gemini API uv tool install "sprawl-graph[gemini]"
anthropic Anthropic Claude API (--backend claude, uses ANTHROPIC_API_KEY) uv tool install "sprawl-graph[anthropic]"
bedrock AWS Bedrock (uses IAM, no API key) uv tool install "sprawl-graph[bedrock]"
azure Azure OpenAI Service (--backend azure, uses AZURE_OPENAI_API_KEY + AZURE_OPENAI_ENDPOINT) uv tool install "sprawl-graph[openai]"
sql SQL schema extraction uv tool install "sprawl-graph[sql]"
postgres Live PostgreSQL introspection (--postgres DSN) uv tool install "sprawl-graph[postgres]"
dm BYOND DreamMaker .dm/.dme AST extraction (may need a C compiler + python3-dev if no wheel matches your platform) uv tool install "sprawl-graph[dm]"
terraform Terraform / HCL .tf/.tfvars/.hcl AST extraction uv tool install "sprawl-graph[terraform]"
pascal Pascal / Delphi .pas/.dpr/.dpk/.inc AST extraction (more accurate calls/inherits edges; falls back to a regex extractor when absent) uv tool install "sprawl-graph[pascal]"
chinese Chinese query segmentation (jieba) uv tool install "sprawl-graph[chinese]"
all Everything above uv tool install "sprawl-graph[all]"

Make your assistant always use the graph

Run this once in your project after building a graph:

Platform Command
Claude Code sprawl claude install
CodeBuddy sprawl codebuddy install
Codex sprawl codex install
OpenCode sprawl opencode install
Kilo Code sprawl kilo install
GitHub Copilot CLI sprawl copilot install
VS Code Copilot Chat sprawl vscode install
Aider sprawl aider install
OpenClaw sprawl claw install
Factory Droid sprawl droid install
Trae sprawl trae install
Trae CN sprawl trae-cn install
Cursor sprawl cursor install
Gemini CLI sprawl gemini install
Hermes sprawl hermes install
Kimi Code sprawl install --platform kimi
Amp sprawl amp install
Agent Skills (cross-framework) sprawl agents install (alias sprawl skills install)
Kiro IDE/CLI sprawl kiro install
Pi coding agent sprawl pi install
Devin CLI sprawl devin install
Google Antigravity sprawl antigravity install

This writes a small config file that tells your assistant to consult the knowledge graph for codebase questions, preferring scoped queries like sprawl query "<question>" over reading the full report or grepping raw files.

  • Hook platforms (Claude Code, Gemini CLI): a hook fires automatically before search-style tool calls (and, on Claude Code, before reading source files one by one via the Read/Glob tools) and nudges your assistant toward the graph path.
  • Instruction-file platforms (Codex, OpenCode, Cursor, etc.): persistent instruction files (AGENTS.md, .cursor/rules/, etc.) provide the same query-first guidance.

GRAPH_REPORT.md is still available for broad architecture review.

CodeBuddy does the same two things as Claude Code: writes a CODEBUDDY.md section telling CodeBuddy to read sprawl-out/GRAPH_REPORT.md before answering architecture questions, and installs PreToolUse hooks (.codebuddy/settings.json) that fire before Bash search commands and file reads, nudging toward sprawl query instead.

Codex writes to AGENTS.md, which is what actually carries the always-on graph guidance on this platform. sprawl codex install also registers a PreToolUse hook in .codex/hooks.json (sprawl hook-check), but that entry is deliberately a no-op: Codex Desktop rejects hookSpecificOutput.additionalContext on PreToolUse, so emitting a nudge there would break Bash tool calls. Unlike Claude Code, where the hook (sprawl hook-guard) does the nudging, on Codex the hook fires and intentionally does nothing, and AGENTS.md is the always-on mechanism.

Kilo Code installs the Sprawl skill to ~/.config/kilo/skills/sprawl/SKILL.md and a native /sprawl command to ~/.config/kilo/command/sprawl.md. sprawl kilo install also writes AGENTS.md plus a native tool.execute.before plugin (.kilo/plugins/sprawl.js + .kilo/kilo.json or .kilo/kilo.jsonc registration) so Kilo gets the same always-on graph reminder behavior through native .kilo config.

Cursor writes .cursor/rules/sprawl.mdc with alwaysApply: true, so Cursor includes it in every conversation automatically, no hook needed.

To remove sprawl from all platforms at once: sprawl uninstall (add --purge to also delete sprawl-out/). Or use the per-platform command (e.g. sprawl claude uninstall).


What's in the report

  • God nodes — the most-connected concepts in your project. Everything flows through these.
  • Surprising connections — links between things that live in different files or modules. Ranked by how unexpected they are.
  • The "why" — inline comments (# NOTE:, # WHY:, # HACK:), docstrings, and design rationale from docs are extracted as separate nodes linked to the code they explain.
  • Suggested questions — 4–5 questions the graph is uniquely positioned to answer.
  • Confidence tags — every inferred relationship is marked EXTRACTED, INFERRED, or AMBIGUOUS. You always know what was found vs guessed.

What files it handles

Type Extensions
Code (36 tree-sitter grammars) .py .ts .mts .cts .js .jsx .tsx .mjs .go .rs .java .c .cpp .cc .cxx .h .hpp .cu .cuh .metal .rb .cs .kt .kts .scala .php .swift .lua .luau .toc .zig .ps1 .psm1 .psd1 .ex .exs .m .mm .jl .vue .svelte .astro .groovy .gradle .dart .v .sv .svh .sql .f .f90 .f95 .f03 .f08 .pas .pp .dpr .dpk .lpr .inc .dfm .lfm .lpk .sh .bash .json .dm .dme .dmi .dmm .dmf .sln .slnx .csproj .fsproj .vbproj .xaml .razor .cshtml (.dm/.dme requires uv tool install sprawl-graph[dm]; .mts/.cts reuse the TypeScript grammar, .cc/.cxx and CUDA .cu/.cuh and Metal .metal reuse the C++ grammar)
Salesforce Apex .cls .trigger (regex-based; classes, interfaces, enums, methods, triggers, SOQL/DML edges)
Terraform / HCL .tf .tfvars .hcl (requires uv tool install sprawl-graph[terraform])
MCP configs .mcp.json mcp.json mcp_servers.json claude_desktop_config.json — extracts server nodes, package refs, env var requirements
Package manifests apm.yml pyproject.toml go.mod pom.xml — one canonical package node per package (by name) plus depends_on edges, so a package referenced from many manifests is a single hub
Docs .md .mdx .qmd .html .txt .rst .yaml .yml (markdown [text](./other.md) links and [[wikilinks]] become references edges between docs)
Office .docx .xlsx (requires uv tool install sprawl-graph[office])
Google Workspace .gdoc .gsheet .gslides (opt-in; requires gws auth and --google-workspace; Sheets need uv tool install sprawl-graph[google])
PDFs .pdf
Images .png .jpg .webp .gif
Video / Audio .mp4 .mov .mp3 .wav and more (requires uv tool install sprawl-graph[video])
YouTube / URLs any video URL (requires uv tool install sprawl-graph[video])

Code is extracted locally with no API calls (AST via tree-sitter). Everything else goes through your AI assistant's model API.

Setup for Google Workspace ingestion, video/audio transcription, PR triage, and database/graph-export integrations is covered in More below.


Common commands

/sprawl .                        # build graph for current folder
/sprawl ./docs --update          # re-extract only changed files
/sprawl . --cluster-only         # rerun clustering without re-extracting
/sprawl . --cluster-only --resolution 1.5      # more granular communities
/sprawl . --cluster-only --exclude-hubs 99     # suppress utility super-hubs from god-node rankings
/sprawl . --no-viz               # skip the HTML, just the report + JSON
/sprawl . --wiki                 # build a markdown wiki from the graph
sprawl export callflow-html      # Mermaid architecture/call-flow HTML (auto-regenerates on every git commit if hook is installed)

/sprawl query "what connects auth to the database?"
/sprawl path "UserService" "DatabasePool"
/sprawl explain "RateLimiter"

/sprawl add https://arxiv.org/abs/1706.03762   # fetch a paper and add it

sprawl hook install              # auto-rebuild on git commit
sprawl merge-graphs a.json b.json              # combine two graphs

See the full command reference below, or More for PR triage, video/audio transcription, Google Workspace ingestion, and database/graph-export integrations.


Ignoring files

Create a .sprawlignore in your project root — same syntax as .gitignore, including ! negation.

.gitignore is respected automatically. sprawl reads the .gitignore in each directory. If a .sprawlignore is also present, the two are merged — .sprawlignore patterns are evaluated last, so they win on conflicts (including ! negations). Adding a .sprawlignore only ever excludes more; it never re-includes a file your .gitignore already excluded. Subdirectory scoping works the same way as git — an ignore file only affects its own subtree.

Pass --no-gitignore to sprawl extract when git-ignored generated or transpiled code belongs in the graph. This disables .gitignore and .git/info/exclude; .sprawlignore still applies.

# .sprawlignore
node_modules/
dist/
*.generated.py

# only index src/, ignore everything else
*
!src/
!src/**

Team setup

sprawl-out/ is meant to be committed to git so everyone on the team starts with a map.

Recommended .gitignore additions:

sprawl-out/cost.json        # local only
# sprawl-out/cache/         # optional: commit for speed, skip to keep repo small

manifest.json is now portable — keys are stored as relative paths and re-anchored on load, so committing it is safe and avoids a full rebuild on first checkout.

Workflow:

  1. One person runs /sprawl . and commits sprawl-out/.
  2. Everyone pulls — their assistant reads the graph immediately.
  3. Run sprawl hook install to auto-rebuild after each commit (AST only, no API cost). This also sets up a git merge driver so graph.json is never left with conflict markers — two devs committing in parallel get their graphs union-merged automatically.
  4. When docs or papers change, run /sprawl --update to refresh those nodes.

Using the graph directly

# ask a question and get an answer back, not a node dump
sprawl ask "how does workflow execution work"
sprawl ask "why does the layout algorithm use igraph?" --verbose   # + ranked node appendix
sprawl ask "how are errors handled" --no-llm                       # extractive, no model needed

# query the graph from the terminal (raw subgraph)
sprawl query "show the auth flow"
sprawl query "what connects DigestAuth to Response?" --graph sprawl-out/graph.json

# expose the graph as an MCP server (for repeated tool-call access)
python -m sprawl.serve sprawl-out/graph.json
python -m sprawl.serve --graph sprawl-out/graph.json  # --graph flag also accepted

# register with Kimi Code:
kimi mcp add --transport stdio sprawl -- python -m sprawl.serve sprawl-out/graph.json

# or serve over HTTP so a whole team points at one URL (no local sprawl needed):
python -m sprawl.serve sprawl-out/graph.json --transport http --port 8080
python -m sprawl.serve sprawl-out/graph.json --transport http --host 0.0.0.0 --api-key "$SECRET"

The MCP server gives your assistant structured access: ask, query_graph, get_node, get_neighbors, shortest_path, list_prs, get_pr_impact, triage_prs.

ask vs query

query walks the graph and prints the subgraph it reached — every node, truncated at a token budget. That is the right tool when you want the raw structure, and the wrong one when you asked a question: on a 3,464-file project "how does workflow execution work" returned 251 nodes, cut to 51 by the budget, with components/ui/button.tsx sitting among the real hits purely because it neighbours something whose label matched.

ask answers instead. It reuses the same retrieval, adds a lexical pass over the excavated intent fields (summary / why / gotchas / scars, written by sprawl excavate), then ranks with BM25 gated on topical similarity so a node that is merely adjacent to a hit cannot outrank one that is actually about the question — on that same query button.tsx lands at rank 2381 of 2462. If a hand-written doc covers the question it returns that excerpt first and the graph synthesis second. The answer is ≤250 words with file:line citations, assembled from ≤2000 tokens of context; the node list lives behind --verbose. With no LLM backend configured it degrades to an extractive answer built from the graph's own excavated prose rather than failing. sprawl benchmark --ask scores the two against each other.

Shared HTTP server

--transport stdio (the default) spawns one local server per developer. --transport http serves the same tools over the MCP Streamable HTTP transport, so a single shared process can serve the graph for the whole team — clients point their IDE MCP config at http://<host>:8080/mcp instead of running sprawl locally.

Flag Default Purpose
--transport {stdio,http} stdio Transport to serve on
--host 127.0.0.1 HTTP bind host (use 0.0.0.0 to expose beyond localhost)
--port 8080 HTTP bind port
--api-key env SPRAWL_API_KEY Require Authorization: Bearer <key> (or X-API-Key)
--path /mcp HTTP mount path
--json-response off Return plain JSON instead of SSE streams
--stateless off No per-session state (for load-balanced / CI deployments)
--session-timeout 3600 Reap idle stateful sessions after N seconds (0 disables)

The default 127.0.0.1 bind is loopback-only. Set --host 0.0.0.0 and --api-key together when exposing on a shared host. Run it in a container:

docker build -t sprawl .
docker run -p 8080:8080 -v "$(pwd)/sprawl-out:/data" sprawl \
  /data/graph.json --transport http --host 0.0.0.0 --api-key "$SECRET"

WSL / Linux note: Ubuntu ships python3, not python. Use a venv to avoid conflicts:

python3 -m venv .venv && .venv/bin/pip install "sprawl-graph[mcp]"

Codebase biography

sprawl biography writes a long-form, narrative account of a codebase from its own git history and code graph: the oldest surviving code, the most-scarred files and the real incidents behind them (each cited to a commit or PR), the graveyard of abandoned approaches, the single load-bearing hack, and the eras the project has moved through. Output is one self-contained HTML file (or Markdown) — open it in a browser or paste it into a blog post, no server and no build step.

Live demo: FastAPI's biography — generated by sprawl biography https://github.com/tiangolo/fastapi against this project's own history, no edits.

sprawl biography .                              # this project, self-contained HTML
sprawl biography . --md                         # Markdown instead
sprawl biography https://github.com/org/repo    # clone + build a biography for any public repo
sprawl biography . --name "My Project"          # display name (default: directory name)
sprawl biography . --no-github                  # skip GitHub PR/issue enrichment (faster)
sprawl biography . --limit 80                   # excavation budget in files (default 40)
sprawl biography . --render-only                # re-render from an existing graph.json, no re-excavation
sprawl biography . --theme light                # light | dark | auto (default: auto)

First run does extract → excavate → render in one shot. The narrated incidents need an LLM backend (an installed claude or codex CLI works with no API key — see Environment variables for the rest); --no-llm produces a deterministic-only page with churn/scar counts but no written incidents.


Environment variables

These are only needed for headless / CI extraction (sprawl extract). When running via the /sprawl skill inside your IDE, the model API is provided by your IDE session — no extra keys needed.

Variable Used for When required
ANTHROPIC_API_KEY Claude (Anthropic) backend --backend claude
ANTHROPIC_BASE_URL Anthropic-compatible endpoint URL (LiteLLM proxy, gateways, ...) --backend claude (default: https://api.anthropic.com)
ANTHROPIC_MODEL Model name for the Claude backend — for custom endpoints, use the model name/alias your server exposes --backend claude (default: claude-sonnet-4-6)
GEMINI_API_KEY or GOOGLE_API_KEY Google Gemini backend --backend gemini
OPENAI_API_KEY OpenAI or OpenAI-compatible APIs --backend openai (local servers accept any non-empty value)
OPENAI_BASE_URL OpenAI-compatible server URL (llama.cpp, vLLM, LM Studio, ...) --backend openai (default: https://api.openai.com/v1)
OPENAI_MODEL Model name for the OpenAI backend — for self-hosted servers, use the model name/alias your server exposes (check its /v1/models endpoint), e.g. LFM2.5-8B-A1B-UD-Q4_K_XL for llama.cpp --backend openai (default: gpt-4.1-mini)
DEEPSEEK_API_KEY DeepSeek backend --backend deepseek
MOONSHOT_API_KEY Kimi Code backend --backend kimi
OLLAMA_BASE_URL Ollama local inference URL --backend ollama (default: http://localhost:11434)
OLLAMA_MODEL Ollama model name --backend ollama (default: auto-detect)
SPRAWL_OLLAMA_NUM_CTX Override Ollama KV-cache window size optional — auto-sized by default
SPRAWL_OLLAMA_KEEP_ALIVE Minutes to keep Ollama model loaded optional — set 0 to unload after each chunk
AZURE_OPENAI_API_KEY Azure OpenAI Service backend --backend azure
AZURE_OPENAI_ENDPOINT Azure resource endpoint URL --backend azure (required alongside API key)
AZURE_OPENAI_API_VERSION Azure API version override optional — default 2024-12-01-preview
AZURE_OPENAI_DEPLOYMENT or SPRAWL_AZURE_MODEL Azure deployment name optional — default gpt-4o
AWS_* / ~/.aws/credentials AWS Bedrock — standard credential chain --backend bedrock (no API key, uses IAM)
SPRAWL_MAX_WORKERS AST parallelism thread count optional — also --max-workers flag
SPRAWL_MAX_OUTPUT_TOKENS Raise output cap for dense corpora optional — e.g. 32768 for large files
SPRAWL_API_TIMEOUT Per-call timeout in seconds for HTTP, claude-cli, Anthropic SDK, and Bedrock backends (default: 600) optional — also --api-timeout flag
SPRAWL_MAX_RETRIES How many times to retry a rate-limited (429) request before giving up (default: 6; honors Retry-After) optional — raise for strict per-org limits (e.g. kimi); 0 disables
SPRAWL_FORCE Force graph rebuild even with fewer nodes optional — also --force flag
SPRAWL_GOOGLE_WORKSPACE Auto-enable Google Workspace export optional — set to 1
SPRAWL_TRIAGE_BACKEND Backend for sprawl prs --triage optional — auto-detected from available keys
SPRAWL_TRIAGE_MODEL Model override for triage optional — e.g. claude-opus-4-7
SPRAWL_QUERY_LOG_ENABLE Set to 1 to turn on the local query log at ~/.cache/sprawl-queries.log (records each query/path/explain question + corpus path). Off by default — nothing is written unless you opt in (#1797) optional
SPRAWL_QUERY_LOG Enable the query log and write it to this path instead of the default optional — off unless this or _ENABLE is set
SPRAWL_QUERY_LOG_DISABLE Set to 1 to force the query log off (wins over the enable vars) optional
SPRAWL_QUERY_LOG_RESPONSES When the log is enabled, also record full subgraph responses (off by default) optional
SPRAWL_MAX_GRAPH_BYTES Override the 512 MiB graph.json size cap — e.g. 700MB, 2GB, or plain bytes optional — useful for very large corpora
SPRAWL_MAX_CONTEXTS Maximum number of non-default project graphs retained by one multi-project MCP server optional — default: 8; invalid values use 8, and values below 1 use 1
SPRAWL_LLM_TEMPERATURE Override LLM temperature for semantic extraction — e.g. 0.7, or none to omit optional — auto-omitted for o1/o3/o4/gpt-5 reasoning models

Privacy

  • Code files — processed locally via tree-sitter. Nothing leaves your machine. A code-only corpus requires no API key — sprawl extract runs fully offline. On a mixed repo, add --code-only to index just the code and skip the docs/PDFs/images that would otherwise need an LLM.
  • Video / audio — transcribed locally with faster-whisper. Nothing leaves your machine.
  • Docs, PDFs, images — sent to your AI assistant for semantic extraction (via the /sprawl skill, using whatever model your IDE session runs). Headless sprawl extract requires GEMINI_API_KEY / GOOGLE_API_KEY (Gemini), MOONSHOT_API_KEY (Kimi), ANTHROPIC_API_KEY (Claude), OPENAI_API_KEY (OpenAI), DEEPSEEK_API_KEY (DeepSeek), a running Ollama instance (OLLAMA_BASE_URL), AWS credentials via the standard provider chain (Bedrock - no API key needed, uses IAM), or the claude CLI binary (Claude Code - no API key needed, uses your Claude subscription). The --dedup-llm flag uses the same key.
  • Data residency — sprawl extract auto-detects which provider to use based on which API key is set (priority: Gemini → Kimi → Claude → OpenAI → DeepSeek → Azure → Bedrock → Ollama). For code with data-residency requirements, use --backend ollama (fully local) or pass an explicit --backend flag. Kimi (MOONSHOT_API_KEY) routes to Moonshot AI servers in China.
  • No telemetry, no usage tracking, no analytics.
  • Query logging — every sprawl query, sprawl path, sprawl explain, and MCP query_graph call is logged to ~/.cache/sprawl-queries.log in JSON Lines format (timestamp, question, corpus, nodes returned, duration). Full subgraph responses are not stored by default. Set SPRAWL_QUERY_LOG_DISABLE=1 to opt out, or SPRAWL_QUERY_LOG=/dev/null to silence without disabling the code path.

Troubleshooting

sprawl: command not found after installing The CLI is installed but its bin directory isn't on your shell's PATH. Pick the fix for how you installed:

  • uv (uv tool install sprawl-graph): the command lands in uv's tool bin dir (~/.local/bin), which a fresh macOS/zsh setup often doesn't have on PATH. Run uv tool update-shell, then open a new terminal. (Find the dir with uv tool dir --bin.)
  • pipx (pipx install sprawl-graph): run pipx ensurepath, then open a new terminal.
  • pip (pip install sprawl-graph): pip installs scripts to a user bin dir that may not be on PATH — add ~/Library/Python/3.x/bin (macOS) or ~/.local/bin (Linux) to your PATH in ~/.zshrc/~/.bashrc, or just run python -m sprawl.

uvx sprawl … or uv tool run sprawl … fails to resolve sprawl The PyPI package is sprawl-graph; sprawl is only the command it provides. uv tool run treats the first word as a package name, so plain uvx sprawl … looks for a package literally called sprawl — which is someone else's unrelated package, not this one. Name the package explicitly: uvx --from sprawl-graph sprawl install (same as uv tool run --from sprawl-graph sprawl install). Or uv tool install sprawl-graph once and then call sprawl directly.

uv run --with sprawl-graph python -m sprawl silently runs an older install uv run uses your system Python, so if an older sprawl also lives there (e.g. a past pip install sprawl-graph), Python can find that copy first on sys.path and --with sprawl-graph won't override it. It runs with no error, but you get the old version's behavior — e.g. env overrides like OPENAI_BASE_URL are silently ignored, so requests hit the default endpoint and fail with a 401 that looks like a bad key. The fingerprint is a warning: skill is from sprawl <newer>, package is <older> line — that means a different install was loaded, not just a stale skill. Check which copy actually loaded:

python -c "import sprawl; print(sprawl.__file__)"

Then run the installed command directly (it uses the uv-managed copy), or drop the stale system copy:

uvx --from sprawl-graph sprawl extract . --backend openai   # names the package explicitly
pip uninstall sprawl-graph                                   # or remove the old system install

python -m sprawl works but sprawl command doesn't Your shell's PATH doesn't include the bin directory the command was installed to. Prefer uv tool install / pipx install over plain pip, then run uv tool update-shell / pipx ensurepath and open a new terminal (see the install notes above).

/sprawl . causes "path not recognized" in PowerShell PowerShell treats a leading / as a path separator. Use sprawl . (no slash) on Windows.

Graph has fewer nodes after --update or rebuild If a refactor deleted files, the old nodes linger. Pass --force (or set SPRAWL_FORCE=1) to overwrite even when the rebuild has fewer nodes.

extract exits with "extraction was incomplete ... refusing to overwrite" When an extraction pass crashes or a walk can't fully read the corpus, the run would be smaller than a complete one, so sprawl extract refuses to overwrite a larger existing graph with the partial result (protecting your graph.json). Fix the underlying failure and re-run, or pass --allow-partial to overwrite anyway.

Graph has duplicate nodes for the same entity (ghost duplicates) Ghost duplicates (same symbol appearing twice — once from AST extraction with a source location, once from semantic extraction without) are now automatically merged at build time. If you see this in a graph built before v0.8.33, run a full re-extract to clean up:

sprawl extract . --force

Ollama runs out of VRAM / context window exceeded The KV-cache window is auto-sized but may be too large for your GPU. Reduce it:

SPRAWL_OLLAMA_NUM_CTX=8192 sprawl extract ./docs --backend ollama --token-budget 4000

LLM returned invalid JSON / Unterminated string warnings The model's JSON response hit its output-token limit and was cut off mid-string. sprawl auto-recovers (it splits the chunk and re-extracts the halves, and an oversized single document is first sliced at heading/paragraph boundaries so the whole file is still covered), so these warnings are noisy but not data loss. To reduce the churn, raise the output cap or shrink each chunk's output:

SPRAWL_MAX_OUTPUT_TOKENS=16384 sprawl extract . --mode deep   # lift the cap
sprawl extract . --mode deep --token-budget 4000                # smaller input chunks -> smaller output

With a cloud gateway like OpenRouter, prefer --backend openai (set OPENAI_BASE_URL) over the Ollama shim — it's a cleaner OpenAI-compatible path. If the model has its own max-output ceiling, lowering --token-budget is the reliable lever.

Graph HTML is too large to open in a browser (>5000 nodes) Skip HTML generation and use the JSON directly:

sprawl cluster-only ./my-project --no-viz
sprawl query "..."

graph.json has conflict markers after two devs commit at once Run sprawl hook install — it sets up a git merge driver that union-merges graph.json automatically so conflicts never happen.

Extraction returns empty nodes/edges for docs or PDFs Docs, PDFs, and images require an LLM call — code-only corpora need no key. Check that your API key is set and the backend is correct:

ANTHROPIC_API_KEY=sk-... sprawl extract ./docs --backend claude

Skill version mismatch warning in your IDE Your installed sprawl version is different from the skill file. Update:

uv tool upgrade sprawl
sprawl install  # overwrites the skill file

Claude Code prompt cache invalidated after every sprawl extract Sprawl writes output files (graph.json, sprawl-out/) into the workspace. If those paths aren't ignored, every write invalidates Claude Code's prompt cache, forcing a full re-upload at cache-write rates on the next turn. Add them to .claudeignore:

# .claudeignore
graph.json
sprawl-out/

Full command reference

/sprawl                          # run on current directory
/sprawl ./raw                    # run on a specific folder
/sprawl ./raw --mode deep        # more aggressive relationship extraction
sprawl extract ./raw --code-only # index code only — local AST, no API key (skips docs/PDFs/images); an `extract` flag, not a skill flag
/sprawl ./raw --update           # re-extract only changed files
/sprawl ./raw --directed         # preserve edge direction
/sprawl ./raw --cluster-only     # rerun clustering on existing graph
/sprawl ./raw --no-viz           # skip HTML visualization
/sprawl ./raw --obsidian         # generate Obsidian vault
/sprawl ./raw --obsidian --obsidian-dir ~/vault  # write into an existing vault (never overwrites your own notes or .obsidian config)
/sprawl ./raw --wiki             # build agent-crawlable markdown wiki
/sprawl ./raw --svg              # export graph.svg
/sprawl ./raw --graphml          # export for Gephi / yEd
/sprawl ./raw --neo4j            # generate cypher.txt for Neo4j
/sprawl ./raw --neo4j-push bolt://localhost:7687
/sprawl ./raw --falkordb         # generate cypher.txt for FalkorDB
/sprawl ./raw --falkordb-push falkordb://localhost:6379
/sprawl ./raw --watch            # auto-sync as files change
/sprawl ./raw --mcp              # start MCP stdio server

/sprawl add https://arxiv.org/abs/1706.03762
/sprawl add <video-url>
/sprawl add https://... --author "Name" --contributor "Name"

/sprawl query "what connects attention to the optimizer?"
/sprawl query "..." --dfs --budget 1500
/sprawl path "DigestAuth" "Response"
/sprawl explain "SwinTransformer"

sprawl save-result --question "Q" --answer "A" --nodes Foo Bar --outcome useful   # record how a Q&A turned out (work memory; outcome ∈ useful|dead_end|corrected)
sprawl reflect                   # aggregate sprawl-out/memory/ outcomes into reflections/LESSONS.md
sprawl reflect --if-stale        # no-op when LESSONS.md is already newer than every input (cheap to run each session)
sprawl reflect --out docs/LESSONS.md    # write the lessons doc somewhere else
sprawl reflect --graph sprawl-out/graph.json  # group lessons by community + write the work-memory overlay (.sprawl_learning.json)
                                   # the overlay tags nodes preferred/tentative/contested (recency-weighted, with provenance);
                                   # sprawl explain / query then show a "Lesson:" hint, flagged "code changed — re-verify" when the source moved on

sprawl uninstall                 # remove from all platforms in one shot
sprawl uninstall --purge         # also delete sprawl-out/
sprawl uninstall --project --platform codex  # remove project-scoped install files only

sprawl hook install              # post-commit + post-checkout hooks
sprawl hook uninstall
sprawl hook status

# always-on assistant instructions - platform-specific
sprawl claude install            # CLAUDE.md + PreToolUse hook (Claude Code)
sprawl claude uninstall
sprawl codebuddy install         # CODEBUDDY.md + PreToolUse hook (CodeBuddy)
sprawl codebuddy uninstall
sprawl codex install             # AGENTS.md + PreToolUse hook in .codex/hooks.json (Codex)
sprawl opencode install          # AGENTS.md + tool.execute.before plugin (OpenCode)
sprawl kilo install              # native Kilo skill + /sprawl command + AGENTS.md + .kilo plugin
sprawl kilo uninstall
sprawl cursor install            # .cursor/rules/sprawl.mdc (Cursor)
sprawl cursor uninstall
sprawl gemini install            # GEMINI.md + BeforeTool hook (Gemini CLI)
sprawl gemini uninstall
sprawl copilot install           # skill file (GitHub Copilot CLI)
sprawl copilot uninstall
sprawl aider install             # AGENTS.md (Aider)
sprawl aider uninstall
sprawl claw install              # AGENTS.md (OpenClaw)
sprawl claw uninstall
sprawl droid install             # AGENTS.md (Factory Droid)
sprawl droid uninstall
sprawl trae install              # AGENTS.md (Trae)
sprawl trae uninstall
sprawl trae-cn install           # AGENTS.md (Trae CN)
sprawl trae-cn uninstall
sprawl hermes install             # AGENTS.md + ~/.hermes/skills/ (Hermes)
sprawl hermes uninstall
sprawl amp install               # skill file (Amp)
sprawl amp uninstall
sprawl agents install            # ~/.agents/skills/ + AGENTS.md (cross-framework; alias: sprawl skills)
sprawl agents uninstall
sprawl kiro install               # .kiro/skills/ + .kiro/steering/sprawl.md (Kiro IDE/CLI)
sprawl kiro uninstall
sprawl pi install                # skill file (Pi coding agent)
sprawl pi uninstall
sprawl devin install             # skill file + .windsurf/rules/sprawl.md (Devin CLI)
sprawl devin uninstall
sprawl antigravity install       # .agents/rules + .agents/workflows (Google Antigravity)
sprawl antigravity uninstall

sprawl extract ./docs                        # headless LLM extraction for CI (no IDE needed)
sprawl extract ./docs --backend gemini       # explicit backend: gemini, kimi, claude, openai, deepseek, ollama, bedrock, or claude-cli
sprawl extract ./docs --backend gemini --model gemini-3.1-pro-preview
sprawl extract ./docs --backend ollama       # local Ollama (set OLLAMA_BASE_URL / OLLAMA_MODEL) - no API key needed for loopback
OPENAI_BASE_URL=http://localhost:8080/v1 OPENAI_MODEL=my-model sprawl extract ./docs --backend openai   # any OpenAI-compatible server (llama.cpp, vLLM, LM Studio)
ANTHROPIC_BASE_URL=http://localhost:4000 ANTHROPIC_MODEL=my-model sprawl extract ./docs --backend claude   # any Anthropic-compatible endpoint (LiteLLM proxy, gateways)
SPRAWL_OLLAMA_NUM_CTX=32768 sprawl extract ./docs --backend ollama   # override KV-cache window (auto-sized by default)
SPRAWL_OLLAMA_KEEP_ALIVE=0 sprawl extract ./docs --backend ollama    # unload model after each chunk (saves VRAM on small GPUs)
sprawl extract ./docs --backend bedrock      # AWS Bedrock via IAM - no API key, uses AWS credential chain
sprawl extract ./docs --backend claude-cli   # route through Claude Code CLI - no API key, uses your Claude subscription
sprawl extract ./docs --backend azure        # Azure OpenAI (set AZURE_OPENAI_API_KEY + AZURE_OPENAI_ENDPOINT)
sprawl extract ./docs --max-workers 16       # AST parallelism (also SPRAWL_MAX_WORKERS)
sprawl extract --postgres "postgresql://user:pass@host/db"   # introspect live PostgreSQL schema directly
sprawl extract ./my-workspace --cargo        # introspect Rust Cargo workspace dependencies directly
sprawl extract ./docs --token-budget 30000   # smaller semantic chunks for local/small models
sprawl extract ./docs --max-concurrency 2    # fewer parallel LLM calls (useful for local inference)
sprawl extract ./docs --api-timeout 900      # longer HTTP timeout for slow local models (default 600s)
sprawl extract ./docs --google-workspace     # export .gdoc/.gsheet/.gslides via gws before extraction
sprawl extract ./src --no-gitignore          # include git-ignored source; still honor .sprawlignore
sprawl extract ./docs --mode deep            # richer semantic extraction via extended system prompt
sprawl extract ./docs --no-cluster           # raw extraction only, skip clustering
sprawl extract ./docs --timing               # print per-stage wall-clock timings to stderr (also works on cluster-only)
sprawl extract ./docs --force                # overwrite graph.json even if new graph has fewer nodes (use after refactors or to clear ghost duplicates)
sprawl extract ./docs --dedup-llm            # LLM tiebreaker for ambiguous entity pairs (uses same API key)
sprawl extract ./docs --global --as myrepo   # extract and register into the cross-project global graph
SPRAWL_MAX_OUTPUT_TOKENS=32768 sprawl extract ./docs --backend claude  # raise output cap for dense corpora

sprawl ask "how does workflow execution work"   # answer-first: <=250-word cited answer, docs excerpt first when a doc covers it
sprawl ask "why does X exist" --verbose         # append the ranked-node appendix (scores, topical similarity, what was demoted and why)
sprawl ask "how are errors handled" --no-llm    # extractive answer built from excavated prose; no model, no network
sprawl ask "..." --backend claude-cli           # force a synthesis backend (default: auto-detect, degrade to extractive)
sprawl ask "..." --top-k 12 --budget 3000       # widen the evidence set / raise the assembled-context token cap (default 8 / 2000)
sprawl ask "..." --json                         # machine-readable: answer, citations, doc excerpt, full ranking with scores
sprawl benchmark --ask                          # score `ask` against the old `query` on the same questions

sprawl excavate                                 # why-mining pass on cwd: git-history churn/scars/summary + LLM intent (why/scars/gotchas/deprecated), written into graph.json
sprawl excavate ./my-repo                       # excavate a specific repo (graph.json must already exist at ./my-repo/sprawl-out/graph.json — run `sprawl extract` first)
sprawl excavate --no-llm                        # deterministic pass only: churn_score, scar_count, summary — no network/LLM calls at all
sprawl excavate --backend claude-cli            # force a specific distiller backend (claude-cli/codex-cli need no API key; also accepts any `sprawl extract --backend` name, or ollama)
sprawl excavate --limit 50                      # cap how many file-level nodes are processed (cost control while testing)
sprawl excavate --github                        # also mine the GitHub remote: commits → PRs → review threads → linked issues (PR/issue text is the strongest "why" signal; commit messages are the fallback)
sprawl excavate --github --no-llm               # GitHub without an LLM: writes PR-backed scars[] citing real PR numbers, entirely from PR/issue titles
#   auth: GITHUB_TOKEN (or GH_TOKEN), else an authenticated `gh` CLI (`gh auth login` once — no env var needed)
#   cache: every API response is cached in sprawl-out/cache/github/ keyed by (repo, endpoint); a re-run re-fetches nothing
#   resume: completed nodes are recorded in sprawl-out/cache/github/_progress.json and skipped on the next run
#   no remote / no credentials / rate-limited → warns and falls back to git-only mode, never fails the run

sprawl export callflow-html                       # sprawl-out/<project>-callflow.html
sprawl export callflow-html --max-sections 8      # cap generated architecture sections
sprawl export callflow-html --output docs/arch.html
sprawl export callflow-html ./some-repo/sprawl-out

sprawl global add sprawl-out/graph.json --as myrepo   # register a project graph into ~/.sprawl/global-graph.json
sprawl global remove myrepo                         # remove a project from the global graph
sprawl global list                                  # show all registered repos + node/edge counts
sprawl global path                                  # print path to the global graph file

sprawl prs                              # PR dashboard: CI, review, worktree, graph impact
sprawl prs 42                           # deep dive on PR #42
sprawl prs --triage                     # AI triage ranking (auto-detects backend from env)
sprawl prs --worktrees                  # worktree → branch → PR mapping
sprawl prs --conflicts                  # PRs sharing graph communities (merge-order risk)
sprawl prs --base main                  # filter to PRs targeting a specific base branch
sprawl prs --repo owner/repo            # run against a different GitHub repo
SPRAWL_TRIAGE_BACKEND=kimi sprawl prs --triage   # use a specific backend for triage

sprawl clone https://github.com/karpathy/nanoGPT
sprawl merge-graphs a.json b.json --out merged.json
sprawl --version                                    # print installed version
sprawl watch ./src
sprawl check-update ./src
sprawl update ./src
sprawl update ./src --no-cluster  # skip reclustering, write raw AST graph only
sprawl update ./src --force       # overwrite even if new graph has fewer nodes
sprawl cluster-only ./my-project
sprawl cluster-only ./my-project --graph path/to/graph.json  # custom graph location
sprawl cluster-only ./my-project --max-concurrency 16 --batch-size 200  # parallel community labeling (large graphs)
sprawl cluster-only ./my-project --resolution 1.5            # more, smaller communities
sprawl cluster-only ./my-project --exclude-hubs 99           # exclude p99 degree nodes from partitioning
sprawl cluster-only ./my-project --no-label                  # keep "Community N" placeholders
sprawl cluster-only ./my-project --backend=gemini            # backend for community naming
sprawl cluster-only ./my-project --backend=gemini --model gemini-2.5-pro  # specific model
sprawl label ./my-project                                    # (re)name communities with the configured backend
sprawl label ./my-project --backend=openai --model gpt-4o   # force a specific backend and model

sprawl affected "UserService"                   # reverse traversal: what breaks if this changes
sprawl affected "UserService" --depth 3 --relation calls   # deeper, scoped to one edge type
sprawl god-nodes                                # the most-connected nodes (architectural hubs)
sprawl god-nodes --top 20 --json                # more results, machine-readable
sprawl diagnose multigraph                      # report same-endpoint edge collapse risk in graph.json
sprawl check-update ./src                       # check whether a semantic re-extraction is pending (cron-safe)
sprawl tree                                     # D3 collapsible-tree HTML for graph.json (sprawl-out/GRAPH_TREE.html)

sprawl drift                                    # which excavated intent no longer matches its code (git-verified)
sprawl drift --fix                              # re-excavate only the flagged nodes (needs an LLM backend)
sprawl drift --json --no-write                  # report only, machine-readable, leaves graph.json untouched

sprawl biography .                              # long-form narrative biography of a codebase; see above
sprawl biography . --md --out docs/BIOGRAPHY.md
sprawl biography https://github.com/org/repo --branch main
sprawl biography . --full-extract               # semantic (LLM) extraction too, not just AST/code

sprawl capture --note "why we chose X over Y"   # file this session's context into sprawl-out/memory/
sprawl capture --asked "..." --decided "..." --constraint "..."
sprawl capture --dry-run                        # print the note that would be written, without writing it

Community names: inside an agent (Claude Code, Gemini CLI) the agent names communities itself. When you run the bare CLI, cluster-only auto-names them with the configured backend (built-in or custom OpenAI-compatible provider) — pass --no-label to keep Community N, or run sprawl label to (re)generate names on demand.


More

The core story is git-history mining (excavate / drift / biography) and mapping a large codebase into a queryable knowledge graph. sprawl also covers a few adjacent integrations, kept out of the way up top so they don't crowd that story. None of these are required for the core workflow — pull in only what you need.

PR triage

sprawl prs is a PR dashboard: CI state, review status, worktree mapping, and (optionally) AI-ranked review-queue triage, cross-referenced against the graph so you can see which PRs touch the same communities.

sprawl prs                       # PR dashboard: CI state, review status, worktree mapping
sprawl prs 42                    # deep dive on PR #42 with graph impact
sprawl prs --triage              # AI ranks your review queue (uses whatever backend is configured)
sprawl prs --worktrees           # worktree → branch → PR mapping
sprawl prs --conflicts           # PRs sharing graph communities — merge-order risk
sprawl prs --base main           # filter to PRs targeting a specific base branch
sprawl prs --repo owner/repo     # run against a different GitHub repo
SPRAWL_TRIAGE_BACKEND=kimi sprawl prs --triage   # use a specific backend for triage

The MCP server exposes the same capability as the list_prs, get_pr_impact, and triage_prs tools.

Google Workspace ingestion

Google Drive for desktop .gdoc, .gsheet, and .gslides files are shortcut pointers, not document content. To include native Google Docs, Sheets, and Slides in a headless extraction, install and authenticate the gws CLI, then run:

uv tool install "sprawl-graph[google]"  # needed for Google Sheets table rendering
gws auth login -s drive
sprawl extract ./docs --google-workspace

You can also set SPRAWL_GOOGLE_WORKSPACE=1. Sprawl exports shortcuts into sprawl-out/converted/ as Markdown sidecars, then extracts those files.

Video & audio transcription

Video and audio files are transcribed locally with faster-whisper (nothing leaves your machine) and added to the graph like any other source:

uv tool install "sprawl-graph[video]"      # faster-whisper + yt-dlp
/sprawl add <youtube-url>                   # transcribe and add a video

Supported extensions: .mp4 .mov .mp3 .wav and more — see What files it handles.

Database & workspace introspection

sprawl can introspect a live database schema or a Rust workspace's dependency graph directly, instead of parsing source files:

uv tool install "sprawl-graph[postgres]"
sprawl extract --postgres "postgresql://user:pass@host/db"   # introspect live PostgreSQL schema directly
sprawl extract ./my-workspace --cargo                          # introspect Rust Cargo workspace dependencies directly

Graph database export

Push the graph into Neo4j or FalkorDB for Cypher-based querying outside sprawl:

uv tool install "sprawl-graph[neo4j]"       # or [falkordb]
/sprawl ./raw --neo4j                       # generate cypher.txt for Neo4j
/sprawl ./raw --neo4j-push bolt://localhost:7687
/sprawl ./raw --falkordb                    # generate cypher.txt for FalkorDB
/sprawl ./raw --falkordb-push falkordb://localhost:6379

See the full command reference above for every flag on every command, including these.


Learn more

  • How it works — the extraction pipeline, community detection, confidence scoring, benchmarks
  • ARCHITECTURE.md — module breakdown, how to add a language
  • Optional integrations — Docker MCP Toolkit + SQLite
  • The Memory Layer — the upstream author's book on the ideas behind the graphify engine, the architecture end to end

Contributing

Development setup

The project uses uv for dev workflow. Install it once, then:

git clone <this repository>
cd sprawl

# Create the project venv and install sprawl + all extras + the dev group
# (pytest). uv installs the dev dependency group by default; pass --no-dev to
# skip it.
uv sync --all-extras

Verify the editable install:

uv run sprawl --version
uv run python -c "import sprawl; print(sprawl.__file__)"

Running tests

uv run pytest tests/ -q                # run the full suite
uv run pytest tests/test_extract.py -q # one module
uv run pytest tests/ -q -k "python"    # filter by name

macOS note: the test suite includes both sample.f90 and sample.F90 fixtures. These collide on case-insensitive HFS+ / APFS file systems. Run on Linux or in a Docker container if you need to test both Fortran variants simultaneously.

Git workflow

  • Active development happens on the v8 branch.
  • Commit style: fix: <description> / feat: <description> / docs: <description>
  • Before opening a PR, run uv run pytest tests/ -q and confirm it passes.
  • Add a fixture file to tests/fixtures/ and tests to tests/test_languages.py for any new language extractor.

What to contribute

Worked examples are the most useful contribution. Run /sprawl on a real corpus, save the output to worked/{slug}/, write an honest review.md covering what the graph got right and wrong, and open a PR.

Extraction bugs — open an issue with the input file, the cache entry (sprawl-out/cache/), and what was missed or wrong.

See ARCHITECTURE.md for module responsibilities and how to add a language.


Translations

Read this in other languages

🇺🇸 English | 🇨🇳 简体中文 | 🇯🇵 日本語 | 🇰🇷 한국어 | 🇩🇪 Deutsch | 🇫🇷 Français | 🇪🇸 Español | 🇮🇳 हिन्दी | 🇧🇷 Português | 🇷🇺 Русский | 🇸🇦 العربية | 🇮🇷 فارسی | 🇮🇹 Italiano | 🇵🇱 Polski | 🇳🇱 Nederlands | 🇹🇷 Türkçe | 🇺🇦 Українська | 🇻🇳 Tiếng Việt | 🇮🇩 Bahasa Indonesia | 🇸🇪 Svenska | 🇬🇷 Ελληνικά | 🇷🇴 Română | 🇨🇿 Čeština | 🇫🇮 Suomi | 🇩🇰 Dansk | 🇳🇴 Norsk | 🇭🇺 Magyar | 🇹🇭 ภาษาไทย | 🇺🇿 Oʻzbekcha | 🇹🇼 繁體中文 | 🇵🇭 Filipino | 🇮🇱 עברית


Upstream project

sprawl is a fork of graphify. The upstream project has its own site, community and sponsorship channels — see its repository for those. sprawl is maintained independently; issues and discussion for sprawl belong on this repository.

About

Mine a codebase's git history for why the code is the way it is — not just what it is.

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages