Compiler-exact C++ code intelligence for AI agents. It answers the questions your
assistant otherwise guesses at with grep — who calls this function, what it
calls, who references this type, the class hierarchy, the change impact / blast
radius, the call path between two symbols, and where the busiest fan-in/fan-out
symbols in the whole project are — exposed as MCP tools your agent calls
directly (who_calls, what_it_calls, find_references, impact_of, path,
base_classes / subclasses, hotspots), backed by a pre-indexed graph, plus a CLI.
Built from a compiler front-end index (SCIP via scip-clang), not a syntactic
AST: every symbol has a stable, unique identity (USR / mangled name), so edges are
exact — overloads, ptr->method(), free functions, templates, and virtual dispatch
all resolved, with no name-collision merges.
Why it matters for C++. By-name graph tools (tree-sitter) and LSP-based agents (e.g. Serena) degrade on large C++ — templates, macros, overloads — so the agent silently falls back to grep, which merges distinct same-named symbols, counts comments/decls as calls, and misses
ptr->method()/ virtual dispatch. cppgraph is the graph the compiler actually sees — see COMPARISON.md for measured numbers (grep-and-read vs cppgraph, and vs graphify / Serena).
To install, tell your agent: "Follow instructions to install https://github.com/rakiz/cppgraph" — that points it at this section.
If you are that agent, follow these steps exactly. They are deterministic; the scripts fail loudly with actionable messages.
Two phases. Phase A (machine setup): clone, then run setup with the scip-clang source the user chose. Phase B (index a project): run the index with the scope the user chose. You gather each choice by asking the user (in your own question UI), then run the script non-interactively with those choices as flags.
RULE —
!commands can't be answered, so you ask, then pass flags. A Claude Code! …run has no interactive stdin: a prompt inside it gets EOF. So the interactive menus are for a human in a real terminal, not for you. Your loop is: (1) ask the user every choice via your question UI — the scip-clang source (Phase A) and the index scope (Phase B) — surfacing the real options; (2) run the script non-interactively with those choices as flags. Decide nothing yourself; the answers are the user's. Give a realistic time estimate with each heavy step: indexing is the long one (minutes to hours, CPU-bound — ~20 min for ~6 000 TUs on a fast 14-core x86, ~4 h for 6 482 TUs on an 8-core Graviton2 (m6g), ~1 h 42 for 4 301 TUs on an 8-core Graviton3 (m7g.2xlarge)); downloading scip-clang is seconds, compiling it (#504) is ~25–60 min (~26 min measured on m7g.2xlarge).
- Before cloning anything cppgraph-related — even just to read/explore the
source — check for an existing per-machine install first: if
${XDG_DATA_HOME:-$HOME/.local/share}/cppgraph/repoexists, use that checkout (read it,git pullit if stale); never clone a second copy (~/cppgraph, a project dir, a dev checkout). Only clone fresh if absent. - Clone into the per-machine tool dir (always this path — not the user's
project, not a dev checkout; even "install from a local copy" means clone that
copy into this path):
git clone https://github.com/rakiz/cppgraph "${XDG_DATA_HOME:-$HOME/.local/share}/cppgraph/repo" - Get the valid sources from the tool (don't guess the platform):
It prints this machine's OS/arch and the sources that actually apply (e.g. no
! ~/.local/share/cppgraph/repo/scripts/setup.sh --list-sourcesdownloadon ARM-Linux). Ask the user to pick from exactly those — costs: download-patched ~1 min (patched binary, no Docker; macOS arm64 + Linux aarch64 today), download ~1 min (stock, unpatched), build (patched) ~25–60 min Docker, emulate (slower indexing). Windows → WSL2; Intel Mac → only emulate. - Run it with their choice (
!runs it non-interactively;--scip-sourceis what makes that work — without it a piped run stops withACTION NEEDED):! ~/.local/share/cppgraph/repo/scripts/setup.sh --scip-source <download-patched|download|build|emulate>setup.shis the sole entry point — it creates the venv, obtains scip-clang, registers the MCP server, installs the bundled skill and/cppgraphslash command into detected agent tools, and (in a real terminal) offers to index the current project. Never tell the user to runcppgraph setup/.venv/bin/cppgraph …/ anything under a dev checkout: none of that exists on a fresh machine untilsetup.shhas run. - Never
pip install cppgraph(or add it as a dependency) into a target project's own venv/virtualenv. It's a single per-machine tool with one dedicated venv (~/.local/share/cppgraph/repo/.venv); invoke its binaries from there directly (optionally aliascppgraphto that binary), never re-install it into another project's environment — that creates a second divergent install with its own dependency versions that can conflict with the committed generated bindings (observed: a project venv's pinnedprotobufvs cppgraph's). Setup also links the binary into~/.local/binautomatically (skipped when acppgraphis already on PATH), so a barecppgraphworks after setup; the shell alias below remains an alternative (e.g. for older installs predating the link). Convenience alias for the user's shell rc (~/.bashrc/~/.zshrc):alias cppgraph="${XDG_DATA_HOME:-$HOME/.local/share}/cppgraph/repo/.venv/bin/cppgraph"
You drive Phase B; the user only picks from choices you offer — never typing a scope, path, or flag, and never seeing these commands. The conversation: (a) a yes/no — "Index this project now? (~estimate)"; if yes, (b) a selectable scope choice (whole tree / a subtree) + a yes/no on tests; then (c) you run it and report progress; (d) when done, they open a new Claude Code session.
- Get the scope options: run
~/.local/share/cppgraph/repo/scripts/index.sh --plan-json(a barecppgraphworks when setup linked it into~/.local/bin; otherwise use the full path~/.local/share/cppgraph/repo/.venv/bin/cppgraph, the fallback on older installs) from the project dir; it auto-locates thecompile_commands.json— if it reports none, generate one, see AGENTS.md → "Fallback", after the user's OK. It returns the compdb breakdown, the questions (subtree / tests / attribution), andartifacts(whether a.scip/.graph.dbalready exists). - Offer the choices — a single-select of whole tree + each subtree (with TU
counts) and a yes/no on tests; never ask the user to type a subdirectory or flag.
If
artifactsshows the project is already indexed, ask whether to keep it (default) or rebuild. - Run it with their answers (non-interactive, so
!works):Non-destructive by default: an existing! ~/.local/share/cppgraph/repo/scripts/index.sh <compdb> -y --filter <sub> [--no-tests] [--attributed-refs] --run.scip/.graph.dbis kept, not overwritten — only what's missing is built. Add--from-scratchonly if the user asked to rebuild. The chosen scope is recorded and reused by later updates. (A human in a real terminal can instead runscripts/index.shwith no flags for the interactive wizard.) Then tell the user to open a new Claude Code session from their project directory (that's how the server finds this project's graph) and ask "what calls X?", "impact of changing Y?", "show the dependency graph of Z".
Uninstalling: run ! ~/.local/share/cppgraph/repo/scripts/uninstall.sh (it asks
per item). There is no uninstall flag on setup.sh, and you must never improvise
rm -rf on ~/.local/share/cppgraph or a project's .cppgraph/ — the latter holds
a .scip index that can take hours to rebuild; uninstall.sh keeps project data by
default and warns before touching anything precious.
Humans: the same flow, step by step, is in QUICKSTART.md. Every CLI subcommand, grouped by purpose with one example each, is in CLI_REFERENCE.md.
Early, but functional end-to-end: build, query, incremental update, an MCP
server for LLMs, and visualization. The tool is general — point it at any C++
project's compile_commands.json.
scip-clang compile_commands.json → <name>.compdb.json → <name>.scip → <name>.graph.db
(indexer binary) (target's build) (filtered subset) (SCIP protobuf) (SQLite: query/MCP/viz)
Five stages get you from nothing to a queryable graph. Only the last three are cppgraph's; the timings are rough and dominated by your machine (the indexer runs the C++ front-end once per translation unit, so it scales with cores/CPU).
| # | Stage | Produces | Rough time |
|---|---|---|---|
| 1 | Get scip-clang — copy the prebuilt binary, or compile from source |
the per-machine indexer binary | ~1 min (copy) … ~25–60 min (compile, #504; ~26 min measured on m7g) |
| 2 | Get compile_commands.json — from the target's build system, if not already present |
compile_commands.json |
~3 min (warm) … ~15 min (cold) — the target's build, not cppgraph |
| 3 | Filter what to index — scope to a subtree, optionally drop tests | <name>.compdb.json |
seconds |
| 4 | Index — scip-clang runs the C++ front-end per TU → SCIP |
<name>.scip |
the big one: ~20 min → several hours |
| 5 | Build the store — parse SCIP, intern symbols + edges into SQLite | <name>.graph.db |
~30 s (filtered src/mongo) … ~3.5 min (full repo + refs) |
Then you're ready to query (see the gains below).
Notes:
- Stage 1 is once per machine. macOS arm64 and Linux aarch64 get a
native prebuilt patched binary (
download-patched, ~1 min); Linux x86_64 gets the native stock prebuilt (download, ~1 min, unpatched — #504 there means a local build). All three index natively — no Docker. Docker is only for thebuildsource (a #504 compile, ~25–60 min, Linux-only) or for hosts with no native binary at all (Intel Mac, Windows) viaemulate. "#504 build is Linux-only" ≠ "scip-clang needs Docker on macOS". - Stage 4 is the variable one. Measured extremes: ~20 min on a fast 14-core
x86 for ~6 000 TUs; ~4 h for 6 482 TUs on an 8-core AWS Graviton2
(
m6g.2xlarge— older Neoverse-N1 cores are slow at this); ~1 h 42 for 4 301 TUs on an 8-core Graviton3 (m7g.2xlarge— ~1.4 s/TU vs ~2.2 on Graviton2). the index wizard prints a per-machine estimate right before it starts. Fewer/older cores → proportionally longer; scope to a subtree (stage 3) to cut it down. enclosing_range(#504) is a choice at indexing, with consequences: it enables exact caller attribution and the symbol-granularity usage view (--attributed-refs), but needs the compiled #504 binary and makes the.scipand store larger. Without it you still get an exact graph, just with file-granularity usage. You can add it later without re-indexing (enrich-refs) — but noteenrich-refsre-parses the.scipand rebuilds the graph with references, so it costs a full store build (mongo: ~3.5 min, ~9 GB RAM) — the top of stage 5's range, not its filtered ~30 s floor.
Two components: the builder (scip-clang, external compiled binary) does
the expensive, perf-critical C++ parsing, crash-isolated per TU; cppgraph
(this repo, Python) is the glue + graph — parses SCIP, builds the store, serves
queries, exposes the MCP server, and exports a graph.json for visualization.
Yes, measurably. On a real case (two distinct methods sharing the name
makeResumeToken), a tree-sitter tool drops the real call edges and collapses
431 unrelated Value sites onto one node; cppgraph returns the correct,
separated caller sets. Full write-up with numbers and reproduction steps:
COMPARISON.md (cppgraph vs graphify vs Serena/LSP, on a
large C++ codebase).
Your AI assistant already answers "what calls X?" with grep — but on a large
C++ codebase that's wrong (matches by name: merges distinct symbols, includes
comments/decls, misses ptr->method() / virtual dispatch / templates) and
token-expensive (noisy output, then whole-file reads to disambiguate, all
through the model's context).
grep's raw dump is only a floor: it can't tell a call from a decl, a comment,
or a different same-named symbol, so to answer correctly it must read around
every hit — and that's where the cost explodes. Measured on MongoDB, "who
calls the method makeResumeToken?": grep dumps ~6,600 tokens (98% noise, 4
symbols merged) and needs ~110,000 to read-and-disambiguate; cppgraph find
who_callsingests ~400, exact — 272× leaner. On a rare, uniquely named symbol (grep's best case) cppgraph still wins ~8× once grep reads to verify; on a hot type likeOperationContextreading every hit would take ~1,900× more than cppgraph — a theoretical number nobody ingests, so in practice grep truncates and answers from a partial view, silently missing call sites, where cppgraph stays at ~6k tokens, exact and complete. The trade-off is a one-time index (minutes), amortized over every later query. Full spectrum, noise ratios, and the token-lean defaults: COMPARISON.md (reproduce withscripts/measure_tokens.py --suite).
Same idea versus an LSP-driven agent (e.g. Serena): the questions are the same, but cppgraph is a pre-indexed graph the agent queries in one call, not a live language server it has to drive file-by-file — which is what keeps it token-lean and complete on large C++. Numbers: COMPARISON.md.
| Doc | What's in it |
|---|---|
| AGENTS.md | Working instructions, principles, guardrails — read first |
| DESIGN.md | Architecture, edge model, the call-attribution heuristic + its known limitation |
| COMPARISON.md | Measured comparison vs graphify and Serena on a real design question |
| INSTALL.md | Setting up a new machine (scip-clang, protoc, the venv) |
| viz/README.md | The bundled graph viewer + cppgraph export graph.json format |
| CHANGELOG.md | What's been built so far |
| TODO.md | Open tasks |
- Re-implementing a C++ parser. We consume a compiler index.
- Being a linter or a refactoring engine. This is about understanding structure (who calls what, impact/blast-radius, inheritance) for humans and LLMs.
Secondary to the query tools: cppgraph export <symbol> --depth 2 --out graph.json
(run from the indexed project — graph auto-discovered, <symbol> a plain name or
exact SCIP string), or the visualize MCP tool, writes a bounded neighbourhood and
opens it in a self-contained, offline viewer (viz/cppgraph-viz.html). The
graph.json is graphify-compatible.
export/view (and the MCP visualize) also take --mode path --dst <target> to
draw how two symbols connect — the shortest calls chain by default,
--expand-paths for the full corridor of every route between them — or
--mode cycle to draw the multi-member call cycle containing a symbol.
boundary-violations --out (or the MCP visualize_boundary_violations) draws
the edges that cross your declared layering rules.
Details: viz/README.md.
The call graph can't answer it — a plain struct has no callers. cppgraph keeps an
exact reference index (on by default) so cppgraph references <type> (and the
find_references tool) lists every use site, and export --mode usage draws a
usage graph.
By default that graph is at file granularity ("used somewhere in these
files"). Built with a scip-clang that emits enclosing_range (a source build
carrying PR #504), you can upgrade it to symbol
granularity — "used by these functions" — with cppgraph build --attributed-refs, or add it to an existing graph with cppgraph enrich-refs.
Worth it when you want symbol-level usage — at a cost. Attribution stores one extra symbol id per reference, so the graph grows. Enable it when "which functions use this type?" matters; otherwise the default file granularity is already exact and leaner.
cppgraph statustells you which granularity a graph has and how to upgrade.
MIT. The bundled viewer vendors vis-network (MIT / Apache-2.0); see viz/README.md.