Skip to content
rakizPublic

About

Compiler-exact C++ code intelligence for AI agents (MCP + CLI): who-calls, references, class hierarchy, change-impact — from a pre-indexed SCIP graph, so the agent stops falling back to grep.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

183 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

cppgraph

Compiler-exact C++ code intelligence for AI agents. It answers the questions your assistant otherwise guesses at with grep — who calls this function, what it calls, who references this type, the class hierarchy, the change impact / blast radius, the call path between two symbols, and where the busiest fan-in/fan-out symbols in the whole project are — exposed as MCP tools your agent calls directly (who_calls, what_it_calls, find_references, impact_of, path, base_classes / subclasses, hotspots), backed by a pre-indexed graph, plus a CLI.

Built from a compiler front-end index (SCIP via scip-clang), not a syntactic AST: every symbol has a stable, unique identity (USR / mangled name), so edges are exact — overloads, ptr->method(), free functions, templates, and virtual dispatch all resolved, with no name-collision merges.

Why it matters for C++. By-name graph tools (tree-sitter) and LSP-based agents (e.g. Serena) degrade on large C++ — templates, macros, overloads — so the agent silently falls back to grep, which merges distinct same-named symbols, counts comments/decls as calls, and misses ptr->method() / virtual dispatch. cppgraph is the graph the compiler actually sees — see COMPARISON.md for measured numbers (grep-and-read vs cppgraph, and vs graphify / Serena).

Install with your AI agent

To install, tell your agent: "Follow instructions to install https://github.com/rakiz/cppgraph" — that points it at this section.

If you are that agent, follow these steps exactly. They are deterministic; the scripts fail loudly with actionable messages.

Two phases. Phase A (machine setup): clone, then run setup with the scip-clang source the user chose. Phase B (index a project): run the index with the scope the user chose. You gather each choice by asking the user (in your own question UI), then run the script non-interactively with those choices as flags.

RULE — ! commands can't be answered, so you ask, then pass flags. A Claude Code ! … run has no interactive stdin: a prompt inside it gets EOF. So the interactive menus are for a human in a real terminal, not for you. Your loop is: (1) ask the user every choice via your question UI — the scip-clang source (Phase A) and the index scope (Phase B) — surfacing the real options; (2) run the script non-interactively with those choices as flags. Decide nothing yourself; the answers are the user's. Give a realistic time estimate with each heavy step: indexing is the long one (minutes to hours, CPU-bound — ~20 min for ~6 000 TUs on a fast 14-core x86, ~4 h for 6 482 TUs on an 8-core Graviton2 (m6g), ~1 h 42 for 4 301 TUs on an 8-core Graviton3 (m7g.2xlarge)); downloading scip-clang is seconds, compiling it (#504) is ~25–60 min (~26 min measured on m7g.2xlarge).

Phase A — set up the machine (once)

  1. Before cloning anything cppgraph-related — even just to read/explore the source — check for an existing per-machine install first: if ${XDG_DATA_HOME:-$HOME/.local/share}/cppgraph/repo exists, use that checkout (read it, git pull it if stale); never clone a second copy (~/cppgraph, a project dir, a dev checkout). Only clone fresh if absent.
  2. Clone into the per-machine tool dir (always this path — not the user's project, not a dev checkout; even "install from a local copy" means clone that copy into this path):
    git clone https://github.com/rakiz/cppgraph "${XDG_DATA_HOME:-$HOME/.local/share}/cppgraph/repo"
  3. Get the valid sources from the tool (don't guess the platform):
    ! ~/.local/share/cppgraph/repo/scripts/setup.sh --list-sources
    
    It prints this machine's OS/arch and the sources that actually apply (e.g. no download on ARM-Linux). Ask the user to pick from exactly those — costs: download-patched ~1 min (patched binary, no Docker; macOS arm64 + Linux aarch64 today), download ~1 min (stock, unpatched), build (patched) ~25–60 min Docker, emulate (slower indexing). Windows → WSL2; Intel Mac → only emulate.
  4. Run it with their choice (! runs it non-interactively; --scip-source is what makes that work — without it a piped run stops with ACTION NEEDED):
    ! ~/.local/share/cppgraph/repo/scripts/setup.sh --scip-source <download-patched|download|build|emulate>
    
    setup.sh is the sole entry point — it creates the venv, obtains scip-clang, registers the MCP server, installs the bundled skill and /cppgraph slash command into detected agent tools, and (in a real terminal) offers to index the current project. Never tell the user to run cppgraph setup / .venv/bin/cppgraph … / anything under a dev checkout: none of that exists on a fresh machine until setup.sh has run.
  5. Never pip install cppgraph (or add it as a dependency) into a target project's own venv/virtualenv. It's a single per-machine tool with one dedicated venv (~/.local/share/cppgraph/repo/.venv); invoke its binaries from there directly (optionally alias cppgraph to that binary), never re-install it into another project's environment — that creates a second divergent install with its own dependency versions that can conflict with the committed generated bindings (observed: a project venv's pinned protobuf vs cppgraph's). Setup also links the binary into ~/.local/bin automatically (skipped when a cppgraph is already on PATH), so a bare cppgraph works after setup; the shell alias below remains an alternative (e.g. for older installs predating the link). Convenience alias for the user's shell rc (~/.bashrc / ~/.zshrc):
    alias cppgraph="${XDG_DATA_HOME:-$HOME/.local/share}/cppgraph/repo/.venv/bin/cppgraph"

Phase B — index a project (per-project; the heavy step)

You drive Phase B; the user only picks from choices you offer — never typing a scope, path, or flag, and never seeing these commands. The conversation: (a) a yes/no — "Index this project now? (~estimate)"; if yes, (b) a selectable scope choice (whole tree / a subtree) + a yes/no on tests; then (c) you run it and report progress; (d) when done, they open a new Claude Code session.

  1. Get the scope options: run ~/.local/share/cppgraph/repo/scripts/index.sh --plan-json (a bare cppgraph works when setup linked it into ~/.local/bin; otherwise use the full path ~/.local/share/cppgraph/repo/.venv/bin/cppgraph, the fallback on older installs) from the project dir; it auto-locates the compile_commands.json — if it reports none, generate one, see AGENTS.md → "Fallback", after the user's OK. It returns the compdb breakdown, the questions (subtree / tests / attribution), and artifacts (whether a .scip/.graph.db already exists).
  2. Offer the choices — a single-select of whole tree + each subtree (with TU counts) and a yes/no on tests; never ask the user to type a subdirectory or flag. If artifacts shows the project is already indexed, ask whether to keep it (default) or rebuild.
  3. Run it with their answers (non-interactive, so ! works):
    ! ~/.local/share/cppgraph/repo/scripts/index.sh <compdb> -y --filter <sub> [--no-tests] [--attributed-refs] --run
    
    Non-destructive by default: an existing .scip/.graph.db is kept, not overwritten — only what's missing is built. Add --from-scratch only if the user asked to rebuild. The chosen scope is recorded and reused by later updates. (A human in a real terminal can instead run scripts/index.sh with no flags for the interactive wizard.) Then tell the user to open a new Claude Code session from their project directory (that's how the server finds this project's graph) and ask "what calls X?", "impact of changing Y?", "show the dependency graph of Z".

Uninstalling: run ! ~/.local/share/cppgraph/repo/scripts/uninstall.sh (it asks per item). There is no uninstall flag on setup.sh, and you must never improvise rm -rf on ~/.local/share/cppgraph or a project's .cppgraph/ — the latter holds a .scip index that can take hours to rebuild; uninstall.sh keeps project data by default and warns before touching anything precious.

Humans: the same flow, step by step, is in QUICKSTART.md. Every CLI subcommand, grouped by purpose with one example each, is in CLI_REFERENCE.md.

Status

Early, but functional end-to-end: build, query, incremental update, an MCP server for LLMs, and visualization. The tool is general — point it at any C++ project's compile_commands.json.

Pipeline

scip-clang        compile_commands.json  →  <name>.compdb.json  →  <name>.scip  →  <name>.graph.db
(indexer binary)  (target's build)          (filtered subset)      (SCIP protobuf)  (SQLite: query/MCP/viz)

Five stages get you from nothing to a queryable graph. Only the last three are cppgraph's; the timings are rough and dominated by your machine (the indexer runs the C++ front-end once per translation unit, so it scales with cores/CPU).

# Stage Produces Rough time
1 Get scip-clang — copy the prebuilt binary, or compile from source the per-machine indexer binary ~1 min (copy) … ~25–60 min (compile, #504; ~26 min measured on m7g)
2 Get compile_commands.json — from the target's build system, if not already present compile_commands.json ~3 min (warm) … ~15 min (cold) — the target's build, not cppgraph
3 Filter what to index — scope to a subtree, optionally drop tests <name>.compdb.json seconds
4 Index — scip-clang runs the C++ front-end per TU → SCIP <name>.scip the big one: ~20 min → several hours
5 Build the store — parse SCIP, intern symbols + edges into SQLite <name>.graph.db ~30 s (filtered src/mongo) … ~3.5 min (full repo + refs)

Then you're ready to query (see the gains below).

Notes:

  • Stage 1 is once per machine. macOS arm64 and Linux aarch64 get a native prebuilt patched binary (download-patched, ~1 min); Linux x86_64 gets the native stock prebuilt (download, ~1 min, unpatched — #504 there means a local build). All three index natively — no Docker. Docker is only for the build source (a #504 compile, ~25–60 min, Linux-only) or for hosts with no native binary at all (Intel Mac, Windows) via emulate. "#504 build is Linux-only" ≠ "scip-clang needs Docker on macOS".
  • Stage 4 is the variable one. Measured extremes: ~20 min on a fast 14-core x86 for ~6 000 TUs; ~4 h for 6 482 TUs on an 8-core AWS Graviton2 (m6g.2xlarge — older Neoverse-N1 cores are slow at this); ~1 h 42 for 4 301 TUs on an 8-core Graviton3 (m7g.2xlarge — ~1.4 s/TU vs ~2.2 on Graviton2). the index wizard prints a per-machine estimate right before it starts. Fewer/older cores → proportionally longer; scope to a subtree (stage 3) to cut it down.
  • enclosing_range (#504) is a choice at indexing, with consequences: it enables exact caller attribution and the symbol-granularity usage view (--attributed-refs), but needs the compiled #504 binary and makes the .scip and store larger. Without it you still get an exact graph, just with file-granularity usage. You can add it later without re-indexing (enrich-refs) — but note enrich-refs re-parses the .scip and rebuilds the graph with references, so it costs a full store build (mongo: ~3.5 min, ~9 GB RAM) — the top of stage 5's range, not its filtered ~30 s floor.

Two components: the builder (scip-clang, external compiled binary) does the expensive, perf-critical C++ parsing, crash-isolated per TU; cppgraph (this repo, Python) is the glue + graph — parses SCIP, builds the store, serves queries, exposes the MCP server, and exports a graph.json for visualization.

Does it actually beat by-name tools?

Yes, measurably. On a real case (two distinct methods sharing the name makeResumeToken), a tree-sitter tool drops the real call edges and collapses 431 unrelated Value sites onto one node; cppgraph returns the correct, separated caller sets. Full write-up with numbers and reproduction steps: COMPARISON.md (cppgraph vs graphify vs Serena/LSP, on a large C++ codebase).

Why not just grep?

Your AI assistant already answers "what calls X?" with grep — but on a large C++ codebase that's wrong (matches by name: merges distinct symbols, includes comments/decls, misses ptr->method() / virtual dispatch / templates) and token-expensive (noisy output, then whole-file reads to disambiguate, all through the model's context).

grep's raw dump is only a floor: it can't tell a call from a decl, a comment, or a different same-named symbol, so to answer correctly it must read around every hit — and that's where the cost explodes. Measured on MongoDB, "who calls the method makeResumeToken?": grep dumps ~6,600 tokens (98% noise, 4 symbols merged) and needs ~110,000 to read-and-disambiguate; cppgraph find

  • who_calls ingests ~400, exact — 272× leaner. On a rare, uniquely named symbol (grep's best case) cppgraph still wins ~8× once grep reads to verify; on a hot type like OperationContext reading every hit would take ~1,900× more than cppgraph — a theoretical number nobody ingests, so in practice grep truncates and answers from a partial view, silently missing call sites, where cppgraph stays at ~6k tokens, exact and complete. The trade-off is a one-time index (minutes), amortized over every later query. Full spectrum, noise ratios, and the token-lean defaults: COMPARISON.md (reproduce with scripts/measure_tokens.py --suite).

Same idea versus an LSP-driven agent (e.g. Serena): the questions are the same, but cppgraph is a pre-indexed graph the agent queries in one call, not a live language server it has to drive file-by-file — which is what keeps it token-lean and complete on large C++. Numbers: COMPARISON.md.

Documentation

Doc What's in it
AGENTS.md Working instructions, principles, guardrails — read first
DESIGN.md Architecture, edge model, the call-attribution heuristic + its known limitation
COMPARISON.md Measured comparison vs graphify and Serena on a real design question
INSTALL.md Setting up a new machine (scip-clang, protoc, the venv)
viz/README.md The bundled graph viewer + cppgraph export graph.json format
CHANGELOG.md What's been built so far
TODO.md Open tasks

Non-goals

  • Re-implementing a C++ parser. We consume a compiler index.
  • Being a linter or a refactoring engine. This is about understanding structure (who calls what, impact/blast-radius, inheritance) for humans and LLMs.

Visualize a neighbourhood, call path, or cycle

Secondary to the query tools: cppgraph export <symbol> --depth 2 --out graph.json (run from the indexed project — graph auto-discovered, <symbol> a plain name or exact SCIP string), or the visualize MCP tool, writes a bounded neighbourhood and opens it in a self-contained, offline viewer (viz/cppgraph-viz.html). The graph.json is graphify-compatible. export/view (and the MCP visualize) also take --mode path --dst <target> to draw how two symbols connect — the shortest calls chain by default, --expand-paths for the full corridor of every route between them — or --mode cycle to draw the multi-member call cycle containing a symbol. boundary-violations --out (or the MCP visualize_boundary_violations) draws the edges that cross your declared layering rules. Details: viz/README.md.

Where is this type used?

The call graph can't answer it — a plain struct has no callers. cppgraph keeps an exact reference index (on by default) so cppgraph references <type> (and the find_references tool) lists every use site, and export --mode usage draws a usage graph.

By default that graph is at file granularity ("used somewhere in these files"). Built with a scip-clang that emits enclosing_range (a source build carrying PR #504), you can upgrade it to symbol granularity — "used by these functions" — with cppgraph build --attributed-refs, or add it to an existing graph with cppgraph enrich-refs.

Worth it when you want symbol-level usage — at a cost. Attribution stores one extra symbol id per reference, so the graph grows. Enable it when "which functions use this type?" matters; otherwise the default file granularity is already exact and leaner. cppgraph status tells you which granularity a graph has and how to upgrade.

License

MIT. The bundled viewer vendors vis-network (MIT / Apache-2.0); see viz/README.md.

About

Compiler-exact C++ code intelligence for AI agents (MCP + CLI): who-calls, references, class hierarchy, change-impact — from a pre-indexed SCIP graph, so the agent stops falling back to grep.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages