Skip to content

Latest commit

 

History

35 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Columbus, a compass-carrying explorer charting a map of code

Columbus

Navigate your codebase. Bring back the context you need.

한국어 · Download 1.0.0 · Benchmarks · Installation · Validation

Real spring-core evaluation: three navigation fixes, measured indexing costs, and unresolved graph limits.

Columbus gives coding agents a local map of a repository. Find an entry point, follow its relationships, and read verified source in a bounded response. A compass, a map, and an explorer are the project's visual identity; the evidence still comes from your code.

66 language detection profiles · Incremental SQLite index · 4 graph formats · CLI, skill & optional MCP · MIT

No LLM or API key is needed to index a repository. Columbus does not build or execute the project. Python, Java, and Kotlin use ASTs; 43 profiles use declaration heuristics; other UTF-8 text stays searchable as files. Language coverage.

Source development now includes parent-first AST JSONL trees (columbus tree --label NAME), portable agent-skill installation, and explicit pre-commit refresh integration. AST processing and repository workflow. These additions require a revision newer than the v1.0.0 release assets.

Set sail in three commands

Requires Python 3.11+ and uv. Run the last two commands inside the project you want to explore.

uv tool install https://github.com/ch4570/columbus/releases/download/v1.0.0/columbus-1.0.0-py3-none-any.whl
columbus explore
columbus explore checkout

explore shows a small repository map. Add a symbol or keyword to retrieve relevant source, with a default limit of 2,000 estimated tokens. Both commands synchronize the local index automatically. Use --repo /path/to/project from another directory. Running columbus alone shows a short guide.

Already use pipx? Replace uv tool install with pipx install. If the executable is not on PATH, follow the tool's hint (uv tool update-shell or pipx ensurepath). The pinned GitHub wheel is the distribution; these instructions do not assume a PyPI package named Columbus.

Only have Python?

Download the standalone installer, then run it:

curl -fL https://github.com/ch4570/columbus/releases/download/v1.0.0/get-columbus.py -o get-columbus.py
python3 get-columbus.py

On Windows, download get-columbus.py and run py -3 get-columbus.py. The installer verifies the wheel's SHA-256, creates a dedicated user environment, checks its health, and prints the command location. It preserves unrelated installations and shell configuration. Offline installation and troubleshooting.

Used the RepoAtlas preview? Columbus has a new command, skill, configuration name, and cache. Migration instructions.

Choose your route

Your task Command
Understand an unfamiliar repository columbus explore
Locate a feature or exact symbol ID columbus search checkout --format text --limit 5
Inspect declarations before source columbus context checkout --mode signatures --format text
Read the relevant source columbus explore checkout
Follow calls, imports, or inheritance columbus neighbors EXACT_SYMBOL_ID --kinds calls imports inherits --format text
Inspect incoming relationships columbus impact EXACT_SYMBOL_ID --format text
Share an interactive graph columbus graph --level file --format html --output graph.html

Use IDs returned by search. impact is a bounded graph traversal, not proof of every runtime effect. Existing automation commands retain their JSON default; explore defaults to readable text.

Let an agent travel light

For a known name or file, start with a short search or a direct source read. Use a map when the structure is unclear, then inspect relationships only when they help answer the question. Stop when the evidence is sufficient.

columbus init
columbus explore checkout --session checkout
columbus explore calculateTotal --session checkout
columbus stats checkout

init installs .agents/skills/columbus/SKILL.md in the current project. A named session keeps a source receipt and local query measurements under .columbus/sessions/checkout/. Later snippet queries omit source ranges already delivered and continue unread portions. stats reports query count, response bytes, and source bytes.

A receipt records delivered source; it does not restore a model's lost context. Use a new session for a new task, another agent, or after context loss. Telemetry is opt-in and records metadata without queries or source bodies. Add .columbus/ to your project's ignore rules.

Read .agents/skills/columbus/SKILL.md. Start with the smallest useful
search or source range. Use a bounded map only when the structure is
unclear. Reuse one session for this task's source queries and stop
when the evidence answers the question. Verify source before editing,
then run the relevant tests and synchronize the index.

Agent workflow, receipt limits, and MCP setup.

Benchmarks, with the rough seas included

Smaller tool responses are measurable. Lower model-token usage is a separate question. The graphs are generated from recorded JSON; their data and plotting script ship with the source release.

Columbus 1.0: response delivery

Measured Columbus response bytes, comparing map formats and repeated context with receipts

This model-free benchmark measures actual UTF-8 CLI output on the included polyglot fixture. The map comparison verifies the same ordered symbol IDs. The repeated-context comparison uses the same query three times; receipts can deliver previously unread ranges before returning only status metadata. Bytes and the CLI's byte-based token estimate are not model usage or billed cost.

Comparison Before After Change
Same 28 map items: JSON → text 6,639 bytes 3,534 bytes 46.8% smaller
Three context calls: no receipt → receipt 9,612 bytes 4,474 bytes 53.5% smaller
Source across receipt calls 1 → 2 → 3 1,423 bytes 188 → 0 bytes No repeated source after exhaustion

Current measurements, fixture hashes, methodology, and reproduction commands.

Revisit a repository without reparsing it all

Three-run incremental index measurements on a synthetic 1,001-file Python repository

On a deterministic 1,001-file Python fixture, initial indexing parsed 1,001 files, an unchanged refresh parsed and hashed 0, and editing one file reparsed 1. The chart shows the median and min–max range of three runs on macOS arm64. An unchanged scan still checks file metadata; timing is machine-specific and OS caches were not cleared. Fixture, raw runs, and reproduction.

Real agent exploration: the earlier pilot

Historical actual Codex cumulative input tokens, including the case where both answers pass and tool use increases input

These runs used the RepoAtlas 0.4.0 engine, before the Columbus rename. Three fixed questions on the same real source snapshot; ordinary rg and bounded reads versus those tools plus the graph. Each condition ran once with Codex CLI 0.153.4, requested model gpt-5.6-sol, and effort xhigh. A prebuilt graph was supplied; its 1.105-second construction was measured separately.

Question Citation check, ordinary → graph Actual cumulative input, ordinary → graph Change
Graph export protection Pass → Pass 79,358 → 106,334 +34.0%
Configuration invalidation Fail → Pass 157,482 → 138,864 −11.8%
Managed installation conflicts Fail → Pass 70,775 → 88,411 +24.9%

The two baseline citation failures used ellipses instead of contiguous source quotations; they do not prove semantic misunderstanding and are excluded from equal-quality savings claims. The only pair passing both citation checks used 34.0% more input. General model-token savings were not established. In the installation case, command output fell by 49.0% while input increased.

Runtime usage includes startup instructions, repeated context, tool exchanges, and cache behavior. Cached input is already part of input. All six trials and the entire excluded initial empty-index cohort are preserved. This is a small read-only pilot, not a Columbus 1.0 model A/B result. Prompts, answers, usage, and grading · Full historical report.

Bring back a graph

columbus graph --format html --level file --output graph.html
columbus graph --format mermaid --level file --kinds imports calls --output dependencies.mmd
columbus graph --format graphml --language typescript --path 'src/*' --output frontend.graphml
Format Best for
HTML Exploring a self-contained graph offline
Mermaid Editable diagrams in documentation
GraphML Importing into external graph tools
JSON Scripts and structured analysis

Filter by language, path, relationship, direction, and focus symbol. Graphs expose confidence and truncation; unknown languages still produce file nodes. Existing output files are preserved. Graph options.

Under the compass

Discovery finds eligible files. AST and heuristic analyzers add source-backed facts to a local SQLite graph. Retrieval selects declarations or verified source spans within a response budget. CLI, exporters, skill, and MCP share this engine.

The index lives at .columbus/index-v1.sqlite. Source shown to an agent remains untrusted repository content. Architecture and data boundaries.

Guide Contents
Install Wheel, Python-only installer, offline use, updates, preview migration
CLI usage · 한국어 사용법 Queries, formats, budgets, freshness
Agent integration Skill, sessions, receipts, telemetry, MCP
Benchmarks Graphs, data, methodology, reproduction
Validation Executed checks and platform limits
Contributing · Changelog Development and release history

Support · Security · MIT license · Character artwork and provenance

About

Navigate your codebase. Local code graphs, focused context, and reproducible agent benchmarks.

Topics

Resources

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages