Skip to content

Latest commit

 

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Code Risk Intelligence Engine

crie is a command-line (CLI) tool — no server, no web UI — that acts as a universal engineering evaluator for software projects. Point it at a local repository from your terminal and it combines source-code analysis, architecture and dependency intelligence, test and reliability assessment, git-history intelligence, and evidence-driven investigation of arbitrary engineering questions — not a single "AI reads your repo and vibes" tool, and not just a git-history scanner. Every score traces back to deterministic evidence you can inspect; AI, where used at all, only ever synthesizes evidence that was already collected, never invents it.

It answers two questions:

How healthy is this repository? — git churn, ownership concentration (bus factor), dependency structure, co-change coupling: see docs/scoring-model.md.

What does the software itself tell us about maintainability, readability, debuggability, reliability, architecture, testing, and performance readiness? — see docs/software-evaluation.md.

This is not a bug finder. It doesn't look for defects in the code you have today — it estimates where tomorrow's defects and maintenance cost are most likely to accumulate, and lets you interrogate why.

Why not just use a single complexity score?

A file can be low-complexity and still be the riskiest thing in your codebase — if it's a single-owner, high-churn, heavily-depended-on file with no tests and silent exception handling. Complexity alone misses all of that. This tool's scoring model combines independent signals instead of picking one, and every weight in the combination is documented, configurable data, not a hardcoded magic number — see docs/scoring-model.md and docs/software-evaluation.md.

Features

Repository intelligence — git history (churn, edit frequency, file age, ownership/bus factor), intra-repo dependency graph (fan-in, betweenness centrality, import cycles), co-change coupling, optional test coverage gap analysis. Five explainable, config-driven risk dimensions (Overall Risk, Regression Risk, Maintenance Difficulty, Refactoring Priority, Dependency Criticality) with full contributing-factor breakdowns.

Software intelligence — seven domains (Maintainability, Readability, Debuggability, Reliability, Architecture, Testing, Performance Readiness), each a confidence-gated repo-level score with named worst offenders. Automatic project-type detection (library, CLI, backend service, web frontend, ...) and --focus presets that change which analyzers run, not just what's displayed.

Cross-domain correlation — hotspots where multiple concerns converge (maintenance, debugging, performance-risk, testing-gap, architecture, change-risk, knowledge-concentration), and narrative "why is this risky" assessments combining repository and software evidence — never a single unexplained "badness" score.

Custom engineering inquiries — crie ask runs a deterministic question→investigation-plan→evidence pipeline for questions like "How difficult would replacing SQLite with PostgreSQL be?" or "What parts are most dangerous to modify?", optionally synthesized into prose by an AI provider — the model only ever sees the evidence already collected, never your source code, and the whole pipeline still works with AI disabled. See docs/custom-inquiries.md.

Impact analysis — crie impact FILE estimates change radius (direct/indirect dependents, related tests, config references) from the same dependency and co-change data. See docs/impact-analysis.md.

Extensible and honest about uncertainty — a plugin architecture (third-party analyzers/reporters via standard Python entry points), every finding tagged with its provenance (deterministic, historical, semantic, ...), and a confidence-gated fallback to "insufficient evidence" instead of false precision wherever the data doesn't support a sharp number.

Requirements

  • Python 3.11+
  • Git (for history-based signals; the tool still runs without it, with reduced evidence and lower confidence scores)
  • Optional: pip install "code-risk-intelligence-engine[js]" for JavaScript/TypeScript complexity analysis (adds tree-sitter)
  • Optional, for AI-assisted synthesis in crie ask (the tool works fully without this — see Enabling AI synthesis below if crie ask isn't giving you prose answers)

Installation

python -m venv .venv
.venv/Scripts/activate   # .venv/bin/activate on macOS/Linux
pip install -e ".[dev,js,ai]"

Usage

crie is invoked from your terminal as crie <command> [arguments] [flags]. The seven commands, at a glance:

Command What it does
crie analyze [PATH] Run the full evaluation (repository + software) and write reports. The main command.
crie ask PATH "QUESTION" Investigate a free-form engineering question with evidence-backed reasoning.
crie impact PATH TARGET Estimate the blast radius of changing one file.
crie report JSON_PATH Re-render a saved report.json in another format, without re-analyzing.
crie list-analyzers List every registered analyzer and which domain(s) it feeds.
crie init [PATH] Scaffold a commented crie.config.yaml.
crie version Print the installed version.

Full flag reference: docs/cli-reference.md. Config file reference: docs/configuration.md.

Everyday examples

Run a full evaluation on the current directory:

crie analyze .

Focus on one concern — this changes which analyzers actually run, not just what's displayed (debugging, performance, readability, maintainability, architecture, reliability, or testing):

crie analyze . --focus debugging

Gate CI on a risk threshold (exits non-zero if breached):

crie analyze . --fail-on-score 0.9

Ask a specific engineering question:

crie ask . "What parts of this application are most dangerous to modify?"

Check what changing one file would affect before you touch it:

crie impact . src/auth/session.py

Scaffold a config file to customize weights, excludes, or output formats:

crie init

Enabling AI synthesis for crie ask

crie ask always works — the deterministic investigation plan and evidence are the real substance of the answer. AI synthesis is an optional layer on top that turns that evidence into prose. If you ran crie ask and just got the plan + evidence with no prose, it's not broken — it's not configured yet. Three steps:

1. Install the extra:

pip install "code-risk-intelligence-engine[ai]"

2. Get an API key from console.anthropic.com/settings/keys (requires an Anthropic account with billing set up — synthesis calls are billed to your account like any API usage).

3. Set it as an environment variable — the syntax differs by shell, and using the wrong one for your shell is the most common reason this doesn't seem to work:

# PowerShell (Windows) — for the current session only:
$env:ANTHROPIC_API_KEY = "sk-ant-..."
# bash/zsh (macOS/Linux, or Git Bash on Windows) — for the current session only:
export ANTHROPIC_API_KEY="sk-ant-..."

To make it persist across terminal sessions instead of just the current one, set it in your shell profile ($PROFILE for PowerShell, ~/.bashrc/~/.zshrc for bash/zsh) or your OS's environment variable settings.

Then verify it's picked up:

crie ask . "Why is this project difficult to debug?"

If it's still not working, crie tells you why rather than just saying "unavailable": a configured-but-failing provider (bad key, no billing, network error) prints AI synthesis failed: <the actual reason>, which is a different message from AI synthesis unavailable (not configured at all) — so you can tell "I haven't set this up" apart from "I set it up wrong." See docs/custom-inquiries.md for the full troubleshooting guide.

Documentation

Tests

pytest
ruff check .
mypy crie

Integration tests run the full pipeline against the synthetic repositories under examples/, which have deliberately crafted git histories and risk patterns. AI-provider tests use a mocked provider — no network calls or API key required to run the suite.

Project status

This is a young project. The scoring model's default weights are justified by general software-engineering research (see docs/scoring-model.md and docs/software-evaluation.md) but haven't been validated against a large corpus of real-world repositories — treat rankings and hotspots as a strong starting point for investigation, not a certified verdict. Contributions, new analyzers, and feedback from running this against real codebases are welcome.

Explaining findings in plain language

Every hotspot and cross-domain correlation finding from crie analyze carries two explanations: a technical one (citing exact metrics/percentiles) and a plain-language one aimed at a non-engineer stakeholder — "here's specifically what's wrong here, said two ways." The plain-language version is deterministic and template-based by default, so crie analyze's reproducibility guarantee is untouched; pass --explain-ai to sharpen it with the same AI-synthesis mechanism crie ask uses (opt-in, gracefully falls back if AI isn't configured). See docs/custom-inquiries.md.

About

CLI that estimates where a codebase's future maintenance cost will accumulate — churn, bus factor, dependency coupling — with every score traceable to inspectable evidence.

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages