crie answers two different questions. The original one:
How healthy is this repository? (git history, ownership, coupling, dependency structure — see scoring-model.md)
And a second, broader one, covered here:
What does the software itself tell us about maintainability, readability, debuggability, reliability, architecture, testing, and performance readiness?
Both run as part of crie analyze by default. They're not separate tools bolted
together — the software-evaluation domains reuse the same Evidence vocabulary,
the same percentile normalization, and (for several domains) the exact same raw
signals the repository-risk scoring already collected.
| Domain | What it measures | Primary source |
|---|---|---|
| Maintainability | Change cost: complexity, coupling, size, churn, technical-debt comments | Bridged from existing evidence + maintainability.py |
| Readability | Naming, function length, nesting, duplicated code shape | readability.py |
| Debuggability | Silent exception handling, lost error context, missing observability | debuggability.py |
| Reliability | Unmanaged resources, missing timeouts, mutable defaults, silent failures | reliability.py |
| Architecture | Fan-in/fan-out, instability, import cycles, "god modules" | Bridged from dependency_graph.py + architecture.py |
| Testing | Whether a plausible test file exists per source file | testing.py |
| Performance Readiness | Static patterns only — nested loops, accumulator-in-loop, blocking calls in async code | performance_readiness.py |
Each domain rolls up to one repo-level score (0–100, LOC-weighted across files) plus a confidence figure and a list of the worst-offending files. A domain score is a different kind of number from the risk dimensions (Overall Risk, Regression Risk, ...): a dimension is a per-file weighted-sum risk score; a domain is "how healthy is this concern across the whole codebase."
Maintainability and Architecture would be pure duplication if built from scratch —
complexity, coupling, and dependency-graph evidence already measure most of what
those domains care about. crie/scoring/bridge.py re-projects that existing
evidence onto the domain axis via a fixed table:
CATEGORY_DOMAIN_BRIDGE = {
Category.COMPLEXITY: [Domain.MAINTAINABILITY, Domain.READABILITY],
Category.COUPLING: [Domain.MAINTAINABILITY, Domain.ARCHITECTURE],
Category.DEPENDENCY: [Domain.ARCHITECTURE],
Category.SIZE: [Domain.MAINTAINABILITY],
Category.CHURN: [Domain.MAINTAINABILITY],
Category.TEST_COVERAGE_GAP: [Domain.TESTING],
}This runs once per analysis, before normalization, so bridged evidence is scored consistently with everything else. Debuggability, Reliability, Testing, and Performance Readiness have no equivalent existing signal — those are genuinely new analyzers.
A domain score is only shown as a number when enough files actually produced
evidence for it (evidence coverage ≥ 35%, see MIN_CONFIDENCE_FOR_NUMERIC_SCORE in
crie/scoring/software_domains.py). Below that, reports show "insufficient
evidence" instead of a number — a domain that would otherwise be, say, "62" from
two files out of two hundred is not meaningfully "62% healthy," and pretending
otherwise would be false precision. The raw score and confidence are still both
present in JSON output, so nothing is hidden from tooling that wants to use it
anyway.
Before scoring, crie classifies the repository's type — library,
cli_application, web_frontend, backend_service, desktop_application,
mobile_application, framework, developer_tool, systems_software,
monorepo, educational_example, or unknown — using deterministic signals
already on disk (crie/profile/detector.py): pyproject.toml scripts/dependencies,
package.json bin/framework dependencies, Cargo.toml [[bin]]/[lib],
go.mod + import scanning, Dockerfile EXPOSE, and workspace files for
monorepos. Candidate profiles accumulate votes from these signals; a tie or no
signal at all returns unknown — detection failing never silently skips
evaluation. Override it explicitly with --profile NAME.
All seven domains apply to every profile. They're universal engineering
concerns — a CLI tool's testing and architecture matter as much as a backend
service's. Profile instead affects a handful of specific finding types (e.g.
print_debugging is suppressed for a CLI application's own entry point, where
printing is the correct behavior, not leftover debug output) and which domains a
report visually leads with under --focus.
--focus NAME (debugging, performance, readability, maintainability,
architecture, reliability, testing) changes which analyzers actually run —
not just which results are shown — via crie/config/focus.py's data-driven
FOCUS_PRESETS. It reuses the same analyzer-restriction mechanism --analyzers
uses; --analyzers, if also given, wins as the more specific override.
crie analyze . --focus debuggingSee configuration.md for the full flag reference and cli-reference.md for exact preset contents.
crie/scoring/correlation.py runs after both scoring passes complete and produces
two things, neither of which feeds back into any score (correlation is a reporting
synthesis, never a re-scoring pass — see that module's docstring for why that
separation matters):
- Hotspots — files where multiple concerns converge, one of seven kinds
(
maintenance_hotspot,debugging_hotspot,performance_risk_hotspot,testing_gap,architecture_hotspot,change_risk_hotspot,knowledge_hotspot), each with the underlying evidence bullets rather than a single unexplained "badness" number. - Correlation findings — one narrative sentence per top-risk file, built from a template (deterministic, not AI) that names which signals combined to produce that risk, e.g. "app/legacy_processor.py is structurally complex and lacks a detected test file — this combination of signals raises its practical risk beyond what any single dimension shows."
This is also where "risky module with no detected tests" is computed — it needs
both testing.py's per-file test-presence evidence and churn/complexity evidence
from other analyzers, which don't exist at the same moment any single analyzer
runs (analyzers run independently; evidence is only assembled afterward).
Every domain score, hotspot, and correlation finding also carries a Low/Moderate/High/Critical severity label alongside its raw number — see scoring-model.md.
- Domain-specific checks (debuggability, reliability, testing-presence, performance-readiness) are Python-only in this release. JS/TS files still get complexity and dependency-graph coverage via the bridge, just not new checks in those four domains — see architecture.md for why language support is scoped this way.
- No general concurrency/data-race detection — too unreliable to do statically
without excessive false positives. Only one narrow, defensible check exists
(blocking calls inside
async defbodies). - No architectural layer-violation detection. Flagging "module A shouldn't import module B" requires knowing the intended layering, which isn't something this release infers automatically or accepts as configuration — inferring it from structure alone is too false-positive-prone, and a config-driven version is a reasonable follow-up rather than something built here. What is implemented (fan-in/fan-out, instability, cycles, "god modules") stands on its own.
- Findings explain what was measured precisely (exact counts, percentiles,
specific line-level patterns); hotspots and correlation findings additionally
carry a plain-language explanation of why it specifically matters here, for
a non-engineer audience — deterministic by default, optionally sharpened by AI
via
crie analyze --explain-ai— see custom-inquiries.md. - See scoring-model.md for the equivalent caveats on the original repository-risk model.