Skip to content

Latest commit

 

History

History
155 lines (127 loc) · 8.38 KB

File metadata and controls

155 lines (127 loc) · 8.38 KB

Software evaluation

crie answers two different questions. The original one:

How healthy is this repository? (git history, ownership, coupling, dependency structure — see scoring-model.md)

And a second, broader one, covered here:

What does the software itself tell us about maintainability, readability, debuggability, reliability, architecture, testing, and performance readiness?

Both run as part of crie analyze by default. They're not separate tools bolted together — the software-evaluation domains reuse the same Evidence vocabulary, the same percentile normalization, and (for several domains) the exact same raw signals the repository-risk scoring already collected.

The seven domains

Domain What it measures Primary source
Maintainability Change cost: complexity, coupling, size, churn, technical-debt comments Bridged from existing evidence + maintainability.py
Readability Naming, function length, nesting, duplicated code shape readability.py
Debuggability Silent exception handling, lost error context, missing observability debuggability.py
Reliability Unmanaged resources, missing timeouts, mutable defaults, silent failures reliability.py
Architecture Fan-in/fan-out, instability, import cycles, "god modules" Bridged from dependency_graph.py + architecture.py
Testing Whether a plausible test file exists per source file testing.py
Performance Readiness Static patterns only — nested loops, accumulator-in-loop, blocking calls in async code performance_readiness.py

Each domain rolls up to one repo-level score (0–100, LOC-weighted across files) plus a confidence figure and a list of the worst-offending files. A domain score is a different kind of number from the risk dimensions (Overall Risk, Regression Risk, ...): a dimension is a per-file weighted-sum risk score; a domain is "how healthy is this concern across the whole codebase."

Why some domains reuse existing evidence (the "bridge")

Maintainability and Architecture would be pure duplication if built from scratch — complexity, coupling, and dependency-graph evidence already measure most of what those domains care about. crie/scoring/bridge.py re-projects that existing evidence onto the domain axis via a fixed table:

CATEGORY_DOMAIN_BRIDGE = {
    Category.COMPLEXITY: [Domain.MAINTAINABILITY, Domain.READABILITY],
    Category.COUPLING: [Domain.MAINTAINABILITY, Domain.ARCHITECTURE],
    Category.DEPENDENCY: [Domain.ARCHITECTURE],
    Category.SIZE: [Domain.MAINTAINABILITY],
    Category.CHURN: [Domain.MAINTAINABILITY],
    Category.TEST_COVERAGE_GAP: [Domain.TESTING],
}

This runs once per analysis, before normalization, so bridged evidence is scored consistently with everything else. Debuggability, Reliability, Testing, and Performance Readiness have no equivalent existing signal — those are genuinely new analyzers.

Confidence and "insufficient evidence"

A domain score is only shown as a number when enough files actually produced evidence for it (evidence coverage ≥ 35%, see MIN_CONFIDENCE_FOR_NUMERIC_SCORE in crie/scoring/software_domains.py). Below that, reports show "insufficient evidence" instead of a number — a domain that would otherwise be, say, "62" from two files out of two hundred is not meaningfully "62% healthy," and pretending otherwise would be false precision. The raw score and confidence are still both present in JSON output, so nothing is hidden from tooling that wants to use it anyway.

Project profile detection

Before scoring, crie classifies the repository's type — library, cli_application, web_frontend, backend_service, desktop_application, mobile_application, framework, developer_tool, systems_software, monorepo, educational_example, or unknown — using deterministic signals already on disk (crie/profile/detector.py): pyproject.toml scripts/dependencies, package.json bin/framework dependencies, Cargo.toml [[bin]]/[lib], go.mod + import scanning, Dockerfile EXPOSE, and workspace files for monorepos. Candidate profiles accumulate votes from these signals; a tie or no signal at all returns unknown — detection failing never silently skips evaluation. Override it explicitly with --profile NAME.

All seven domains apply to every profile. They're universal engineering concerns — a CLI tool's testing and architecture matter as much as a backend service's. Profile instead affects a handful of specific finding types (e.g. print_debugging is suppressed for a CLI application's own entry point, where printing is the correct behavior, not leftover debug output) and which domains a report visually leads with under --focus.

Focus presets

--focus NAME (debugging, performance, readability, maintainability, architecture, reliability, testing) changes which analyzers actually run — not just which results are shown — via crie/config/focus.py's data-driven FOCUS_PRESETS. It reuses the same analyzer-restriction mechanism --analyzers uses; --analyzers, if also given, wins as the more specific override.

crie analyze . --focus debugging

See configuration.md for the full flag reference and cli-reference.md for exact preset contents.

Cross-domain correlation and hotspots

crie/scoring/correlation.py runs after both scoring passes complete and produces two things, neither of which feeds back into any score (correlation is a reporting synthesis, never a re-scoring pass — see that module's docstring for why that separation matters):

  • Hotspots — files where multiple concerns converge, one of seven kinds (maintenance_hotspot, debugging_hotspot, performance_risk_hotspot, testing_gap, architecture_hotspot, change_risk_hotspot, knowledge_hotspot), each with the underlying evidence bullets rather than a single unexplained "badness" number.
  • Correlation findings — one narrative sentence per top-risk file, built from a template (deterministic, not AI) that names which signals combined to produce that risk, e.g. "app/legacy_processor.py is structurally complex and lacks a detected test file — this combination of signals raises its practical risk beyond what any single dimension shows."

This is also where "risky module with no detected tests" is computed — it needs both testing.py's per-file test-presence evidence and churn/complexity evidence from other analyzers, which don't exist at the same moment any single analyzer runs (analyzers run independently; evidence is only assembled afterward).

Every domain score, hotspot, and correlation finding also carries a Low/Moderate/High/Critical severity label alongside its raw number — see scoring-model.md.

Known limitations

  • Domain-specific checks (debuggability, reliability, testing-presence, performance-readiness) are Python-only in this release. JS/TS files still get complexity and dependency-graph coverage via the bridge, just not new checks in those four domains — see architecture.md for why language support is scoped this way.
  • No general concurrency/data-race detection — too unreliable to do statically without excessive false positives. Only one narrow, defensible check exists (blocking calls inside async def bodies).
  • No architectural layer-violation detection. Flagging "module A shouldn't import module B" requires knowing the intended layering, which isn't something this release infers automatically or accepts as configuration — inferring it from structure alone is too false-positive-prone, and a config-driven version is a reasonable follow-up rather than something built here. What is implemented (fan-in/fan-out, instability, cycles, "god modules") stands on its own.
  • Findings explain what was measured precisely (exact counts, percentiles, specific line-level patterns); hotspots and correlation findings additionally carry a plain-language explanation of why it specifically matters here, for a non-engineer audience — deterministic by default, optionally sharpened by AI via crie analyze --explain-ai — see custom-inquiries.md.
  • See scoring-model.md for the equivalent caveats on the original repository-risk model.