Skip to content

Fix hub analysis and bounded trails; audit graph comparison evidence - #340

Merged
forhappy merged 173 commits into
mainfrom
audit/code-graph-intelligence
Sep 28, 2026
Merged

forhappy merged 173 commits into
mainfrom
audit/code-graph-intelligence

Conversation

@forhappy

@forhappy forhappy commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Audit and implementation work toward more reliable code-graph intelligence. Broad superiority over Graphify and actual god-object detection remain unproven. Compass stays at 0.3.30. These implementation changes are ready for review and merge; this PR does not start a release.

  • Correct hub eligibility, deterministic ranking and source identities; expose incident relation counts without treating degree as proof of excessive responsibility.
  • Correct bounded typed trails, legacy paths and directed shortest-path outcomes; preserve work/depth limits, ambiguity and feasible paths.
  • Correct source-backed receiver, route-parent, Java constructor/varargs and Rust indexed-receiver cases; preserve negative/shadowing evidence and invalidate affected disposable caches.
  • Preserve natural-query subjects and operands, community identities, exact neighbor identities and parallel occurrence records. Avoid silently resolving ambiguous targets.
  • Add source-verified declaration/member explanations and exact symbol lookup with explicit source file, line and kind constraints. Preserve source quotas, incomplete results and unverified/stale-source status.
  • Publish replayable real-repository comparisons, native regressions, scorer tests, all known limitations and competitor advantages.

Latest held-out three-language checkpoint (2026-09-28)

Commits a269c1c9 and 3b7db6a2 register a source-selected Anyhow/Rust, MarkItDown/Python, and Vue/TypeScript panel, correct Owner::member suffix lookup, and remove repeated TypeScript receiver work that caused Vue extraction to time out. The registered original Vue build exceeded its 1,800-second limit; after the performance fix, a Vue-only replay completed Compass extraction in 75.996 seconds versus Graphify's 4.246 seconds. This is a development replay, not a frozen full-panel speed comparison. Both tools failed all four Vue text questions because both omit the named baseCompile callback declaration.

Independent source-site review of thirteen selected direct calls found Compass 10/13, Graphify 7/13 exact owner, target and occurrence links: Anyhow 2/2 versus 0/2; MarkItDown 8/8 versus 7/8; Vue 0/3 for both. This selected sample does not estimate whole-graph precision or establish broad superiority. The initial Anyhow callers Chain::new text result was a Compass no_match; after the suffix fix, a post-registration replay returns the exact ErrorImpl::chain call. Original registered results remain unchanged. See the detailed review, registration, and source-site reviewer.

Current-head local verification passed: formatting, workspace lib/bin Clippy and tests, 114 TypeScript evidence tests, 44 CLI code-query tests, product boundary, and the complete fixtures-only code-graph gate including release-mode React qualification. GitHub CI for this head is still running. The previous head's Windows ARM and x86 native jobs passed after the snapshot-test fix. The remaining Vue callback-owner gap and need for a fresh full twelve-question panel are documented for a separate optimization session.

Detailed evidence and history

The authoritative checkpoint history, acceptance matrix, failed attempts, invalidated claims and remaining work are in the audit. Protocols, per-case judgments, hashes and replay tools are in benchmarks/agent_query. Compatibility and migration effects are documented in COMPATIBILITY.md and MIGRATION.md.

This description replaces an accumulated 65,436-character checkpoint log. Historical results remain in the audit and committed review files; they are not pooled into one accuracy or superiority score. Real-source panels include Cobra/Go, Flask/Python, Gson/Java, Zod/TypeScript and Axum/Rust, plus Chi, Click, jsoup, Redux, WalkDir and a separate fd sample. Graphify 0.9.67 package/environment provenance and frozen inputs are retained externally. Reused development questions are identified as such.

Current comparison checkpoints

These rows measure different protocols and must not be combined:

Evidence Compass Graphify Interpretation
Earlier release development text oracles 50/50 44/50 Five gains concern source excerpts; text recall is not full answer precision
Separate source-first fd text questions 12/12 9/12 Both pass the earlier Graphify-only callees row after source-proven Rust receiver flow; one known repository sample
Source-first Go/Python/Rust selected questions, registered text 9/12 8/12 Frozen text checks on three previously unused repositories
Same questions, source-supported direct-call review 9/12 6/12 Removes two Graphify structural paths that do not show the registered direct calls; neither score establishes broad superiority
Source-reviewed fd call occurrences 16/16 10/16 Selected sample, not population precision
Explicit neighbor identity with one extra resolver allowed 14/14 14/14 Both resolve; Compass payloads are substantially larger
Community collaborator pair co-location 13/15 12/15 Cross-task co-location is 18/60 versus 12/60; neither proves functional clustering quality
Exact source-constrained lookup plus identical bounded source-window policy 14/20 15/20 Graphify retains the overall lead in source evidence; neither authors the responsibility answers
Registered shared-state access sites after Rust/Java corrections 8/20 0/20 Four Rust and four Java sites recovered; 12 remain missing; not a god-object classification score

The source-window comparison uses the same 8,000-byte quota per subject and 35,673 total source bytes per tool. Compass's neighbor responses are much larger (372,829 versus 7,534 full MCP response bytes). Its separate native member-source mode supports 13/20 facts, but Graphify has no equivalent flag; that is not a paired win. The original unscoped-window result, 11/20 versus 15/20, remains recorded.

Graphify retains advantages in several Java/Rust explanation cases and earlier shared-pass token measurements. Scores have been corrected when source review exposed oracle, subject, ambiguity or provenance mistakes. No selected-fact recall or graph-consistency score establishes broad assertion precision.

Source-first Go/Python/Rust checkpoint

Registration commit ab93cd34 fixed 12 selected declaration, caller, callee and one-hop path questions on clean pinned Litestream, FastAPI and Celld sources before either tool ran. Commit 0c106a6e adds the source-and-edge review, verifier and audit account. Raw commands, graphs and responses remain in the external source-first-3lang-01 run.

The registered text score is Compass 9/12, Graphify 8/12. The separate source-supported score is Compass 9/12, Graphify 6/12: Graphify's Litestream and FastAPI path outputs contain the requested names but follow three- and four-hop structural routes instead of the direct source calls. Both graphs contain the FastAPI call, yet the registered short-name requests collide with seven methods; Compass safely reports ambiguity and Graphify's path chooses a wrong source after an ambiguity warning. In a separate unscored control, Compass full names and exact IDs recover the direct call in explain, callees and path; Graphify file-qualified explain finds it, while qualified and exact-ID path requests still select unrelated routes in this example. Celld's selected questions pass for both tools. Two verifier runs are byte-identical. The sample does not prove whole-graph quality, authored explanations, god-object diagnosis, performance or broad superiority. No production code changed in this checkpoint.

FastAPI owner-qualified navigation correction

Commit 2a058e1b makes typed ask/call/trail queries and legacy explain/path resolve a unique owner/member suffix such as APIRoute.get_route_handler against the stored fastapi.routing.APIRoute::get_route_handler. Typed resolution checks the complete bounded exact leaf-name posting; duplicates and truncation stay ambiguous, and a nonexistent owner cannot borrow a same-named method. Exact IDs and full names retain precedence. JSON schemas, graph artifacts and package version 0.3.30 are unchanged.

On the same frozen FastAPI graph, a post-registration replay manifest pins the new binary and five raw outputs. Owner-qualified ask callees and call-path answers now reach the exact line-1232 call without diagnostics; explain identifies the line-1225 declaration and path renders one extracted call hop. MissingRoute.get_route_handler fails. This is development work after the 12-question registration; the original 9/12 versus 8/12 text and 9/12 versus 6/12 source-supported scores stay unchanged. Graphify was not rerun for this patch.

Verification: 44/44 query CLI tests, focused native JSON/store and ambiguity tests, workspace formatting, workspace library/binary Clippy and tests, CLI product tests, and the product boundary gate passed. Unmodified compass-query --all-targets --all-features Clippy stops on a pre-existing redundant closure in tests/explanation_members.rs; rerunning with only that lint allowed passed all targets. Broad query superiority, authored explanation quality and god-object diagnosis remain unproven.

Rust fd collected-builder checkpoint

Commit 68f71422 recovers source-proven Rust method receivers through standard Result<Vec<_>> collection, Ok match bindings, vector loops, and iter().map closures. The native regression also rejects custom aliases and a constructor with a different return type. AST cache semantics advance to 11; package version stays 0.3.30.

The unchanged registered fd suite and pinned source checkout were replayed with freshly rebuilt graphs. Compass now passes 12/12 selected text oracles versus Graphify 9/12; the prior Graphify-only execute_batch callees row now passes for both. Exact Compass call edges reach CommandBuilder::push, finish, and exit_code at source lines 104, 111, and 116. The replay manifest pins source, suite, binary, and graph hashes and lists all six newly source-reviewed calls. Another 213 new reference keys lack independent precision review. On nine shared passes, estimated answer medians are 288 versus 114 tokens. This one known Rust sample does not establish broad superiority, god-object diagnosis, or performance parity.

Formatting, workspace library/binary Clippy and tests, targeted integration suites and Clippy, and fixtures-only code-graph qualification passed. The audit retains the earlier Graphify win and the corrected replay as separate checkpoints.

Java field-access production checkpoint

Production commits 09ca58e6 and f2cacfeb emit bounded lexical Java field-access evidence and preserve exact occurrences, field-only targets and unresolved receivers. Source member types shadow imports/package prefixes; duplicate receiver types cannot be selected by field availability. Compiler/native counterexamples caught both defects before publication; the initial capture is explicitly superseded. AST cache semantics advance from 8 to 9; rebuild graphs for these facts. Package version remains 0.3.30.

The unchanged five-repository comparison now supports 8/20 registered accesses for Compass versus 0/20 for Graphify, up from 4/20 after the Rust correction. All four selected Java sites are recovered. Chi, Click, Redux and WalkDir graphs are byte-identical; jsoup retains all 6,116 nodes and 21,110 old edges and adds 3,896 field references. Existing jsoup/Redux warnings each report two omitted edges. All added occurrence/endpoint records are checked, including AST ownership for 26 field initializers; this is not comprehensive semantic precision.

The pre-registered compiler fixture now recovers 39/45 occurrences versus Graphify 0/45, retaining six misses and all eight negative controls. A separate fixture-specific check verifies supported compiler field/caller pairs and source anchors. Inherited, anonymous/local-class and one qualified static-field access remain missing.

The four registered known-ID public neighbor requests recover the fields and selected-line anchors for Compass (4/4) versus Graphify (0/4); all eight calls succeed. Compass returns 9,883 text / 108,999 wire bytes versus 3,313 / 3,708. No identity-discovery, authored-answer, efficiency or held-out win is claimed.

jsoup's partition changes from 41 to 42 communities, but the unchanged 75 task-pair outcomes stay identical. There is no new community-quality or god-object classification result. Prior Rust and compiler-baseline checkpoints remain in the audit and committed reports.

Latest evidence: compiler source-binding census

Registered the complete 88-file jsoup Java 8 base-source set before compiler binding capture. A public-JDK source oracle independently checks target declarations, source owners and UTF-8 occurrence spans against both frozen native graphs; the JDK is an optional evaluator dependency only.

  • Ordinary source fields: Compass represents 614/616 declarations and verifies 3,047/3,785 references; Graphify represents 0/616 and has no field-contact records.
  • Every one of Compass's 3,047 returned in-scope contacts agrees with the compiler. Graphify has no positive-contact precision denominator.
  • Both represent 131/131 enum constants, but neither recovers their 529 reference occurrences. Compass also misses 738 ordinary references; inherited state, static imports and anonymous fields are documented examples.
  • 44 external fields, 55 array lengths and 31 class literals are separate from the 4,314 internal-reference denominator. Initial captures that mixed these non-source categories remain explicitly superseded.
  • The oracle agrees with all 44 previous scope cases/45 bytecode-checked occurrences and a new 20-record adversarial fixture. Repeated complete captures and saved reviews replay exactly; 176 benchmark tests and the product boundary pass.

This is source-declaration precision evidence for one known development build configuration. It does not score authored explanations, longer walks, functional communities or god-object defects, and is not held-out generalization. No production code changed in this census checkpoint; native Rust/JS/platform/packaging gates were not rerun. Full judgments, missing occurrences, source/build/dependency hashes and replay commands are retained in java_real_field_review.json and the audit.

Verification scope

Latest production checkpoint (f2cacfeb): formatting; 27 Java language integration tests, 244 resolver integration tests, 33 cache contracts; workspace and focused-test Clippy; 1,106 workspace tests passed, two ignored; nine product tests; product boundary; full production fixture qualification, including Markdown and independent React release-binary fixtures; 162 benchmark tests.

All five repository rebuilds, compiler fixture results, raw public responses, full graph deltas and registered task-pair replays are retained with hashes. The qualifying/evaluated debug binary and production sources match. All 227 Graphify package files verify unchanged. Failed native cases, an initial test type mismatch, interrupted/superseded qualification and a corrected diagnostic-script error remain in external artifacts.

Existing fixture omission, linker and unused-mut warnings remain visible. Full hosted CI/platform/packaging/browser results are not claimed. No merge or release was performed.

Still required

  • Fix and qualify the remaining Java/TypeScript access gaps, Go/Python state declarations and other extraction/resolution gaps.
  • Assess authored responsibility explanations and actual god-object defects against defensible source judgments.
  • Broaden edge precision, longer directed walks, ambiguity/unreachable/work-exhaustion cases and functional community evaluation.
  • Re-review invalidated historical hierarchy scorecards and retain competitor wins.
  • Confirm improvements on fresh held-out repositories/questions before making a broad superiority claim.

Explanation source-budget audit (e445c5c)

Registered all five source quotas before capturing the same 20 responsibility facts on the five pinned language repositories. Public resolver/neighbor workflow and source-window rules are identical for both tools; latest Compass graphs are used.

Source quota per subject Compass evidence facts Graphify evidence facts
2,000 bytes 8/20 8/20
4,000 bytes 8/20 8/20
8,000 bytes 14/20 15/20
16,000 bytes 19/20 18/20
32,000 bytes 20/20 18/20

These are source-witness scores, not authored explanations or a broad superiority claim. The historical Click header allowance is retained; literal coverage is one fact lower per tool. At the largest quota Compass consumes 64,644 actual source bytes versus 53,747 and still returns substantially larger graph payloads. All 8,000-byte windows reproduce the earlier results. Source-order starvation and the fixed last-window cap explain remaining gaps in this workflow.

All 20 public calls, 153 membership anchors and 50 window/scoring arms verify independently; repeated verification is byte-identical. All 187 benchmark tests and the product boundary pass. No product code/version changes; native/platform/JS checks were not rerun. Registration, complete score table and artifact hashes are committed; raw captures are retained under external explanation-budget-01. Authored answers, god-object judgments, longer walks, community quality and held-out confirmation remain open.

Optional explanation member focus (293582c; audit 9f17e0c)

Added compass explain OWNER --source-members --member-focus TEXT. It prioritizes distinct normalized matches in recorded callable names, then source order, within existing discovery/source/verification bounds. No source is read to rank, unmatched members remain eligible, and the output names the lexical matches. Default output reproduces the frozen baseline byte-for-byte on all five subjects.

The registered Compass-only before/after experiment passes each full original question with the same 8,000-byte quota. Supported facts improve 14/20 → 15/20, gaining Chi routing evidence with no fact losses on this development panel. Retained source stays 28,129 bytes; stdout grows 58,506 → 60,942 bytes. WalkDir loop/contents-first, Click decorator/owner and Redux enhancer/store evidence gaps remain. Six Redux excerpts correctly retain missing-digest/unverified status.

This is not a new paired Graphify result: its recorded native interface lacks an equivalent flag. The separate paired neighbor-window control remains Compass 14/20 versus Graphify 15/20. No authored-answer, god-object or overall superiority claim is made.

Validation: 19 targeted source/member tests, 42 CLI contract tests, 1,106 workspace lib/bin tests (two existing ignores), nine product tests, formatting, workspace Clippy, product boundary, CLI build and 187 benchmark tests pass. All 15 real CLI invocations and their member order, source spans, provenance statuses and scores verify independently; repeated verification is byte-identical. Native source and binary hashes are retained under external member-focus-02; failed attempts are preserved. Version remains 0.3.30.

Paired member focus: negative result (ce8b739)

Registered a common name-focus helper for both public interfaces before fresh captures. Both tools receive identical questions, identity constraints and source quotas. Returned fields/bindings remain included; duplicate labels cannot increase scores. Source interval ends are fixed before reranking.

Per-subject source quota Source order: Compass / Graphify Common focus: Compass / Graphify
2,000 bytes 8/20 / 8/20 6/20 / 8/20
4,000 bytes 8/20 / 8/20 9/20 / 10/20
8,000 bytes (primary) 14/20 / 15/20 12/20 / 14/20
16,000 bytes 19/20 / 18/20 19/20 / 18/20
32,000 bytes 20/20 / 18/20 20/20 / 18/20

Reject this common window policy as the default. Both tools gain Chi routing but lose its With/shared-state fact. Compass additionally loses two Redux facts; Graphify loses WalkDir contents-first evidence. A promoted Redux getState binding window spends 3,759 bytes beyond the createStore boundary. This result is separate from native callable-span focus; the native before/after gain does not transfer reliably to common anchor windows.

All 202 benchmark tests and product boundary pass. All 20 fresh public responses and 153 anchors match the prior capture, all 50 source-order controls reproduce, and all 100 window/scoring cases verify independently with byte-identical repeated replay. All gains, losses, quotas, costs and raw-evidence hashes are retained in paired_member_focus_review.json and external paired-member-focus-01. No product code/version changes in this checkpoint; broader superiority remains unproven.

Longer call paths: route-selection gap (c39e549)

Registered five new source-grounded endpoint pairs across Go, Python, Java, TypeScript and Rust before capture. Positive witnesses span four to seven calls; conditional calls and overloads are explicit. These remain known development repositories.

Public workflow Compass Graphify
Source-assisted resolver IDs 2/5 0/5
Native short names 1/5 1/5

Both native interfaces return valid six-call Click chains through different branches. Graphify's stdin/reader route was source-reviewed after capture and accepted, with no preregistered occurrence credit. The ID arm includes cross-interface interoperability failures and must not be presented as pure path quality.

Compass's jsoup and WalkDir outputs use structural shortcuts instead of complete call chains. Post-capture graph diagnostics find the registered call edges already present in Compass, identifying route selection as the next production target. Neither tool exposes Redux's named combination function at the frozen coordinate. No graph-oracle IDs replace failed public lookups.

All 204 benchmark tests and the product boundary pass. Twenty resolver calls, 42 CLI requests, all 18 depth diagnostics, source/graph/binary hashes and alternate occurrences are retained and separately verified; repeated verification is byte-identical. The review preserves failed verifier attempts and corrected occurrence-credit bookkeeping. No production changes or native Rust/JS/platform/packaging rerun in this checkpoint. Version stays 0.3.30.

Evidence: longer_path_registration.json, longer_path_witnesses.json, longer_path_label_control_registration.json, longer_path_review.json; raw external longer-paths-01. Next: explicit calls-only trail policy, then a separate named-function-expression identity fix. Broader superiority, authored explanations, genuine god-object judgments, functional community quality and held-out confirmation remain unproven.

Calls-only trail correction

  • Add optional calls-only trails through typed queries, CLI node --calls-only, MCP get_node, and explicit ask call-path syntax. Preserve structural defaults, ambiguity handling, occurrence evidence and bounded work.
  • Replay the frozen five-question panel: source-assisted static call paths improve from 2/5 to 4/5; native short names remain 1/5. All four resolved CLI/ask/MCP bodies agree, 20 resolver outcomes remain unchanged, and nine default stdout/stderr controls are byte-identical.
  • The jsoup alternative uses a missing-body fallback although the caller initializes a shell with a body. Only static edge support is established; execution feasibility remains unproven. The post-capture adjudication and stopped verification are retained, with no frozen-route/runtime credit.
  • Passed: 56 query, 43 CLI, 11 MCP, 1,106 workspace (two existing ignored), nine product and 204 benchmark tests; formatting, workspace Clippy, product boundary and CLI build. Repeated evidence verification is byte-identical.
  • Evidence: benchmarks/agent_query/calls_only_trails_review.json and the audit document. No equal-feature Graphify comparison, held-out or overall superiority claim. Version remains 0.3.30; PR was draft at that checkpoint.

Shared public-neighbor path evaluation

  • Register and implement a common bounded FIFO search using public call-filtered neighbors, with identical endpoint questions and workflow limits. Compass exposes destination IDs; Graphify needs additional label lookups. Ambiguities remain incomplete branches and costs include those lookups.
  • On the five known repository questions: source-supported static call paths are Compass 4/5, Graphify 0/5. Graphify hits intermediate label ambiguity in Chi/Click, no outgoing calls at the jsoup start, and unresolved Redux/WalkDir endpoints. Its native Click path success remains in the separate control. This does not isolate internal path algorithms or establish overall superiority.
  • All 110 public requests replay exactly; all 20 endpoint resolver outcomes are unchanged; all 76 neighbor responses match frozen graph projections. Every returned Compass parallel call record is source-checked. Only Chi/WalkDir have a frozen occurrence at every step; no runtime-feasibility credit is given.
  • Measured response costs: Compass 72 requests / 1,558,220 bytes; Graphify 38 / 11,001. These are not equal-work efficiency results.
  • All 218 benchmark tests and product boundary pass on repeat. Fourteen new tests include an independent oracle over all 4,096 four-node directed graphs at three depth bounds. Retain the first suite attempt's existing subprocess cleanup permission error. No product change or Rust/JS/platform rerun in this evaluation checkpoint; version stays 0.3.30.
  • Evidence: benchmarks/agent_query/public_neighbor_path_review.json, registration, harness, tests and audit document. Authored explanations, god-object judgments, functional communities and held-out superiority remain open.

Externally rated Blob cohort

  • Added a preregistered MLCQ v1.1 Java cohort (CC-BY-4.0): four multi-reviewer positive and four unanimous-none classes from pinned CloudStack and Eclipse revisions. All 55 selected-revision classes and 104 review rows remain visible, including uncertain ratings. Complete source spans were reviewed before queries.
  • Preserved the original limits: the frozen unoptimized Compass development binary on CloudStack timed out at 1,200 seconds; Compass Eclipse produced a 477 MB partial graph above the registered 256 MiB artifact limit. Compass has no native ranking score under this protocol.
  • Graphify admitted both graphs, but none of the four positive or four none classes appeared in its top 100. All eight classes exist at exact graph/source anchors and resolve via public lookups. No precision-at-k or overall superiority claim follows.
  • Verified 18 raw tool request/response pairs. All eight Graphify unfiltered neighbor projections match stored displayed triples. The independent verifier repeats byte-identically; ambiguous hub labels remain unresolved.
  • Added conservative retrieval scoring and unfiltered projection checks. All 237 benchmark tests and the product-boundary gate pass. No production Rust/JavaScript changes or version bump in this checkpoint.
  • A live profile identifies repeated SQL prefix scanning before dollar-delimiter syntax rejection as a follow-up candidate. No optimization result is claimed yet. Broader god-object diagnosis, explanation/community quality and held-out confirmation remain open.

SQL prefix guard and validation checkpoint

  • Reject impossible dollar-delimiter syntax before scanning the SQL statement prefix. Preserve the existing statement-sensitive identifier handling; no graph schema, version or dependency change.
  • Passed four SQL unit tests, 34 SQL domain integration tests, formatting, workspace Clippy, 1,108 workspace tests (two existing ignored), nine CLI product tests, the code-graph fixture qualification, 237 benchmark tests and the product-boundary gate. The fixture gate first failed its missing parser-bundle preflight; reuse of this worktree's already provisioned matching bundle passed, with both outcomes retained.
  • Disclose that the original Blob extraction observations used an unoptimized development binary, with debug information and incremental compilation disabled. Those original timeout/admission outcomes remain unchanged and do not establish release performance.
  • Freeze the SQL file selection, limits, baseline hash and matching candidate build settings before timing. The matched build, SQL and product checks pass. All 24 alternating fresh-output observations are complete and independently verified twice with byte-identical reports. Three smaller files have complete old/new graph equality. On the 411,080-byte file, all baseline runs hit the 180-second cap; all candidate runs finish (174.175-second median) with identical graphs. No largest-file old/new equality or latency ratio can be established from unavailable baseline graphs.
  • Separately register a 1 GiB artifact diagnostic for the same eight Blob classes after the original admission failures. It does not replace the original protocol, retune cutoffs, or establish held-out evidence. Completed follow-up results are recorded below.
  • Evidence: sql_prefix_scan_registration.json, sql_prefix_build_context.json, and mlcq_admission_followup_registration.json under benchmarks/agent_query; external validation run sql-prefix-scan-01. Version remains 0.3.30; no merge or release.

Completed SQL diagnosis and retained full-repository outcome

  • Published every observation in benchmarks/agent_query/sql_prefix_scan_review.json, plus the audit and performance documentation. Rechecked source selection, unchanged copies, binary/build/log hashes, exact alternating order, raw process timings, RSS and every successful graph. Repeated verifier digest: 28a7ee6b8e16ec4c4a4485559eff152318b3e2f5329ea31a3e8ddc87eca3a19b.
  • Retain small-file regressions: baseline/candidate medians are 1.010/1.152 s, 1.004/1.077 s and 2.462/2.284 s. The smallest-file regression exceeds the 10% review threshold; this development diagnosis is not promoted as a qualified performance baseline.
  • Both binaries are unoptimized builds with matched settings; unrelated host activity was not isolated. No general speedup, release-performance or Graphify superiority claim.
  • The same frozen candidate completed the registered whole-CloudStack attempt at the original 1,200-second limit: timeout at 1,200.448 seconds, exit -9, no admitted graph. The separately declared 1 GiB Blob query diagnostic completed on the three admitted graphs. Original failures remain unchanged.

Observed residual SQL cost

  • A separate two-second sample during the full CloudStack follow-up repeatedly enters SQL State::add_access_matches and Regex::new. Source inspection confirms that fixed read/write patterns are recompiled per statement. This short snapshot does not attribute the entire run or establish an optimization result.
  • The full-repository wall time includes this profiling overhead and remains an operational observation. The completed 24 SQL timing observations were not sampled or changed.
  • Registered a release-mode comparison in benchmarks/agent_query/sql_regex_cache_registration.json, reusing all four complete SQL files and retaining all failures and graph-equality requirements. The copied release baseline now matches an unchanged release rebuild byte-for-byte. Cache implementation and release timing results remain pending.

Completed larger-artifact Blob follow-up

  • Preserve the same eight reviewed classes and native ranking cutoffs under the declared 1 GiB graph cap. Compass retrieves 0/2 admitted positive classes, with two positives unavailable after the CloudStack timeout; Graphify retrieves 0/4. The corresponding none-class outcomes have the same counts. No classifier decision is inferred from ranking absence.
  • Both tools resolve all four Eclipse classes at exact reviewed source coordinates. The none-rated SimpleValueProperty has more unfiltered neighbor records than either reviewed positive Eclipse class in both tools, illustrating the limitation of connectivity as a responsibility measure.
  • Compass's 100 returned hub IDs, degrees and stored anchors match the graph, as do all 100 connectivity summaries. All four Compass neighbor projections match 122 complete record appearances. Graphify retains 37 ambiguous hub labels across 200 rows; all eight neighbor projections match 234 displayed triples. This is graph consistency, not broad source precision or superiority.
  • Corrected an auditor defect: unfiltered native responses serialize relationFilter as an empty string, while the mock and auditor expected null. Two regressions failed before the fix. Retain the first invalid verification, every raw capture and unchanged primary ranking outcomes. Corrected verification repeats byte-identically (a659988abe0c9e64449b8265f06e066425fb6963e340176ccfde8f0a1e258d3d). All 27 tool calls are checked; 238 benchmark tests and product boundary pass.
  • Preserve Compass Eclipse's 62 omitted edges. Disclose that the frozen Graphify CloudStack build skipped all 119 SQL files because its optional parser dependency was absent, and reported three other files with syntax errors. Whole-repository requests did not produce equal language coverage; no controlled extraction speed ranking follows.
  • Evidence: benchmarks/agent_query/mlcq_admission_followup_review.json, original review coverage disclosure, native-contract auditor regression and the audit document. Native product remains the tested prefix guard, version 0.3.30. God-object classification, authored explanations, functional communities and final independent confirmation remain open.

SQL regex-cache release comparison

  • Cache the six fixed SQL read/write access regexes once per process; preserve pattern order, captures, alias/CTE handling and emitted records. Version remains 0.3.30.
  • All 24 registered release observations succeeded on four complete CloudStack SQL files. Full graph JSON, including array order and every field, is equal across the six runs of each file. Repeated verification is byte-identical.
  • Baseline → candidate medians: 938 bytes, 0.596 → 0.570 s; 1,949 bytes, 0.580 → 0.594 s; 10,288 bytes, 0.657 → 0.666 s; 411,080 bytes, 13.121 → 1.650 s. Small-file regressions remain visible.
  • Baseline and candidate use matching release settings and unchanged build manifests, lockfile, toolchain and repository configuration. An unchanged release rebuild confirmed the frozen baseline before the edit. Host work and ordinary OS caches remain limitations. This is a four-file development panel, not whole-product qualification or a Graphify speed ranking.
  • Passed formatting, 5 SQL unit tests, 34 domain tests, workspace Clippy, 1,109 workspace tests (2 ignored), 9 product tests, complete code-graph fixture qualification, 238 benchmark tests, product boundary and the release build.
  • sql_regex_cache_review.json retains every outcome and artifact digest. sql_regex_cache_whole_registration.json declares the separate complete-source follow-up, now completed under the original 1,200-second extraction bound and declared 1 GiB artifact cap (results below). Both earlier timeouts remain unchanged. God-object classification, authored explanation quality, functional communities and independent final confirmation remain open.

Complete-source release and Blob audit results

  • Both frozen release binaries completed the full pinned CloudStack source scope and emitted identical 671,502,979-byte graphs: 173,689 nodes, 502,158 edges, 2,153 communities. Both graphs carry a warning for 1 omitted node and 45 omitted edges, with 0 identity collisions. They exceed the original 256 MiB artifact cap and fit the declared 1 GiB follow-up cap.
  • The pre-cache release baseline already succeeds. Recovery cannot be credited solely to the regex cache. Earlier unoptimized timeouts and original admission failures remain unchanged.
  • The candidate took longer in the single complete-source observation: 280.053 s versus 193.966 s. Cause unproven; no whole-repository speedup claim. A separately registered three-repetition alternating panel will retain CPU time, peak RSS, graph equality and every outcome.
  • The registered release-server audit completed all four sessions and verified 36 raw tool calls twice with byte-identical reports. Both tools resolve all 8 reviewed classes, but neither returns any of the 4 consensus-positive classes at cutoffs 10, 50 or 100. Both also return 0/4 none-rated classes; missing rows are not classifier decisions or population precision.
  • All 200 Compass hub identities/degrees/anchors and connectivity summaries match the stored graphs. Graphify has 163 uniquely resolved displayed hub identities with matching degree; 37 ambiguous labels remain unresolved. All 8 Compass neighbor projections match 266 full record appearances; all 8 Graphify projections match 234 displayed triples. These are graph-consistency checks, not source-edge precision.
  • Eclipse keeps its original frozen Compass graph and 62-edge omission warning, queried by the same release server as CloudStack. Graphify's environment and graphs remain frozen, including the 119 omitted CloudStack SQL files. No controlled Graphify speed ranking follows.
  • Reports: sql_regex_cache_whole_review.json and mlcq_release_followup_review.json. Repeat registration: sql_regex_cache_whole_repeat_registration.json. No production code changed after the validated cache commit; version remains 0.3.30. Actual god-object diagnosis, explanation quality, functional communities and independent final confirmation remain open.

Complete-source timing repeat

  • All six registered alternating release runs complete with identical 671,502,979-byte partial graphs. Baseline wall times: 235.574, 215.288, 213.526 s; candidate: 204.082, 210.928, 203.430 s. Medians are 215.288 s and 204.082 s (ratio 0.948); the registered >10% median regression flag is false.
  • Repeated verification checks frozen executables, clean pinned source, raw streams/resource measurements and every complete graph hash; both reports are byte-identical. Peak RSS, CPU times, UTC timestamps and host load are retained in sql_regex_cache_whole_repeat_review.json.
  • Preserve the initial 193.966/280.053 s observation, original timeouts and graph-admission failures. Timing variation remains unexplained. Ordinary OS caches and unrelated host work were uncontrolled; our other compilation/extraction/query/large-graph verification was paused during the panel. This is one known development repository, not a general performance or Graphify superiority result.
  • Every graph retains one omitted node and 45 omitted edges. No production code or version change in this checkpoint.

Java enum-reference census follow-up

Production a8b4a490 preserves enum constants, constant-specific body ownership, lexical shadowing and selector-derived switch labels. Source types retain precedence over imports; ambiguous or unsupported receivers remain unresolved. AST cache semantics advance from 9 to 10; package version remains 0.3.30.

  • Registered complete jsoup compiler census: 355/529 enum references (previously 0) and 3,106/3,785 ordinary field references (previously 3,047). All 3,461 returned in-cohort contacts agree with compiler target, owner and span. All prior verified occurrences remain. Frozen Graphify retains zero such contacts, with no positive-contact precision denominator.
  • Keep 679 ordinary and 174 enum misses. Of 670 added graph references, 256 are test-source references outside the compiler cohort; source-token/provenance checks do not prove those bindings. Both jsoup graphs retain two omitted edges.
  • All eight fresh Go/Python/TypeScript/Rust preservation controls succeed; each candidate is byte-identical to its baseline and preserved original. Repeated census/control verification reports match exactly. The 227 Graphify package files and 58 package versions remain unchanged.
  • jsoup communities move from 42 to 40. A post-change replay of all 75 existing source task-pair outcomes remains unchanged; this does not establish functional community quality.
  • Native checks pass: formatting, 38 Java language integrations, all 533 resolver tests, 33 cache contracts, focused/workspace Clippy, 1,109 workspace tests (two existing ignored), nine product tests, complete graph-fixture qualification, 238 benchmark tests, product boundary and release build. Suites overlap. No standalone browser/platform/packaging matrix was rerun locally.
  • Retain initial native failures, the invalid-initializer correction, source/import regression and scope-test correction. The full resolver suite and focused Clippy pass after the scope correction.

Full provenance and limitations: enum follow-up report and audit. These known development results do not establish held-out performance, authored answer quality, god-object classification or overall superiority. PR was draft at that checkpoint; no merge or release.

Java dotted nominal-owner checkpoint

Production be2fb542, review b54c20ee: joins established dotted Java nominal
receivers to canonical nested declarations, retaining ambiguity, exact source
precedence and bounded lookup. The unchanged compiler-backed jsoup census now
verifies 3,115/3,785 ordinary field and 368/529 enum references. All 3,461
previously verified occurrences remain; all 3,483 returned in-cohort contacts
match the compiler. There are still 670 ordinary and 161 enum misses. Frozen
Graphify returns zero of these contacts, with no positive-contact precision
denominator. These are known development results.

All eight other-language controls reproduce complete baseline/original graphs.
Of 49 added references, 27 are outside the compiler cohort (21 test, six Java 11
overlay) and remain semantically unscored. All 75 existing community task-pair
outcomes remain unchanged despite a 40-to-43 jsoup community count change.
Reusing all 206 v10 AST cache entries yields the fresh candidate graph exactly.
Census/control/community verifier repeats agree byte-for-byte.

Formatting, focused/workspace Clippy, 536 resolver tests, 1,109 workspace tests
(two existing ignored), nine product tests, full code-graph fixture gate, 238
benchmark tests, product boundary and release build pass. Suites overlap.
Standalone browser/platform/packaging matrices were not rerun. Version stays
0.3.30; the PR was draft at that checkpoint. Overall superiority, god-object diagnosis, authored
explanation quality and functional community quality remain unproven.

Snapshot-coherence checkpoint

Commit f48ccad binds compact, full and typed JSON graph views to one bounded content snapshot and rebuilds disposable caches under digest-based format markers. Package version remains 0.3.30; no historical realization or graph schema changes.

  • Native qualification passes formatting, focused/workspace Clippy, 1,112 workspace library/binary tests (two existing ignored), nine product tests, full code-graph fixtures, 238 benchmark tests, product boundary and release build. Seven separate core integration tests fail identically on the unchanged committed baseline; they remain open, with no passing claim.
  • The two-binary synthetic replay verifies 432 raw calls over 48 replacement cases, with zero mismatches. The candidate uses the replaced graph in all 72 post-replacement observations.
  • Six fresh sessions per frozen graph verify 828 raw calls on Go, Python, Java, TypeScript, Rust and the separate full CloudStack diagnostic. All 46 available fixed questions return identical complete tool responses; 18 Redux path observations remain unavailable because the exact endpoint is missing. No fallback endpoint was chosen.
  • Cost is mixed: 31 warm question medians regress by more than 10%, covering the six lightweight questions on each language graph and CloudStack graph_stats. Available hub, neighbor and path questions improve locally. Median peak RSS rises on all five language mixed-query sessions. Five of six CloudStack session memory measurements are unavailable, so no paired large-graph memory claim is supported.

Full cache review includes per-query cold/warm measurements, response equality, resource gaps and limitations. This correction does not establish improved graph source truth, god-object diagnosis, authored explanations, functional communities or overall Graphify superiority. PR was draft at that checkpoint; no merge or release.

Core inference fixture follow-up

Commit ffe50a2 makes Max inference explicit in the seven semantic-preservation integration fixtures (12 build configurations), preserving every assertion and the public Low default. The same seven failures were recorded on the unchanged baseline; all 13 core integration tests now pass. The test-only follow-up changes no production Rust input or frozen cache-evaluation binary. Formatting, focused Clippy, workspace Clippy, and workspace library/binary tests also pass under pinned offline parser inputs. Fixture review.

Class-span diagnostic

Registration 3ae08aa preceded the frozen-graph scan; review 6b3f692 preserves the negative outcome. Across CloudStack and Eclipse, the exact file/name/span rule uniquely joins 48/55 reviewed classes, including all eight consensus primary cases. Seven uncertain secondary cases remain unavailable. Ranking all 13,001 source-located class nodes solely by inclusive source span retrieves 0/4 consensus-positive classes at cutoffs 10, 50 and 100; the positive ranks are 131, 411, 426 and 714. This offers no improvement over the original top-100 hub retrieval and no god-object diagnosis. The first bounded scanner failure and corrected repeat verification are retained. No Compass or Graphify extraction/query was rerun. Complete review.

Full-class relation evidence inventory

Registration 19877907 preceded a bounded read-only scan of the same frozen CloudStack and Eclipse Java graphs. Review b97ded50 verifies all 13,001 class nodes, direct class/member ownership and method call/field-reference coverage, with all 55 prior review rows retained (48 unique joins, seven unavailable; all eight consensus cases uniquely joined). CloudStack has 84,998 uniquely owned members, Eclipse 69,878; neither graph has a member with multiple distinct direct class owners. Both have unowned members and unresolved relation targets.

All eight consensus cases have zero observed foreign-field references under the registered rule. CloudStack also has 1,042 reads and 6,161 writes edges outside that rule; Eclipse's older producer has no such categories. This is stored-graph evidence, not independent source truth or absence of foreign-state access. No WMC/ATFD/TCC score, god-object classifier, new Graphify result or held-out win was computed. Two verifier runs agree byte-for-byte. Registration and review.

CloudStack access-relation boundary

Registered 990d0286 and verified 46d5c98f: all 1,042 reads and 6,161 writes edges in the frozen CloudStack graph join SQL query/view/procedure nodes to database tables. There are zero method-to-field reads/writes edges, so these categories cannot repair the Java foreign-field evidence gap. Five endpoint-kind pairs account for every edge; two verification runs agree byte-for-byte. The original 55 review joins and frozen graphs are unchanged. Registration and review. No source-truth, classifier or Graphify win is claimed.

Rated Java class member source check

Registration ebf393e2 fixed the eight consensus-rated files and frozen graph hashes before a JDK 17 syntax parse. Review c989292d independently matches 95/95 direct methods and 65/65 direct fields to exactly one Compass class member by kind, normalized display name and source range; there are no extra graph methods or fields in these eight classes. Five explicit source constructors are reported but outside this registered graph comparison. The first literal-name comparison failed for all methods because Compass displays foo as .foo(); both raw runs remain recorded. Repeated oracle and verifier outputs agree byte-for-byte. One positive-rated class is a test class. This checks declaration coverage in known Java cases, not call/field target correctness, cohesion, god-object classification or held-out quality. Registration and full review.

Explicit Java field-token source check

Registered 29208f03; review 6086f7f1 uses a separate JDK syntax parser and exact UTF-8 byte anchors for explicit this.field in the same eight rated classes. Five direct-method tokens qualify and all five have exactly one matching Compass references edge to the source-matched field, with no competing edge at that token. Four constructor tokens and one this.getClass() invocation are retained outside the denominator. This is a tiny selected subset, not general field-reference accuracy, a cohesion measure or a Graphify/god-object win. The initial collector preflight failure and repeated byte-identical oracle/verifier outputs are recorded. Registration and review.

Compiler-bound direct-field census

Registration a523de10 expanded the eight rated Java classes to every direct-method field token that a public JDK 17 compiler could bind to a directly declared field. Review b4a499d1 verifies 265/265 compiler-bound own-field sites against exact Compass method, field and byte anchor; all 265 stored own-field records in these selected classes are represented, with no extra record at those sites. The five previous explicit-this controls retain their edge IDs. Three same-named candidates bound to nonfield symbols.

The compiler intentionally lacked project dependencies and emitted 1,047 errors. An initial 1,000-error cap was corrected within the registered 10,000-diagnostic bound; candidate bindings were unchanged. This is selected own-field evidence, not recall for unresolved/inherited/foreign fields, god-object diagnosis, held-out quality or a new Graphify comparison. Registration and full review.

Source-first three-language hub checkpoint

Registration 4d5ca44d pinned six source-located Go/Python/Rust declarations, the clean source commits, frozen graph hashes, both server environments and public god_nodes(top_n=100) before any hub call. Commit dc704ca2 adds the bounded collector, independently replayed review and audit account. All six public MCP calls succeeded; the two review runs are byte-identical. Raw transcripts remain in external run source-first-hubs-01.

Compass retrieves 2/6 selected witnesses in the top 10, 3/6 in the top 50 and 3/6 in the top 100; Graphify retrieves 1/6, 3/6 and 4/6. Both return FastAPI at rank 1; Graphify returns Celld PeerAuth at rank 55 while Compass leaves it outside the top 100. Compass has explicit graph IDs, matching source anchors and graph-consistent degrees for all 300 returned rows. Graphify emits labels only: 247/300 are globally unique and all of those have graph-consistent degrees; ambiguous labels remain unresolved. The selected source roles are not god-object defect labels, and the extracted graphs differ. No classifier accuracy, superiority, merge or release claim follows. Package version stays 0.3.30.

Rated Java class-only degree diagnostic

Registration 3ebdf975 froze the same class-only distinct endpoint-pair degree rule and four existing CloudStack/Eclipse graph hashes before computing ranks. Commit 52450894 adds the bounded scanner, review, and audit account. The 16 exact graph identities for eight externally rated classes all remain eligible. Two scans produced byte-identical results.

Neither tool retrieves any of the four consensus-positive Java classes at top 10, 50 or 100 after filtering to source-located classes. Compass positive ranks are 340, 728, 999 and 1,529; Graphify's are 463, 1,315, 1,403 and 1,740. Neither retrieves a consensus-none class at those cutoffs either. The graphs differ in extraction coverage and direction, and this is known development data, not held-out classifier evaluation. Class-only degree cannot be claimed as a repair for god-object detection. No production code or package version changed.

Paired rated-Java member declaration coverage

Registration 3996067c froze the independent JDK syntax census of eight rated Java classes (95 direct methods, 65 direct fields), source hashes, exact class IDs and four graph hashes before this comparison. Commit dd028419 adds the bounded scanner, byte-identical repeated review and audit account.

Under the registered exact file/name/source-span join, Compass has unique directly owned nodes for 95/95 methods and 65/65 fields. Graphify has 90/95 methods and 0/65 field-compatible declaration nodes in these files. Its five method misses are later overloads whose names share a node anchored at the first overload; some ownership edges cite later overload lines without distinct identities. Graphify still has field-context references, so this does not claim it extracts no field information. This selected development result supports declaration coverage only. It does not establish field-access correctness, class cohesion, god-object diagnosis, whole-repository accuracy or broad superiority. No production code or package version changed.

Stored hub member evidence

Commit ee21fae8 adds memberEvidence to directed, typed class/struct rows in public MCP god_nodes. It counts uniquely owned stored methods, fields and method-to-own-field reference records; ambiguous ownership is excluded. sourceCoverage: "unverified" and the contract note make clear that stored reference records include every confidence state. Hub eligibility, degree and order are unchanged; this is not god-object classification.

A two-pass replay on frozen Litestream, FastAPI and Celld graphs confirms all 300 prior hub rows are unchanged after removing the additive field, with identical new results between passes. Member evidence appears in 33, 34 and 26 of their respective top-100 rows. Celld PeerAuth remains outside Compass's top 100. A separate frozen Eclipse top-5,000 read finds three rated Java classes whose selected direct-member and own-field counts match the prior independent source/compiler checks; the fourth rated class is outside the request. The Go/Python field zeros and selected Java checks do not establish cross-language completeness, reference precision, cohesion or broad superiority. Reports and replay tool are in benchmarks/agent_query/; raw responses stay in the external evaluation run.

Focused graph/MCP regressions, workspace formatting, Clippy and library/binary tests, CLI product tests, product boundary and fixtures-only code-graph qualification passed. Package version remains 0.3.30; PR was draft at that checkpoint.

Go struct field declaration census

  • Publish exact, source-anchored named Go struct field declarations and ownership edges; advance the disposable AST cache to v12.
  • A registered all-file Litestream Go parser census found 993 named fields in 167 structs. The candidate graph matched 993/993 exact identities and unique owners, versus 0/993 in each frozen baseline graph. All 993 added nodes and containment edges are source matched; no old node or edge was removed. Graphify also retains 323 field-type context records, which are distinct from named declarations.
  • The fixture gained one valid field node and edge, shifting clustered topology to 238 communities and two exact cross-community edges. The two policy minima were updated to the observed fixture values; every other threshold remains unchanged.
  • Verification: workspace lib/bin tests and Clippy, 46 universal-evidence tests, 33 file contracts, 9 CLI product tests, Python probe tests, product boundary, formatting, and the full fixtures-only code-graph gate including React qualification.
  • Scope: one known Go development repository. This does not establish field-access target correctness, god-object classification, or overall superiority over Graphify.

Go direct receiver-field access follow-up

Commit 5553c8c5 adds source-anchored Go field contacts and resolves exact owner-qualified field declarations before their type aliases. The preregistered Litestream source oracle has 2,742 direct receiver-field sites; all 2,742 join to unique method-to-field references at the exact source byte in the candidate graph, versus zero in the frozen Compass and Graphify graphs under this field-target join. All 4,650 added member-access links target fields, and no old node ID or relationship was removed. The 1,908 added links outside this narrow oracle remain uncredited and require separate precision review.

The audit graph exceeded its preregistered 16 MiB read cap, so the review records a post-capture increase to 24 MiB without filtering records. The cold and warm graph bytes and two full probe reports are identical. Public god_nodes(top_n=500) shows 113 own-field reference records for Store and 13 for HeartbeatClient; these are stored contacts, not god-object labels.

Focused Go, resolver and cache tests, workspace format/Clippy/lib/bin tests, and the product-boundary check passed. After mounted-volume capacity was restored, the full fixtures-only gate passed, including its release-mode React phase. An initial resumed run stopped on a missing snapshot during lifecycle checks; the complete diagnostic repeat passed, so that intermittent failure remains open for investigation. Broad superiority and god-object diagnosis remain unclaimed. Compass stays at 0.3.30 and the PR was draft at that checkpoint.

Typed Go parameter-field audit

Registration d82d94b9 and review c1b0d78f independently verify 276/276 explicitly typed method-parameter field contacts on the pinned Litestream source at exact graph sites. The frozen pre-access Compass graph has no matching links; Graphify has no exact declaration targets for the 98 fields, though it retains 19 field-type context records. These 276 links are disjoint from the prior 2,742 direct-receiver links: 3,018/4,650 added Go member-access links now have narrow source-oracle credit, and 1,632 remain uncredited. Both full source and graph probes were byte-identical across two runs. This remains one known Go development repository and does not prove whole-graph precision or broad superiority.

Final gate and Windows test follow-up

Commit bfe9068c records the complete local fixtures-only gate and adjusts the graph-artifact snapshot test to use the path-rename sequence already exercised by a passing Windows test. Both Windows CI architectures failed the earlier overwrite attempt with Access is denied; the revised test retains the opened-handle snapshot assertion. Local formatting, compass-model tests, workspace Clippy, and workspace library/binary tests pass on the revised patch. CI for that head was pending at that checkpoint.

@forhappy
forhappy marked this pull request as ready for review September 28, 2026 07:12
@forhappy
forhappy merged commit 8ad5b23 into main Sep 28, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant