Fix hub analysis and bounded trails; audit graph comparison evidence - #340
Merged
Merged
Conversation
forhappy
marked this pull request as ready for review
September 28, 2026 07:12
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Audit and implementation work toward more reliable code-graph intelligence. Broad superiority over Graphify and actual god-object detection remain unproven. Compass stays at 0.3.30. These implementation changes are ready for review and merge; this PR does not start a release.
Latest held-out three-language checkpoint (2026-09-28)
Commits
a269c1c9and3b7db6a2register a source-selected Anyhow/Rust, MarkItDown/Python, and Vue/TypeScript panel, correctOwner::membersuffix lookup, and remove repeated TypeScript receiver work that caused Vue extraction to time out. The registered original Vue build exceeded its 1,800-second limit; after the performance fix, a Vue-only replay completed Compass extraction in 75.996 seconds versus Graphify's 4.246 seconds. This is a development replay, not a frozen full-panel speed comparison. Both tools failed all four Vue text questions because both omit the namedbaseCompilecallback declaration.Independent source-site review of thirteen selected direct calls found Compass 10/13, Graphify 7/13 exact owner, target and occurrence links: Anyhow 2/2 versus 0/2; MarkItDown 8/8 versus 7/8; Vue 0/3 for both. This selected sample does not estimate whole-graph precision or establish broad superiority. The initial Anyhow
callers Chain::newtext result was a Compassno_match; after the suffix fix, a post-registration replay returns the exactErrorImpl::chaincall. Original registered results remain unchanged. See the detailed review, registration, and source-site reviewer.Current-head local verification passed: formatting, workspace lib/bin Clippy and tests, 114 TypeScript evidence tests, 44 CLI code-query tests, product boundary, and the complete fixtures-only code-graph gate including release-mode React qualification. GitHub CI for this head is still running. The previous head's Windows ARM and x86 native jobs passed after the snapshot-test fix. The remaining Vue callback-owner gap and need for a fresh full twelve-question panel are documented for a separate optimization session.
Detailed evidence and history
The authoritative checkpoint history, acceptance matrix, failed attempts, invalidated claims and remaining work are in the audit. Protocols, per-case judgments, hashes and replay tools are in benchmarks/agent_query. Compatibility and migration effects are documented in COMPATIBILITY.md and MIGRATION.md.
This description replaces an accumulated 65,436-character checkpoint log. Historical results remain in the audit and committed review files; they are not pooled into one accuracy or superiority score. Real-source panels include Cobra/Go, Flask/Python, Gson/Java, Zod/TypeScript and Axum/Rust, plus Chi, Click, jsoup, Redux, WalkDir and a separate fd sample. Graphify 0.9.67 package/environment provenance and frozen inputs are retained externally. Reused development questions are identified as such.
Current comparison checkpoints
These rows measure different protocols and must not be combined:
The source-window comparison uses the same 8,000-byte quota per subject and 35,673 total source bytes per tool. Compass's neighbor responses are much larger (372,829 versus 7,534 full MCP response bytes). Its separate native member-source mode supports 13/20 facts, but Graphify has no equivalent flag; that is not a paired win. The original unscoped-window result, 11/20 versus 15/20, remains recorded.
Graphify retains advantages in several Java/Rust explanation cases and earlier shared-pass token measurements. Scores have been corrected when source review exposed oracle, subject, ambiguity or provenance mistakes. No selected-fact recall or graph-consistency score establishes broad assertion precision.
Source-first Go/Python/Rust checkpoint
Registration commit
ab93cd34fixed 12 selected declaration, caller, callee and one-hop path questions on clean pinned Litestream, FastAPI and Celld sources before either tool ran. Commit0c106a6eadds the source-and-edge review, verifier and audit account. Raw commands, graphs and responses remain in the externalsource-first-3lang-01run.The registered text score is Compass 9/12, Graphify 8/12. The separate source-supported score is Compass 9/12, Graphify 6/12: Graphify's Litestream and FastAPI path outputs contain the requested names but follow three- and four-hop structural routes instead of the direct source calls. Both graphs contain the FastAPI call, yet the registered short-name requests collide with seven methods; Compass safely reports ambiguity and Graphify's path chooses a wrong source after an ambiguity warning. In a separate unscored control, Compass full names and exact IDs recover the direct call in
explain,calleesandpath; Graphify file-qualifiedexplainfinds it, while qualified and exact-IDpathrequests still select unrelated routes in this example. Celld's selected questions pass for both tools. Two verifier runs are byte-identical. The sample does not prove whole-graph quality, authored explanations, god-object diagnosis, performance or broad superiority. No production code changed in this checkpoint.FastAPI owner-qualified navigation correction
Commit
2a058e1bmakes typedask/call/trail queries and legacyexplain/pathresolve a unique owner/member suffix such asAPIRoute.get_route_handleragainst the storedfastapi.routing.APIRoute::get_route_handler. Typed resolution checks the complete bounded exact leaf-name posting; duplicates and truncation stay ambiguous, and a nonexistent owner cannot borrow a same-named method. Exact IDs and full names retain precedence. JSON schemas, graph artifacts and package version 0.3.30 are unchanged.On the same frozen FastAPI graph, a post-registration replay manifest pins the new binary and five raw outputs. Owner-qualified
askcallees and call-path answers now reach the exact line-1232 call without diagnostics;explainidentifies the line-1225 declaration andpathrenders one extracted call hop.MissingRoute.get_route_handlerfails. This is development work after the 12-question registration; the original 9/12 versus 8/12 text and 9/12 versus 6/12 source-supported scores stay unchanged. Graphify was not rerun for this patch.Verification: 44/44 query CLI tests, focused native JSON/store and ambiguity tests, workspace formatting, workspace library/binary Clippy and tests, CLI product tests, and the product boundary gate passed. Unmodified
compass-query --all-targets --all-featuresClippy stops on a pre-existing redundant closure intests/explanation_members.rs; rerunning with only that lint allowed passed all targets. Broad query superiority, authored explanation quality and god-object diagnosis remain unproven.Rust
fdcollected-builder checkpointCommit
68f71422recovers source-proven Rust method receivers through standardResult<Vec<_>>collection,Okmatch bindings, vector loops, anditer().mapclosures. The native regression also rejects custom aliases and a constructor with a different return type. AST cache semantics advance to 11; package version stays 0.3.30.The unchanged registered
fdsuite and pinned source checkout were replayed with freshly rebuilt graphs. Compass now passes 12/12 selected text oracles versus Graphify 9/12; the prior Graphify-onlyexecute_batchcallees row now passes for both. Exact Compass call edges reachCommandBuilder::push,finish, andexit_codeat source lines 104, 111, and 116. The replay manifest pins source, suite, binary, and graph hashes and lists all six newly source-reviewed calls. Another 213 new reference keys lack independent precision review. On nine shared passes, estimated answer medians are 288 versus 114 tokens. This one known Rust sample does not establish broad superiority, god-object diagnosis, or performance parity.Formatting, workspace library/binary Clippy and tests, targeted integration suites and Clippy, and fixtures-only code-graph qualification passed. The audit retains the earlier Graphify win and the corrected replay as separate checkpoints.
Java field-access production checkpoint
Production commits
09ca58e6andf2cacfebemit bounded lexical Java field-access evidence and preserve exact occurrences, field-only targets and unresolved receivers. Source member types shadow imports/package prefixes; duplicate receiver types cannot be selected by field availability. Compiler/native counterexamples caught both defects before publication; the initial capture is explicitly superseded. AST cache semantics advance from 8 to 9; rebuild graphs for these facts. Package version remains 0.3.30.The unchanged five-repository comparison now supports 8/20 registered accesses for Compass versus 0/20 for Graphify, up from 4/20 after the Rust correction. All four selected Java sites are recovered. Chi, Click, Redux and WalkDir graphs are byte-identical; jsoup retains all 6,116 nodes and 21,110 old edges and adds 3,896 field references. Existing jsoup/Redux warnings each report two omitted edges. All added occurrence/endpoint records are checked, including AST ownership for 26 field initializers; this is not comprehensive semantic precision.
The pre-registered compiler fixture now recovers 39/45 occurrences versus Graphify 0/45, retaining six misses and all eight negative controls. A separate fixture-specific check verifies supported compiler field/caller pairs and source anchors. Inherited, anonymous/local-class and one qualified static-field access remain missing.
The four registered known-ID public neighbor requests recover the fields and selected-line anchors for Compass (4/4) versus Graphify (0/4); all eight calls succeed. Compass returns 9,883 text / 108,999 wire bytes versus 3,313 / 3,708. No identity-discovery, authored-answer, efficiency or held-out win is claimed.
jsoup's partition changes from 41 to 42 communities, but the unchanged 75 task-pair outcomes stay identical. There is no new community-quality or god-object classification result. Prior Rust and compiler-baseline checkpoints remain in the audit and committed reports.
Latest evidence: compiler source-binding census
Registered the complete 88-file jsoup Java 8 base-source set before compiler binding capture. A public-JDK source oracle independently checks target declarations, source owners and UTF-8 occurrence spans against both frozen native graphs; the JDK is an optional evaluator dependency only.
This is source-declaration precision evidence for one known development build configuration. It does not score authored explanations, longer walks, functional communities or god-object defects, and is not held-out generalization. No production code changed in this census checkpoint; native Rust/JS/platform/packaging gates were not rerun. Full judgments, missing occurrences, source/build/dependency hashes and replay commands are retained in
java_real_field_review.jsonand the audit.Verification scope
Latest production checkpoint (
f2cacfeb): formatting; 27 Java language integration tests, 244 resolver integration tests, 33 cache contracts; workspace and focused-test Clippy; 1,106 workspace tests passed, two ignored; nine product tests; product boundary; full production fixture qualification, including Markdown and independent React release-binary fixtures; 162 benchmark tests.All five repository rebuilds, compiler fixture results, raw public responses, full graph deltas and registered task-pair replays are retained with hashes. The qualifying/evaluated debug binary and production sources match. All 227 Graphify package files verify unchanged. Failed native cases, an initial test type mismatch, interrupted/superseded qualification and a corrected diagnostic-script error remain in external artifacts.
Existing fixture omission, linker and unused-mut warnings remain visible. Full hosted CI/platform/packaging/browser results are not claimed. No merge or release was performed.
Still required
Explanation source-budget audit (e445c5c)
Registered all five source quotas before capturing the same 20 responsibility facts on the five pinned language repositories. Public resolver/neighbor workflow and source-window rules are identical for both tools; latest Compass graphs are used.
These are source-witness scores, not authored explanations or a broad superiority claim. The historical Click header allowance is retained; literal coverage is one fact lower per tool. At the largest quota Compass consumes 64,644 actual source bytes versus 53,747 and still returns substantially larger graph payloads. All 8,000-byte windows reproduce the earlier results. Source-order starvation and the fixed last-window cap explain remaining gaps in this workflow.
All 20 public calls, 153 membership anchors and 50 window/scoring arms verify independently; repeated verification is byte-identical. All 187 benchmark tests and the product boundary pass. No product code/version changes; native/platform/JS checks were not rerun. Registration, complete score table and artifact hashes are committed; raw captures are retained under external
explanation-budget-01. Authored answers, god-object judgments, longer walks, community quality and held-out confirmation remain open.Optional explanation member focus (293582c; audit 9f17e0c)
Added
compass explain OWNER --source-members --member-focus TEXT. It prioritizes distinct normalized matches in recorded callable names, then source order, within existing discovery/source/verification bounds. No source is read to rank, unmatched members remain eligible, and the output names the lexical matches. Default output reproduces the frozen baseline byte-for-byte on all five subjects.The registered Compass-only before/after experiment passes each full original question with the same 8,000-byte quota. Supported facts improve 14/20 → 15/20, gaining Chi routing evidence with no fact losses on this development panel. Retained source stays 28,129 bytes; stdout grows 58,506 → 60,942 bytes. WalkDir loop/contents-first, Click decorator/owner and Redux enhancer/store evidence gaps remain. Six Redux excerpts correctly retain missing-digest/unverified status.
This is not a new paired Graphify result: its recorded native interface lacks an equivalent flag. The separate paired neighbor-window control remains Compass 14/20 versus Graphify 15/20. No authored-answer, god-object or overall superiority claim is made.
Validation: 19 targeted source/member tests, 42 CLI contract tests, 1,106 workspace lib/bin tests (two existing ignores), nine product tests, formatting, workspace Clippy, product boundary, CLI build and 187 benchmark tests pass. All 15 real CLI invocations and their member order, source spans, provenance statuses and scores verify independently; repeated verification is byte-identical. Native source and binary hashes are retained under external
member-focus-02; failed attempts are preserved. Version remains 0.3.30.Paired member focus: negative result (ce8b739)
Registered a common name-focus helper for both public interfaces before fresh captures. Both tools receive identical questions, identity constraints and source quotas. Returned fields/bindings remain included; duplicate labels cannot increase scores. Source interval ends are fixed before reranking.
Reject this common window policy as the default. Both tools gain Chi routing but lose its With/shared-state fact. Compass additionally loses two Redux facts; Graphify loses WalkDir contents-first evidence. A promoted Redux getState binding window spends 3,759 bytes beyond the createStore boundary. This result is separate from native callable-span focus; the native before/after gain does not transfer reliably to common anchor windows.
All 202 benchmark tests and product boundary pass. All 20 fresh public responses and 153 anchors match the prior capture, all 50 source-order controls reproduce, and all 100 window/scoring cases verify independently with byte-identical repeated replay. All gains, losses, quotas, costs and raw-evidence hashes are retained in paired_member_focus_review.json and external paired-member-focus-01. No product code/version changes in this checkpoint; broader superiority remains unproven.
Longer call paths: route-selection gap (c39e549)
Registered five new source-grounded endpoint pairs across Go, Python, Java, TypeScript and Rust before capture. Positive witnesses span four to seven calls; conditional calls and overloads are explicit. These remain known development repositories.
Both native interfaces return valid six-call Click chains through different branches. Graphify's stdin/reader route was source-reviewed after capture and accepted, with no preregistered occurrence credit. The ID arm includes cross-interface interoperability failures and must not be presented as pure path quality.
Compass's jsoup and WalkDir outputs use structural shortcuts instead of complete call chains. Post-capture graph diagnostics find the registered call edges already present in Compass, identifying route selection as the next production target. Neither tool exposes Redux's named combination function at the frozen coordinate. No graph-oracle IDs replace failed public lookups.
All 204 benchmark tests and the product boundary pass. Twenty resolver calls, 42 CLI requests, all 18 depth diagnostics, source/graph/binary hashes and alternate occurrences are retained and separately verified; repeated verification is byte-identical. The review preserves failed verifier attempts and corrected occurrence-credit bookkeeping. No production changes or native Rust/JS/platform/packaging rerun in this checkpoint. Version stays 0.3.30.
Evidence: longer_path_registration.json, longer_path_witnesses.json, longer_path_label_control_registration.json, longer_path_review.json; raw external longer-paths-01. Next: explicit calls-only trail policy, then a separate named-function-expression identity fix. Broader superiority, authored explanations, genuine god-object judgments, functional community quality and held-out confirmation remain unproven.
Calls-only trail correction
node --calls-only, MCPget_node, and explicitaskcall-path syntax. Preserve structural defaults, ambiguity handling, occurrence evidence and bounded work.benchmarks/agent_query/calls_only_trails_review.jsonand the audit document. No equal-feature Graphify comparison, held-out or overall superiority claim. Version remains 0.3.30; PR was draft at that checkpoint.Shared public-neighbor path evaluation
benchmarks/agent_query/public_neighbor_path_review.json, registration, harness, tests and audit document. Authored explanations, god-object judgments, functional communities and held-out superiority remain open.Externally rated Blob cohort
SQL prefix guard and validation checkpoint
sql_prefix_scan_registration.json,sql_prefix_build_context.json, andmlcq_admission_followup_registration.jsonunderbenchmarks/agent_query; external validation runsql-prefix-scan-01. Version remains 0.3.30; no merge or release.Completed SQL diagnosis and retained full-repository outcome
benchmarks/agent_query/sql_prefix_scan_review.json, plus the audit and performance documentation. Rechecked source selection, unchanged copies, binary/build/log hashes, exact alternating order, raw process timings, RSS and every successful graph. Repeated verifier digest:28a7ee6b8e16ec4c4a4485559eff152318b3e2f5329ea31a3e8ddc87eca3a19b.Observed residual SQL cost
State::add_access_matchesandRegex::new. Source inspection confirms that fixed read/write patterns are recompiled per statement. This short snapshot does not attribute the entire run or establish an optimization result.benchmarks/agent_query/sql_regex_cache_registration.json, reusing all four complete SQL files and retaining all failures and graph-equality requirements. The copied release baseline now matches an unchanged release rebuild byte-for-byte. Cache implementation and release timing results remain pending.Completed larger-artifact Blob follow-up
SimpleValuePropertyhas more unfiltered neighbor records than either reviewed positive Eclipse class in both tools, illustrating the limitation of connectivity as a responsibility measure.relationFilteras an empty string, while the mock and auditor expected null. Two regressions failed before the fix. Retain the first invalid verification, every raw capture and unchanged primary ranking outcomes. Corrected verification repeats byte-identically (a659988abe0c9e64449b8265f06e066425fb6963e340176ccfde8f0a1e258d3d). All 27 tool calls are checked; 238 benchmark tests and product boundary pass.benchmarks/agent_query/mlcq_admission_followup_review.json, original review coverage disclosure, native-contract auditor regression and the audit document. Native product remains the tested prefix guard, version 0.3.30. God-object classification, authored explanations, functional communities and final independent confirmation remain open.SQL regex-cache release comparison
sql_regex_cache_review.jsonretains every outcome and artifact digest.sql_regex_cache_whole_registration.jsondeclares the separate complete-source follow-up, now completed under the original 1,200-second extraction bound and declared 1 GiB artifact cap (results below). Both earlier timeouts remain unchanged. God-object classification, authored explanation quality, functional communities and independent final confirmation remain open.Complete-source release and Blob audit results
sql_regex_cache_whole_review.jsonandmlcq_release_followup_review.json. Repeat registration:sql_regex_cache_whole_repeat_registration.json. No production code changed after the validated cache commit; version remains 0.3.30. Actual god-object diagnosis, explanation quality, functional communities and independent final confirmation remain open.Complete-source timing repeat
sql_regex_cache_whole_repeat_review.json.Java enum-reference census follow-up
Production
a8b4a490preserves enum constants, constant-specific body ownership, lexical shadowing and selector-derived switch labels. Source types retain precedence over imports; ambiguous or unsupported receivers remain unresolved. AST cache semantics advance from 9 to 10; package version remains 0.3.30.Full provenance and limitations: enum follow-up report and audit. These known development results do not establish held-out performance, authored answer quality, god-object classification or overall superiority. PR was draft at that checkpoint; no merge or release.
Java dotted nominal-owner checkpoint
Production
be2fb542, reviewb54c20ee: joins established dotted Java nominalreceivers to canonical nested declarations, retaining ambiguity, exact source
precedence and bounded lookup. The unchanged compiler-backed jsoup census now
verifies 3,115/3,785 ordinary field and 368/529 enum references. All 3,461
previously verified occurrences remain; all 3,483 returned in-cohort contacts
match the compiler. There are still 670 ordinary and 161 enum misses. Frozen
Graphify returns zero of these contacts, with no positive-contact precision
denominator. These are known development results.
All eight other-language controls reproduce complete baseline/original graphs.
Of 49 added references, 27 are outside the compiler cohort (21 test, six Java 11
overlay) and remain semantically unscored. All 75 existing community task-pair
outcomes remain unchanged despite a 40-to-43 jsoup community count change.
Reusing all 206 v10 AST cache entries yields the fresh candidate graph exactly.
Census/control/community verifier repeats agree byte-for-byte.
Formatting, focused/workspace Clippy, 536 resolver tests, 1,109 workspace tests
(two existing ignored), nine product tests, full code-graph fixture gate, 238
benchmark tests, product boundary and release build pass. Suites overlap.
Standalone browser/platform/packaging matrices were not rerun. Version stays
0.3.30; the PR was draft at that checkpoint. Overall superiority, god-object diagnosis, authored
explanation quality and functional community quality remain unproven.
Snapshot-coherence checkpoint
Commit f48ccad binds compact, full and typed JSON graph views to one bounded content snapshot and rebuilds disposable caches under digest-based format markers. Package version remains 0.3.30; no historical realization or graph schema changes.
Full cache review includes per-query cold/warm measurements, response equality, resource gaps and limitations. This correction does not establish improved graph source truth, god-object diagnosis, authored explanations, functional communities or overall Graphify superiority. PR was draft at that checkpoint; no merge or release.
Core inference fixture follow-up
Commit ffe50a2 makes Max inference explicit in the seven semantic-preservation integration fixtures (12 build configurations), preserving every assertion and the public Low default. The same seven failures were recorded on the unchanged baseline; all 13 core integration tests now pass. The test-only follow-up changes no production Rust input or frozen cache-evaluation binary. Formatting, focused Clippy, workspace Clippy, and workspace library/binary tests also pass under pinned offline parser inputs. Fixture review.
Class-span diagnostic
Registration 3ae08aa preceded the frozen-graph scan; review 6b3f692 preserves the negative outcome. Across CloudStack and Eclipse, the exact file/name/span rule uniquely joins 48/55 reviewed classes, including all eight consensus primary cases. Seven uncertain secondary cases remain unavailable. Ranking all 13,001 source-located class nodes solely by inclusive source span retrieves 0/4 consensus-positive classes at cutoffs 10, 50 and 100; the positive ranks are 131, 411, 426 and 714. This offers no improvement over the original top-100 hub retrieval and no god-object diagnosis. The first bounded scanner failure and corrected repeat verification are retained. No Compass or Graphify extraction/query was rerun. Complete review.
Full-class relation evidence inventory
Registration
19877907preceded a bounded read-only scan of the same frozen CloudStack and Eclipse Java graphs. Reviewb97ded50verifies all 13,001 class nodes, direct class/member ownership and method call/field-reference coverage, with all 55 prior review rows retained (48 unique joins, seven unavailable; all eight consensus cases uniquely joined). CloudStack has 84,998 uniquely owned members, Eclipse 69,878; neither graph has a member with multiple distinct direct class owners. Both have unowned members and unresolved relation targets.All eight consensus cases have zero observed foreign-field
referencesunder the registered rule. CloudStack also has 1,042readsand 6,161writesedges outside that rule; Eclipse's older producer has no such categories. This is stored-graph evidence, not independent source truth or absence of foreign-state access. No WMC/ATFD/TCC score, god-object classifier, new Graphify result or held-out win was computed. Two verifier runs agree byte-for-byte. Registration and review.CloudStack access-relation boundary
Registered
990d0286and verified46d5c98f: all 1,042readsand 6,161writesedges in the frozen CloudStack graph join SQL query/view/procedure nodes to database tables. There are zero method-to-fieldreads/writesedges, so these categories cannot repair the Java foreign-field evidence gap. Five endpoint-kind pairs account for every edge; two verification runs agree byte-for-byte. The original 55 review joins and frozen graphs are unchanged. Registration and review. No source-truth, classifier or Graphify win is claimed.Rated Java class member source check
Registration
ebf393e2fixed the eight consensus-rated files and frozen graph hashes before a JDK 17 syntax parse. Reviewc989292dindependently matches 95/95 direct methods and 65/65 direct fields to exactly one Compass class member by kind, normalized display name and source range; there are no extra graph methods or fields in these eight classes. Five explicit source constructors are reported but outside this registered graph comparison. The first literal-name comparison failed for all methods because Compass displaysfooas.foo(); both raw runs remain recorded. Repeated oracle and verifier outputs agree byte-for-byte. One positive-rated class is a test class. This checks declaration coverage in known Java cases, not call/field target correctness, cohesion, god-object classification or held-out quality. Registration and full review.Explicit Java field-token source check
Registered
29208f03; review6086f7f1uses a separate JDK syntax parser and exact UTF-8 byte anchors for explicitthis.fieldin the same eight rated classes. Five direct-method tokens qualify and all five have exactly one matching Compassreferencesedge to the source-matched field, with no competing edge at that token. Four constructor tokens and onethis.getClass()invocation are retained outside the denominator. This is a tiny selected subset, not general field-reference accuracy, a cohesion measure or a Graphify/god-object win. The initial collector preflight failure and repeated byte-identical oracle/verifier outputs are recorded. Registration and review.Compiler-bound direct-field census
Registration
a523de10expanded the eight rated Java classes to every direct-method field token that a public JDK 17 compiler could bind to a directly declared field. Reviewb4a499d1verifies 265/265 compiler-bound own-field sites against exact Compass method, field and byte anchor; all 265 stored own-field records in these selected classes are represented, with no extra record at those sites. The five previous explicit-thiscontrols retain their edge IDs. Three same-named candidates bound to nonfield symbols.The compiler intentionally lacked project dependencies and emitted 1,047 errors. An initial 1,000-error cap was corrected within the registered 10,000-diagnostic bound; candidate bindings were unchanged. This is selected own-field evidence, not recall for unresolved/inherited/foreign fields, god-object diagnosis, held-out quality or a new Graphify comparison. Registration and full review.
Source-first three-language hub checkpoint
Registration
4d5ca44dpinned six source-located Go/Python/Rust declarations, the clean source commits, frozen graph hashes, both server environments and publicgod_nodes(top_n=100)before any hub call. Commitdc704ca2adds the bounded collector, independently replayed review and audit account. All six public MCP calls succeeded; the two review runs are byte-identical. Raw transcripts remain in external runsource-first-hubs-01.Compass retrieves 2/6 selected witnesses in the top 10, 3/6 in the top 50 and 3/6 in the top 100; Graphify retrieves 1/6, 3/6 and 4/6. Both return
FastAPIat rank 1; Graphify returns CelldPeerAuthat rank 55 while Compass leaves it outside the top 100. Compass has explicit graph IDs, matching source anchors and graph-consistent degrees for all 300 returned rows. Graphify emits labels only: 247/300 are globally unique and all of those have graph-consistent degrees; ambiguous labels remain unresolved. The selected source roles are not god-object defect labels, and the extracted graphs differ. No classifier accuracy, superiority, merge or release claim follows. Package version stays 0.3.30.Rated Java class-only degree diagnostic
Registration
3ebdf975froze the same class-only distinct endpoint-pair degree rule and four existing CloudStack/Eclipse graph hashes before computing ranks. Commit52450894adds the bounded scanner, review, and audit account. The 16 exact graph identities for eight externally rated classes all remain eligible. Two scans produced byte-identical results.Neither tool retrieves any of the four consensus-positive Java classes at top 10, 50 or 100 after filtering to source-located classes. Compass positive ranks are 340, 728, 999 and 1,529; Graphify's are 463, 1,315, 1,403 and 1,740. Neither retrieves a consensus-none class at those cutoffs either. The graphs differ in extraction coverage and direction, and this is known development data, not held-out classifier evaluation. Class-only degree cannot be claimed as a repair for god-object detection. No production code or package version changed.
Paired rated-Java member declaration coverage
Registration
3996067cfroze the independent JDK syntax census of eight rated Java classes (95 direct methods, 65 direct fields), source hashes, exact class IDs and four graph hashes before this comparison. Commitdd028419adds the bounded scanner, byte-identical repeated review and audit account.Under the registered exact file/name/source-span join, Compass has unique directly owned nodes for 95/95 methods and 65/65 fields. Graphify has 90/95 methods and 0/65 field-compatible declaration nodes in these files. Its five method misses are later overloads whose names share a node anchored at the first overload; some ownership edges cite later overload lines without distinct identities. Graphify still has field-context references, so this does not claim it extracts no field information. This selected development result supports declaration coverage only. It does not establish field-access correctness, class cohesion, god-object diagnosis, whole-repository accuracy or broad superiority. No production code or package version changed.
Stored hub member evidence
Commit
ee21fae8addsmemberEvidenceto directed, typed class/struct rows in public MCPgod_nodes. It counts uniquely owned stored methods, fields and method-to-own-field reference records; ambiguous ownership is excluded.sourceCoverage: "unverified"and the contract note make clear that stored reference records include every confidence state. Hub eligibility, degree and order are unchanged; this is not god-object classification.A two-pass replay on frozen Litestream, FastAPI and Celld graphs confirms all 300 prior hub rows are unchanged after removing the additive field, with identical new results between passes. Member evidence appears in 33, 34 and 26 of their respective top-100 rows. Celld
PeerAuthremains outside Compass's top 100. A separate frozen Eclipse top-5,000 read finds three rated Java classes whose selected direct-member and own-field counts match the prior independent source/compiler checks; the fourth rated class is outside the request. The Go/Python field zeros and selected Java checks do not establish cross-language completeness, reference precision, cohesion or broad superiority. Reports and replay tool are inbenchmarks/agent_query/; raw responses stay in the external evaluation run.Focused graph/MCP regressions, workspace formatting, Clippy and library/binary tests, CLI product tests, product boundary and fixtures-only code-graph qualification passed. Package version remains 0.3.30; PR was draft at that checkpoint.
Go struct field declaration census
Go direct receiver-field access follow-up
Commit
5553c8c5adds source-anchored Go field contacts and resolves exact owner-qualified field declarations before their type aliases. The preregistered Litestream source oracle has 2,742 direct receiver-field sites; all 2,742 join to unique method-to-field references at the exact source byte in the candidate graph, versus zero in the frozen Compass and Graphify graphs under this field-target join. All 4,650 added member-access links target fields, and no old node ID or relationship was removed. The 1,908 added links outside this narrow oracle remain uncredited and require separate precision review.The audit graph exceeded its preregistered 16 MiB read cap, so the review records a post-capture increase to 24 MiB without filtering records. The cold and warm graph bytes and two full probe reports are identical. Public
god_nodes(top_n=500)shows 113 own-field reference records forStoreand 13 forHeartbeatClient; these are stored contacts, not god-object labels.Focused Go, resolver and cache tests, workspace format/Clippy/lib/bin tests, and the product-boundary check passed. After mounted-volume capacity was restored, the full fixtures-only gate passed, including its release-mode React phase. An initial resumed run stopped on a missing snapshot during lifecycle checks; the complete diagnostic repeat passed, so that intermittent failure remains open for investigation. Broad superiority and god-object diagnosis remain unclaimed. Compass stays at 0.3.30 and the PR was draft at that checkpoint.
Typed Go parameter-field audit
Registration
d82d94b9and reviewc1b0d78findependently verify 276/276 explicitly typed method-parameter field contacts on the pinned Litestream source at exact graph sites. The frozen pre-access Compass graph has no matching links; Graphify has no exact declaration targets for the 98 fields, though it retains 19 field-type context records. These 276 links are disjoint from the prior 2,742 direct-receiver links: 3,018/4,650 added Go member-access links now have narrow source-oracle credit, and 1,632 remain uncredited. Both full source and graph probes were byte-identical across two runs. This remains one known Go development repository and does not prove whole-graph precision or broad superiority.Final gate and Windows test follow-up
Commit
bfe9068crecords the complete local fixtures-only gate and adjusts the graph-artifact snapshot test to use the path-rename sequence already exercised by a passing Windows test. Both Windows CI architectures failed the earlier overwrite attempt withAccess is denied; the revised test retains the opened-handle snapshot assertion. Local formatting,compass-modeltests, workspace Clippy, and workspace library/binary tests pass on the revised patch. CI for that head was pending at that checkpoint.