Skip to content

[codex] test one-hop witnessed STR-1 frontiers #2127

Description

@TheGreenCedar

STR-1: one-hop witnessed structural frontiers

ETR-1 validly selected no frontier at eb968e4f68e8999ce575ba25bee9d1c13b56cfdf. Its evaluation is 545a7d2ba1731d32bda7a4f1daf346d11b81f32bbbb732a3bb91673fad11c982; independent reconstruction is 2fc235d14d3477b8c7985d0a5078f2e4ab2e4a1c56c668a2d79b7d198a3e51b3. Both remain immutable. This contract exercises the latest user-authorized, single structural alternative. No third retrieval architecture is authorized after a valid failure.

Fixed inputs and authority

Stack the benchmark-only child on ETR head eb968e4f68e8999ce575ba25bee9d1c13b56cfdf, tree 8109976c4b3bede4a2c02961b71594be2fa89760. PR #2126 and its worktree remain the ETR receipt/replay record; PR #2120 remains unchanged. Integration head at registration is cd062c70527013fec221312537ca1bda3bb64519; no STR implementation or existing issue/PR was found there or in issue/PR title searches.

  • ETR preparation: 30b84d4d848f96bd4fe799f2e0f28b9114971da0e47bf98ebe54fe36242199fd
  • Fragment vectors: 7f604b30b823066bd5b0ed71106d10577c28495abd270444bc8ad5b7a63cb70a
  • ETR run: c14da697d03707c0096f5f2fd7a97bff2ab5b4a6f9326c4a9a03da2066d545f2
  • ETR validation: dd15c5842b68857c8c684b680cf3d1e307cd4b1a874736441c594490723e7145
  • Questions: 8e7219a59c973c02f8ea93120bb680da46a75b8272153986c76e55bfb73ca3b6
  • Annotations: 52b0cc223292bc70f1e4fa3f52b67bf42a91e4d4b9ed997aa12c648c068e9ade
  • Retained source-only oracle: db9abe6ce37ab5ddb7900049192e4712e359b5ad5587a9ca982c23796fd137ef
  • Model: 666db8df27c88570cdc07adca28646260038b8ca65354911d57b936ebf56efaa
  • Tokenizer: 7465b93c945b7a266481e6785aa13e505c625562c1c046c4b762bb4da4d46082
  • Model contract: cb0e3c00290f1eb21ecdcd873521d03331069b1efa766fcd1e493e6d4299b4b7

All ETR fragment identities, exact source, row costs, metadata, repository commits and 72 wording rows are retained. Graph inputs are the existing immutable core publications named by each retained witness preparation in definition-repair-source-c9c935d8-v1. Their complete database, pointer, preparation and source bindings are frozen before any candidate output. No rebuilding or repairing parsers/collectors against this corpus. A mismatched graph/source binding invalidates execution; absence of indexed coverage is a measured gap.

One fixed intervention

Control is the complete frozen ETR unconditioned frontier, including each S/D/H/L pool, score ordering and execution receipt. It is not replaced by the broad fragment oracle.

Candidate reconstructs and authenticates the same BM25 top-16 S with natural underfill. Before request timing it pins the named immutable core and builds read-only node/occurrence/incident indexes from that complete publication. It retains all source-bound graph metadata, not a question-derived file subset. The frozen fragment vector index is already prepared.

For each original seed in its BM25 order:

  1. Anchor at every indexed non-file node whose explicit source line span intersects that exact seed fragment, and at both effective endpoints of any eligible relationship whose occurrence intersects the seed. Missing spans remain missing. No names or kind-dependent prompt rules influence anchoring.
  2. Discover the incoming and outgoing incident relationships of those original anchors, once. A successor never becomes a discovery root. Every accepted edge has explicit certainty = Certain, effective endpoints present in that core, and an indexed occurrence with a nonempty authenticated source range. The edge's own explicit file/line occurrence is acceptable line-addressed provenance; otherwise an edge-owned reference occurrence supplies the range. No inferred certainty, synthesized relationship, shortest path or recursion.
  3. Map both endpoints' explicit source spans and the relationship occurrence to every intersecting exact frozen fragment. Preserve direction, edge ID, raw/effective endpoint IDs, occurrence, source digest and discovery root privately. Unknown/outside-universe endpoints and occurrences produce boundary gaps; never add source outside the retained universe or expand a fragment.
  4. Deduplicate by fragment ID. Exclude every S member and all previously retained candidate successors. Rank eligible fragments by normalized dot product with one freshly encoded unchanged raw question, then fragment ID only. Retain at most eight per original seed, at most 128 total and 144 including S. No cutoff, root requirement, quotas, name bonuses or relation credit.
  5. Authenticate all selected descriptor sources and record exact D/H/L. Every source row is independently selectable under the existing 16-row/16-KiB full projection. No source substitution or partial-atom credit.

Raw-question encoding uses the pinned production client and native completion/token evidence. Encode once per wording when seeds exist, with no truncation. The same encoded vector ranks all structural candidates in that wording. No host queries, source-conditioned query, graph-to-text document, new model, production packet, or selector call occurs.

Execution and timing

The real synthetic canary uses the graph export, native query encoder, structural runner, authoritative validator and source-only evaluator. It covers mixed line endings/UTF-8, more than eight candidates, multiple seeds with interleaving IDs, incoming/outgoing discovery, underfill, a second-hop trap, uncertain and unwitnessed edges, outside-universe endpoints, duplicated witnesses, cancellation and invalid inputs.

The supervisor binds clean commit/tree, binaries, graph inputs, method, preparation, native events, process configuration, terminal status and outputs independently before annotation access. The validator reconstructs graph/source mappings, query scores, exclusions and every pool. A rehashed self-description is not authority. Annotations remain unopened by this experiment until all 72 candidate outputs, frozen control joins and execution bindings pass validation. Any failed row invalidates the whole run; no selective aggregation.

Request timing encloses BM25, seed authentication, raw query tokenization/encoding, graph frontier construction, candidate scoring/mapping and remaining source authentication. It records exclusive intervals and unaccounted time. Graph/vector index preparation, output serialization, validation and oracle evaluation are separate. The control's previously validated timing remains labelled retained ETR timing; it is not a new balanced installed timing comparison. Candidate adoption requires measured absolute prepared-state p95 at most 1.25 seconds. A global 30-minute request-run deadline and cancellation yield invalid/not-evaluated, never partial quality.

Evaluation and terminal rule

Reuse the accepted exact source-only optimizer and reproduce its retained fixtures. Optimize each acceptable alternative separately over that arm's L only, charging exact row and complete metadata costs, at most sixteen rows and 16 KiB. No compulsory discovery roots, relation atoms or invented evidence.

Aggregate three phrasings within each case, then cases within each group/corpus. A newly gained atom counts for a case only in at least two phrasings; a control case is incomplete for that test only in at least two phrasings.

Structural sufficiency requires mean feasible source recall >=85%, every group >=70%, complete acceptable-set rate >=75%, and 100% source/address authentication. Material value requires >=10 percentage points improvement in recall or complete-set rate over the paired retained control, no regression in the other, no group loss beyond two points, and a new required atom in at least half of eligible incomplete cases. Quality is evaluated before latency. Only passing both quality gates and p95 <=1.25 seconds selects structural_frontier_selected. A valid quality failure is no_frontier_selected and stops this structural alternative and automatic packet program. A valid latency failure likewise cannot authorize selection; no new latency-repair allowance is created for STR-1.

Invalid execution permits only restoration of this frozen contract and a full rerun from zero. It cannot change questions, graph/parser coverage, representation, model, traversal, limits, thresholds or optimizer. One implementer and the existing independent verifier execute the complete hostile matrix and reconstruct the exact final result. No additional reviewers.

A pass authorizes preparation of a separate SEL-1 contract over the frozen pools only. It does not authorize selector implementation, public routing/schema/skill changes, new corpora, production integration, calibration, versioning or release. All prior stops and downstream product gates remain unchanged.

Refs #2116; Refs #2119; Refs #2125; Refs #2126.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions