Skip to content

refactor: build two-phase System One progressive code reader - #17

Closed
BestNathan wants to merge 277 commits into
mainfrom
fix/code-locator-three-stage-requests
Closed

BestNathan wants to merge 277 commits into
mainfrom
fix/code-locator-three-stage-requests

Conversation

@BestNathan

@BestNathan BestNathan commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Summary

Rework the System One Code Locator around a two-phase architecture instead of a fixed directory/file/line-or-symbol pipeline.

Phase 1 — File Locator

  • enumerate the repository directory tree once;
  • preserve directory pruning by exposing only direct files of retained directories;
  • raise the default directory/file thresholds to 0.50 / 0.65;
  • cap the handoff to the progressive reader at 16 files by default;
  • keep source bodies out of Phase 1.

Phase 2 — Progressive Reader

  • read file stat first and generate a bounded ReadRange action space;
  • let System One choose where to read instead of pre-reading / pre-expanding the whole file;
  • process multiple files concurrently: one Choice question per active file, four files per batch by default, all questions in one System One request;
  • append read content to ReaderState before scoring it;
  • score each new observation with Noul;
  • use observation scores and coverage to generate the next action frontier;
  • only high-relevance observations create local before/after expansion actions;
  • keep separate reader-action and observation thresholds.

Why

The original pilot made 598 model calls because the harness batched large candidate arrays and scored line frontiers per file. Later line/region/symbol variants still pre-expanded too much file state before the model had chosen to inspect it.

The new invariant is:

State
  -> grounded ActionSpace
  -> System One decision
  -> read_file effect
  -> Observation
  -> State transition

Source content is now an observation produced by a selected action, not a precomputed candidate set.

Real TypeSafe baseline

Run 35828994234 against BestNathan/nession staging completed successfully with:

  • 380 directories exposed -> 40 retained;
  • 159 direct files exposed -> 16 handed to the reader;
  • 4 files per reader batch;
  • 38 reads / observations;
  • 27 evidence observations;
  • 29 semantic model calls;
  • 340,395 input tokens;
  • 13,205 output tokens;
  • 7.20s end-to-end runtime.

Examples from the trace show the desired behavior: high-scoring websocket ranges cause contiguous local expansion, while low-scoring small wrapper/index files stop after the initial probe.

Research follow-ups

  • Phase-1 threshold calibration;
  • reader-action vs observation threshold calibration;
  • reader batch-size study;
  • Noul vs Choice comparisons at equivalent semantic decision points;
  • richer observation-driven action generators;
  • gold relevant-file/range benchmark.

Implementation, tests, workflow controls, design notes, and the two-phase Nession baseline are included in this PR.

Multi-hotspot follow-up

The first two-phase trace exposed a harness defect: only the single strongest relevance hotspot could expand, so a second relevant region inside the same file could be discovered but not explored locally.

This PR now preserves up to three disconnected RelevantRegion hotspots per file and generates bounded before/after actions for each.

Real TypeSafe run 35831659317 validates the fix on crates/nession-agent/src/server/websocket.rs:

  • second hotspot: 1585-1724, relevance 0.70;
  • next frontier now included 1445-1584 and 1725-1864 around that hotspot;
  • System One selected 1445-1584 with probability 0.62;
  • the resulting observation scored 0.78 relevance.

The previous single-hotspot baseline instead had no neighbor action for the second hotspot and jumped to an unrelated gap that scored 0.22.

Research notes:

  • research/code-locator/docs/pilots/nession-websocket-two-phase-trace-analysis-2026-09-23.md
  • research/code-locator/docs/pilots/nession-websocket-multi-hotspot-reader-2026-09-23.md

Evidence-driven reader budget

The fixed four-round reader could stop immediately after discovering a useful new hotspot on round four. The reader now uses a soft/hard budget:

  • soft rounds: 4;
  • hard rounds: 8;
  • each file must independently produce a new observation above the observation threshold to earn another post-soft-limit round.

Real run 35833091673 first validated the soft/hard mechanism at batch scope: server/handler.rs found a 0.70 hotspot on round four and received a round-five local expansion instead of terminating.

That trace also exposed sibling over-expansion. The continuation rule was therefore refined to file scope. Real run 35833457198 validates the refinement:

  • cli/client/connection.rs scored 0.64 on round four and stopped immediately;
  • server_client.rs independently continued with 0.82 / 0.81 / 0.70 through rounds four to six, then stopped after a 0.25 round-seven observation;
  • server/handler.rs still earned round five from its round-four hotspot.

Research notes:

  • research/code-locator/docs/pilots/nession-websocket-soft-hard-round-budget-2026-09-23.md
  • research/code-locator/docs/pilots/nession-websocket-file-scoped-round-budget-2026-09-23.md

Claude Code accuracy reference

A new manual dual-run workflow compares the System One locator against Claude Code using the existing ds environment on the same exact Nession revision and verbatim task.

Reference run 35834529068 (BestNathan/nession@b76fe4921a63023a69ce91399328ab53d3526664) found:

  • Claude reference: 5 primary, 9 supporting, 4 context files;
  • System One primary-file recall: 4/5 = 80%;
  • System One primary evidence-region recall: 16/20 = 80%;
  • primary evidence-line coverage: 85.8%;
  • final evidence precision proxy vs Claude reference: 75%;
  • System One runtime 12.71s vs Claude Code 67.19s in this run.

The main primary miss was web/src/platform/socket/MessageRouter.ts: System One scored it 0.64 at rank 21, below the 0.65 threshold and outside the top-16 frontier. However, the already-observed WebSocketService.ts:1-140 imports ./MessageRouter directly. This indicates the next architectural accuracy improvement should be dynamic cross-file actions discovered from observations, rather than simply lowering Phase-1 thresholds until transitive dependencies fit.

Accuracy study:

  • research/code-locator/docs/pilots/nession-websocket-claude-reference-accuracy-2026-09-23.md
  • .github/workflows/system-one-code-locator-accuracy.yml
  • research/code-locator/src/compare_claude_reference.py

Global file scheduler + Claude execution trace

The Phase-2 reader no longer uses a Phase-1 top-k cap or fixed file batches.

  • every file above the Phase-1 file threshold enters one shared ReaderState;
  • every scheduler epoch scores all eligible files in one System One request;
  • System One decides which subset should receive reads now;
  • unselected files remain in state and may become active later;
  • read budgets are owned per file, not per batch or scheduler epoch.

Real run 35839323377 retained 18 Phase-1 files and selected only three in its first global scheduling request. It completed with 12 model calls, 8 reads, 3 unique files, and 3.78s elapsed. This is recorded as behavior, not judged against another model's final file set.

The Claude Code + ds side of the same run now preserves its full execution trajectory:

  • claude.raw.jsonl — complete stream-json events;
  • execution-path.json — normalized ordered tool calls, inputs and result metadata;
  • execution-summary.md — human-readable path;
  • reference.json — final localization record.

Claude used 42 tools (24 Read, 2 Grep, 16 Bash) across 43 turns, moving from broad websocket search into connection/session/router/handler dependencies. Claude is explicitly treated as a separate System-2 trace, not ground truth or an optimization target.

Research note:

  • research/code-locator/docs/pilots/nession-websocket-global-scheduler-cross-trace-2026-09-23.md

Canonical localization result

System One and Claude Code now emit the same execution-independent code-localization-result contract.

Each result contains:

  • final valuable files;
  • file role and confidence;
  • valuable source ranges;
  • per-range confidence and reason;
  • exact source content for each range;
  • confidence provenance/semantics.

System One confidence comes from Noul relevance and derived evidence aggregation. Claude confidence is model self-assessment. The schema preserves these different semantics instead of assuming calibration.

Execution traces remain separate from final results. The generic comparator now consumes only the two canonical localization-result.json artifacts and reports symmetric file/evidence overlap without designating either side as truth.

Real validation run 35844251415:

  • System One: 3 valuable files / 8 evidence regions;
  • Claude Code: 18 valuable files / 61 evidence regions;
  • shared files: 3;
  • union files: 18.

Contract:

  • research/code-locator/docs/localization-result.md
  • research/code-locator/src/localization_result.py
  • research/code-locator/src/compare_localization_results.py

Two-session Claude confidence + cost contract

Claude Code now separates localization from confidence:

  1. Session A performs repository localization only and is forbidden from emitting confidence.
  2. The Session-A draft is normalized and its evidence source is materialized from the pinned revision.
  3. Session B is a fresh Claude Code session with repository tools disabled. It only assigns overall/file/evidence confidence.
  4. The finalizer rejects any Session-B change to file paths/order or evidence counts/ranges.

The canonical localization result also carries cost metrics: elapsed time, model calls / turns / tool calls where available, input/output/cache/thinking tokens, provider USD cost when reported, and per-stage breakdowns.

Real validation run 35846265667:

  • Claude localization: 13 files / 42 evidence regions / 0 confidence keys;
  • Claude confidence session: 0 tool calls, no file/range changes;
  • Claude localization cost: 67.864s, 51,398 input, 11,575 output, 554,368 cache-read, $0.823549;
  • Claude confidence cost: 24.424s, 32,119 input, 5,616 output, $0.300995;
  • Claude combined: 92.288s, 33 turns, 31 tool calls, $1.124544;
  • System One canonical result in the same run: 3.770s, 15 model calls, 249,871 input, 13,009 output, 10 reads.

Contract: research/code-locator/docs/localization-result.md

Range-only System One Runtime v0

A separate v0 prototype now reduces Phase 2 to a content-agnostic state/action loop:

State -> geometry-only ActionSpace -> System One Noul scores -> threshold-based parallel selection -> ReadRange effects -> raw observations -> new State.

The only runtime effects are ReadRange and StopTask. Navigation labels (seed_head/middle/tail, expand_before/after, jump) are generated from line-count / coverage geometry only. No AST, LSP, symbols, imports, keywords, semantic chunks, or Harness content analysis are used.

Selection policy:

  • Stop only when StopTask >= threshold and Stop is globally highest;
  • otherwise execute every read action at or above the parallel threshold;
  • when no read exceeds the threshold, execute the highest-scoring read.

Real run 35855369932 on BestNathan/nession@b76fe4921a63023a69ce91399328ab53d3526664 demonstrated that range-only navigation can jump directly to distant useful areas. For crates/nession-agent/src/server/websocket.rs, epoch 1 selected 1-140 (.81) and 1445-1584 (.67); epoch 2 then selected 1305-1444 (.77) and 1585-1724 (.79) plus additional jump probes.

The run also exposed the next runtime issues:

  • StopTask never exceeded .23 and the run ended at the 64-epoch safety budget rather than model stop;
  • threshold .65 selected 23 actions in epoch 1, so it is a real concurrency/resource-control parameter;
  • candidate ReadRanges can overlap within one epoch and need geometry-only conflict elimination;
  • durable RuntimeState must be separated from a bounded DecisionView. A 48k-character recency projection of raw observations fixed Jev context-limit failures without semantic summarization.

Pilot: research/code-locator/docs/pilots/nession-websocket-range-runtime-v0-2026-09-23.md

Per-file range runtime v0

Phase 2 has been redesigned around independent file-local state machines after confirming the System One request context constraint makes one shared multi-file raw-observation state a poor fit.

Current v0:

  • Phase 1 selects plausible files;
  • every selected file gets its own FileRuntime;
  • runtimes share only the immutable user goal;
  • action generation is range-only and content-agnostic (head/middle/tail, expand_before/after, jump, StopFile);
  • the Harness uses no AST/LSP/symbol/import/keyword parsing;
  • the .65 action threshold controls concurrency only;
  • if no read reaches the threshold, the best read still executes;
  • only explicit StopFile can produce model_stop; low scores never stop the Harness;
  • overlapping above-threshold reads are deterministically deconflicted;
  • final evidence scoring happens after navigation and cannot affect the loop.

Real validation run 35857120750 on the Nession websocket task:

  • 17 Phase-1 files / 17 independent FileRuntime instances;
  • 138 reads / 145 model calls / 33.243s;
  • 11 final valuable files / 46 evidence regions;
  • 16 files reached 100% coverage and ended via action_space_exhausted;
  • handler.rs reached 94.2% coverage before the 32-epoch safety budget;
  • 0 early model_stop decisions.

The run validates the per-file context architecture and also shows the next research target clearly: current StopFile semantics are conservative, so System One tends to keep exploring until coverage is exhausted rather than declaring sufficiency early.

Research trace:

  • research/code-locator/docs/pilots/nession-websocket-per-file-range-runtime-v0-2026-09-23.md

StopFile control-choice experiment

Prompt-only StopFile experiments were not enough. Treating StopFile and ReadRange as comparable Noul scores biased the runtime toward exhaustive scanning because they represent different semantics.

The current FileRuntime now sends one mixed System One request per epoch:

  • Choice: StopFile vs ContinueFile — control-flow sufficiency;
  • Noul: one utility score per ReadRange — information-gathering value.

Policy:

  • if Choice selects StopFile, terminate the file runtime;
  • otherwise execute all non-overlapping ReadRanges >= 0.65;
  • if no ReadRange reaches 0.65, execute the top-1 ReadRange;
  • the 0.65 threshold is therefore only a read-concurrency threshold;
  • if no ReadRange exists, terminate mechanically as action_space_exhausted without asking the model.

StopFile is explicitly defined as representative-result sufficiency, not full-file coverage. The goal is to return enough representative valuable ranges to characterize the file, not exhaustively enumerate every relevant range.

Clean run 35864316780:

  • 17 FileRuntime instances;
  • 4 genuine early model_stop decisions;
  • 12 action_space_exhausted;
  • 1 budget_exhausted;
  • 134 reads / 130 model calls / 25.952s;
  • early-stop coverage: server websocket 91.3%, WebSocketService 77.5%, CLI connection 86.0%, web client registry 99.3%.

Compared descriptively with the earlier per-file baseline, those four common files used 24 reads before and 17 reads now. Evidence-line overlap ranged from 0.775 to 1.000; the runs are stochastic, so this is not treated as a controlled accuracy metric.

The remaining open stop problem is concentrated in very large files (agent/server/websocket.rs, server_client.rs, handler.rs), where System One still prefers Continue until exhaustion or the safety budget.

Two experimental System One localization algorithms

Two algorithms discussed separately were implemented as isolated experiments without replacing the current per-file range-runtime baseline.

A — Explore-Guided Evidence Filtering

  • reuses the existing FileRuntime state/action-space mechanics;
  • changes ReadRange scoring from topical/useful relevance to NEW MATERIAL EVIDENCE;
  • explicitly penalizes redundancy, wrappers, logging/debug plumbing, and repeated facts;
  • Stop/Continue is framed around closing concrete evidence gaps;
  • post-loop filtering keeps a minimal evidence set.

Real run 35882424922 on BestNathan/nession@97b9d2b49c5137064e910fc180aabc01cfb1021f: 17 Phase-1 files, 135 reads, 8 valuable files, 32 evidence regions, 164 model calls, 1.83M input tokens, 40.84s.

B — Adaptive Semantic Zoom Search

  • stratifies each file into 16 coarse regions and reads 32-line midpoint probes;
  • uses one Choice distribution over the region frontier;
  • keeps up to top-3 regions until cumulative probability mass reaches 0.80;
  • reserves one geometry-only exploration slot (~75/25 exploit/explore);
  • recursively splits selected regions and probes children;
  • converges when high-probability regions are <=40 lines and dominant roots are stable.

Real run 35881894374 on the same subject revision: 17 Phase-1 files, 417 probes, 12 probability-frontier convergences, 5 round-budget exits, 11 valuable files, 194 evidence regions, 184 model calls, 1.02M input tokens, 43.36s.

Manual A/B workflow: .github/workflows/system-one-localization-algorithms.yml.

Design note: research/code-locator/docs/system-one-localization-algorithms.md.

Copy link
Copy Markdown
Owner Author

Implemented the immediate follow-up fixes from the PR #17 review.

StopFile / large-file exploration

  • Removed unconditional low-score fallback_top1 from the per-file range runtime.
  • Added explicit control/action reconciliation for the two contradictory states:
    • StopFile while a concrete ReadRange is still >= the parallel threshold.
    • ContinueFile while every concrete ReadRange is below the threshold.
  • Conflict resolution is still model-driven: System One chooses between finalizing the file and executing the displayed best concrete read.
  • A below-threshold read can only run after explicit reconciliation authorization.
  • DecisionView now includes bounded exploration-yield telemetry:
    • recent control choice;
    • max concrete read utility per epoch;
    • selection mode;
    • consecutive low-utility epoch streak.
      This exposes diminishing returns without Harness source-semantic analysis.

Canonical cross-trace

  • The System One side of .github/workflows/system-one-code-locator-accuracy.yml now runs system_one_range_runtime.py, so StopFile experiments and Claude comparison use the same System One algorithm.
  • The range runtime now emits canonical localization-result.json directly.
  • Canonical results can carry frozen subject.repository + subject.revision.
  • The generic comparator now fails closed when either subject identity is missing or the revisions differ.

Quality-evaluator regression

  • Added regression coverage for evaluator output containing prose / fenced JSON instead of a bare JSON object.
  • Added an explicit failure-path test when no valid quality-evaluation object is present.
  • This covers the parser class that caused run 35869966400 to fail during scorecard finalization.

Validation

The automatic offline code-locator validation passed on the code head after these changes (run 35886450064, including unit tests and deterministic fixture). The final documentation-only commits triggered a fresh CI cycle and are currently queued/running.

The remaining important experiment is a fresh manual System One / Claude cross-trace on the updated workflow, followed by the blind quality evaluator. That run should be used to measure whether fewer large-file reads preserve downstream localization quality; older cross-trace overlap numbers came from the previous progressive-reader algorithm and should not be used for that judgment.

Copy link
Copy Markdown
Owner Author

Fresh end-to-end cross-trace completed successfully after fixing the confidence-session schema normalization.

Valid run

  • Run: 35890516881
  • Harness commit: 475355202066fdc2268f5cf1c7b67be146c5a0b4
  • Frozen subject: BestNathan/nession@7ac9b6e0c2bb43c52f83e7dd706c0c0dc0d7a1df
  • System One, Claude canonicalization, symmetric comparison, and blind quality evaluation all succeeded.

StopFile reconciliation repeatability

Three fresh System One executions on the same frozen subject/task after the reconciliation change produced:

Run Reads Model calls Valuable files Evidence regions Elapsed
35889657137 94 96 11 42 25.4s
35890090004 86 88 11 40 23.9s
35890516881 96 96 12 47 22.7s

Mean: 92 reads / 93.3 model calls / 24.0s.

For context, the earlier StopFile clean run 35864316780 used 134 reads / 130 model calls and still exhausted the large files. It was a different subject revision, so this is descriptive rather than a controlled pre/post A/B, but the new behavior is stable across three identical-subject reruns.

Latest large-file behavior:

  • server/handler.rs: 24 reads, 48.8% coverage, model_stop.
  • agent server/websocket.rs: 16 reads, 73.9% coverage, model_stop.
  • connection/server_client.rs: 16 reads, 78.4% coverage, model_stop.
  • No file hit budget_exhausted.

The previous exhaustive tail is therefore no longer the dominant termination mode for the large files.

Current range runtime vs Claude Code

Same frozen revision, same task:

  • System One: 22.743s, 96 model calls, 96 reads, 724,127 input tokens, 17,427 output tokens.
  • Claude Code: 96.347s, 40 turns, 38 tool calls, 75,469 input tokens, 21,045 output tokens, 624,384 cache-read tokens, $1.215662 provider cost.
  • Shared files: 8.
  • Union files: 17.
  • File Jaccard: 47.1%.
  • System One evidence overlapped by Claude: 48.9% regions / 20.5% lines.
  • Claude evidence overlapped by System One: 66.7% regions / 69.9% lines.

System One-only files include the agent server_client.rs, server web_client_registry.rs, useRealtimeUpdates.ts, and relay terminal server path. Claude-only files include MessageRouter.ts, socket/types.ts, visibility reconnect, relay adapter, and terminal agent path.

Blind quality review

The evaluator received one anonymized candidate per fresh session and the exact frozen repository revision.

  • System One: 71/100, can_proceed=true.
  • Claude Code: 83/100, can_proceed=true.

System One dimensions:

  • completeness 6.5
  • relevance precision 7.5
  • evidence grounding 7.0
  • redundancy efficiency 6.5
  • downstream actionability 7.0
  • organization/prioritization 5.0
  • risk/uncertainty coverage 10.0

The evaluator's main System One gaps were exactly the architectural direction suggested by the cross-trace analysis:

  • web/src/platform/socket/MessageRouter.ts
  • crates/nession-server/src/server/command_broker.rs
  • client_registry.rs
  • producer-side PTY/read loops
  • the missing reconnect/registration section inside server_client.rs
  • several observation-referenced dependencies.

The primary quality deficit is no longer inability to stop. It is now cross-file dependency discovery + evidence shaping:

  1. observation-driven FollowFile / dependency actions;
  2. task-specific evidence reasons instead of template reasons;
  3. primary/supporting/context prioritization;
  4. tighter evidence boundaries instead of fixed ~140-line retained windows.

Confidence-session infrastructure issue found and fixed

Two attempted runs exposed a real Claude confidence formatting behavior: all file assessments were present, but files[] was nested under overall. Raw artifacts confirmed this was a schema nesting error, not missing judgments.

The finalizer now performs one narrow lossless normalization only when top-level files is absent and overall.files is a list. The raw Claude artifact is also uploaded with if: always(), so future malformed outputs remain inspectable.

This run validates the entire canonical-result -> comparison -> blind-quality pipeline end to end.

Copy link
Copy Markdown
Owner Author

The System One code-localization research has now been extracted into the standalone repository BestNathan/system-one-code-explore.

The standalone repository contains the validated range runtime, canonical localization result, Claude Code-style cross-trace, blind quality evaluator, historical pilot records, and the next research roadmap.

Future algorithm work should continue there. The first follow-up is issue #1: observation-driven FollowFile actions for progressive cross-file state-space discovery.

Copy link
Copy Markdown
Owner Author

Closing this research branch as superseded by the repository split.

This PR remains the primary historical record for the post-#16 Code Locator evolution (two-phase reader, file-local runtimes, stop control, bounded DecisionView, canonical localization result, and alternative exploration algorithms), but it should not be merged into Narness.

Active implementation and ongoing experiments now live in BestNathan/system-one-code-explore. Narness keeps only the stable architectural conclusions and provenance; #26 converges the topic and points to the active repository.

@BestNathan BestNathan closed this Sep 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant