refactor: build two-phase System One progressive code reader - #17
BestNathan wants to merge 277 commits into
Conversation
|
Implemented the immediate follow-up fixes from the PR #17 review. StopFile / large-file exploration
Canonical cross-trace
Quality-evaluator regression
ValidationThe automatic offline code-locator validation passed on the code head after these changes (run The remaining important experiment is a fresh manual System One / Claude cross-trace on the updated workflow, followed by the blind quality evaluator. That run should be used to measure whether fewer large-file reads preserve downstream localization quality; older cross-trace overlap numbers came from the previous progressive-reader algorithm and should not be used for that judgment. |
|
Fresh end-to-end cross-trace completed successfully after fixing the confidence-session schema normalization. Valid run
StopFile reconciliation repeatabilityThree fresh System One executions on the same frozen subject/task after the reconciliation change produced:
Mean: 92 reads / 93.3 model calls / 24.0s. For context, the earlier StopFile clean run Latest large-file behavior:
The previous exhaustive tail is therefore no longer the dominant termination mode for the large files. Current range runtime vs Claude CodeSame frozen revision, same task:
System One-only files include the agent Blind quality reviewThe evaluator received one anonymized candidate per fresh session and the exact frozen repository revision.
System One dimensions:
The evaluator's main System One gaps were exactly the architectural direction suggested by the cross-trace analysis:
The primary quality deficit is no longer inability to stop. It is now cross-file dependency discovery + evidence shaping:
Confidence-session infrastructure issue found and fixedTwo attempted runs exposed a real Claude confidence formatting behavior: all file assessments were present, but The finalizer now performs one narrow lossless normalization only when top-level This run validates the entire canonical-result -> comparison -> blind-quality pipeline end to end. |
|
The System One code-localization research has now been extracted into the standalone repository The standalone repository contains the validated range runtime, canonical localization result, Claude Code-style cross-trace, blind quality evaluator, historical pilot records, and the next research roadmap. Future algorithm work should continue there. The first follow-up is issue #1: observation-driven |
|
Closing this research branch as superseded by the repository split. This PR remains the primary historical record for the post-#16 Code Locator evolution (two-phase reader, file-local runtimes, stop control, bounded DecisionView, canonical localization result, and alternative exploration algorithms), but it should not be merged into Narness. Active implementation and ongoing experiments now live in |
Summary
Rework the System One Code Locator around a two-phase architecture instead of a fixed directory/file/line-or-symbol pipeline.
Phase 1 — File Locator
Phase 2 — Progressive Reader
Why
The original pilot made 598 model calls because the harness batched large candidate arrays and scored line frontiers per file. Later line/region/symbol variants still pre-expanded too much file state before the model had chosen to inspect it.
The new invariant is:
Source content is now an observation produced by a selected action, not a precomputed candidate set.
Real TypeSafe baseline
Run 35828994234 against BestNathan/nession staging completed successfully with:
Examples from the trace show the desired behavior: high-scoring websocket ranges cause contiguous local expansion, while low-scoring small wrapper/index files stop after the initial probe.
Research follow-ups
Implementation, tests, workflow controls, design notes, and the two-phase Nession baseline are included in this PR.
Multi-hotspot follow-up
The first two-phase trace exposed a harness defect: only the single strongest relevance hotspot could expand, so a second relevant region inside the same file could be discovered but not explored locally.
This PR now preserves up to three disconnected
RelevantRegionhotspots per file and generates bounded before/after actions for each.Real TypeSafe run 35831659317 validates the fix on
crates/nession-agent/src/server/websocket.rs:1585-1724, relevance0.70;1445-1584and1725-1864around that hotspot;1445-1584with probability0.62;0.78relevance.The previous single-hotspot baseline instead had no neighbor action for the second hotspot and jumped to an unrelated gap that scored
0.22.Research notes:
research/code-locator/docs/pilots/nession-websocket-two-phase-trace-analysis-2026-09-23.mdresearch/code-locator/docs/pilots/nession-websocket-multi-hotspot-reader-2026-09-23.mdEvidence-driven reader budget
The fixed four-round reader could stop immediately after discovering a useful new hotspot on round four. The reader now uses a soft/hard budget:
Real run 35833091673 first validated the soft/hard mechanism at batch scope:
server/handler.rsfound a 0.70 hotspot on round four and received a round-five local expansion instead of terminating.That trace also exposed sibling over-expansion. The continuation rule was therefore refined to file scope. Real run 35833457198 validates the refinement:
cli/client/connection.rsscored 0.64 on round four and stopped immediately;server_client.rsindependently continued with 0.82 / 0.81 / 0.70 through rounds four to six, then stopped after a 0.25 round-seven observation;server/handler.rsstill earned round five from its round-four hotspot.Research notes:
research/code-locator/docs/pilots/nession-websocket-soft-hard-round-budget-2026-09-23.mdresearch/code-locator/docs/pilots/nession-websocket-file-scoped-round-budget-2026-09-23.mdClaude Code accuracy reference
A new manual dual-run workflow compares the System One locator against Claude Code using the existing
dsenvironment on the same exact Nession revision and verbatim task.Reference run 35834529068 (
BestNathan/nession@b76fe4921a63023a69ce91399328ab53d3526664) found:The main primary miss was
web/src/platform/socket/MessageRouter.ts: System One scored it 0.64 at rank 21, below the 0.65 threshold and outside the top-16 frontier. However, the already-observedWebSocketService.ts:1-140imports./MessageRouterdirectly. This indicates the next architectural accuracy improvement should be dynamic cross-file actions discovered from observations, rather than simply lowering Phase-1 thresholds until transitive dependencies fit.Accuracy study:
research/code-locator/docs/pilots/nession-websocket-claude-reference-accuracy-2026-09-23.md.github/workflows/system-one-code-locator-accuracy.ymlresearch/code-locator/src/compare_claude_reference.pyGlobal file scheduler + Claude execution trace
The Phase-2 reader no longer uses a Phase-1 top-k cap or fixed file batches.
Real run 35839323377 retained 18 Phase-1 files and selected only three in its first global scheduling request. It completed with 12 model calls, 8 reads, 3 unique files, and 3.78s elapsed. This is recorded as behavior, not judged against another model's final file set.
The Claude Code +
dsside of the same run now preserves its full execution trajectory:claude.raw.jsonl— complete stream-json events;execution-path.json— normalized ordered tool calls, inputs and result metadata;execution-summary.md— human-readable path;reference.json— final localization record.Claude used 42 tools (24 Read, 2 Grep, 16 Bash) across 43 turns, moving from broad websocket search into connection/session/router/handler dependencies. Claude is explicitly treated as a separate System-2 trace, not ground truth or an optimization target.
Research note:
research/code-locator/docs/pilots/nession-websocket-global-scheduler-cross-trace-2026-09-23.mdCanonical localization result
System One and Claude Code now emit the same execution-independent
code-localization-resultcontract.Each result contains:
System One confidence comes from Noul relevance and derived evidence aggregation. Claude confidence is model self-assessment. The schema preserves these different semantics instead of assuming calibration.
Execution traces remain separate from final results. The generic comparator now consumes only the two canonical
localization-result.jsonartifacts and reports symmetric file/evidence overlap without designating either side as truth.Real validation run 35844251415:
Contract:
research/code-locator/docs/localization-result.mdresearch/code-locator/src/localization_result.pyresearch/code-locator/src/compare_localization_results.pyTwo-session Claude confidence + cost contract
Claude Code now separates localization from confidence:
The canonical localization result also carries cost metrics: elapsed time, model calls / turns / tool calls where available, input/output/cache/thinking tokens, provider USD cost when reported, and per-stage breakdowns.
Real validation run 35846265667:
Contract:
research/code-locator/docs/localization-result.mdRange-only System One Runtime v0
A separate v0 prototype now reduces Phase 2 to a content-agnostic state/action loop:
State -> geometry-only ActionSpace -> System One Noul scores -> threshold-based parallel selection -> ReadRange effects -> raw observations -> new State.The only runtime effects are
ReadRangeandStopTask. Navigation labels (seed_head/middle/tail,expand_before/after,jump) are generated from line-count / coverage geometry only. No AST, LSP, symbols, imports, keywords, semantic chunks, or Harness content analysis are used.Selection policy:
StopTask >= thresholdand Stop is globally highest;Real run 35855369932 on
BestNathan/nession@b76fe4921a63023a69ce91399328ab53d3526664demonstrated that range-only navigation can jump directly to distant useful areas. Forcrates/nession-agent/src/server/websocket.rs, epoch 1 selected1-140 (.81)and1445-1584 (.67); epoch 2 then selected1305-1444 (.77)and1585-1724 (.79)plus additional jump probes.The run also exposed the next runtime issues:
Pilot:
research/code-locator/docs/pilots/nession-websocket-range-runtime-v0-2026-09-23.mdPer-file range runtime v0
Phase 2 has been redesigned around independent file-local state machines after confirming the System One request context constraint makes one shared multi-file raw-observation state a poor fit.
Current v0:
FileRuntime;head/middle/tail,expand_before/after,jump,StopFile);.65action threshold controls concurrency only;StopFilecan producemodel_stop; low scores never stop the Harness;Real validation run
35857120750on the Nession websocket task:action_space_exhausted;handler.rsreached 94.2% coverage before the 32-epoch safety budget;model_stopdecisions.The run validates the per-file context architecture and also shows the next research target clearly: current
StopFilesemantics are conservative, so System One tends to keep exploring until coverage is exhausted rather than declaring sufficiency early.Research trace:
research/code-locator/docs/pilots/nession-websocket-per-file-range-runtime-v0-2026-09-23.mdStopFile control-choice experiment
Prompt-only StopFile experiments were not enough. Treating StopFile and ReadRange as comparable Noul scores biased the runtime toward exhaustive scanning because they represent different semantics.
The current FileRuntime now sends one mixed System One request per epoch:
Choice: StopFile vs ContinueFile — control-flow sufficiency;Noul: one utility score per ReadRange — information-gathering value.Policy:
action_space_exhaustedwithout asking the model.StopFile is explicitly defined as representative-result sufficiency, not full-file coverage. The goal is to return enough representative valuable ranges to characterize the file, not exhaustively enumerate every relevant range.
Clean run
35864316780:model_stopdecisions;action_space_exhausted;budget_exhausted;Compared descriptively with the earlier per-file baseline, those four common files used 24 reads before and 17 reads now. Evidence-line overlap ranged from 0.775 to 1.000; the runs are stochastic, so this is not treated as a controlled accuracy metric.
The remaining open stop problem is concentrated in very large files (
agent/server/websocket.rs,server_client.rs,handler.rs), where System One still prefers Continue until exhaustion or the safety budget.Two experimental System One localization algorithms
Two algorithms discussed separately were implemented as isolated experiments without replacing the current per-file range-runtime baseline.
A — Explore-Guided Evidence Filtering
Real run
35882424922onBestNathan/nession@97b9d2b49c5137064e910fc180aabc01cfb1021f: 17 Phase-1 files, 135 reads, 8 valuable files, 32 evidence regions, 164 model calls, 1.83M input tokens, 40.84s.B — Adaptive Semantic Zoom Search
Real run
35881894374on the same subject revision: 17 Phase-1 files, 417 probes, 12 probability-frontier convergences, 5 round-budget exits, 11 valuable files, 194 evidence regions, 184 model calls, 1.02M input tokens, 43.36s.Manual A/B workflow:
.github/workflows/system-one-localization-algorithms.yml.Design note:
research/code-locator/docs/system-one-localization-algorithms.md.