Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
277 commits
Select commit Hold shift + click to select a range
cbd7aee
refactor: keep region frontier compact and semantic
BestNathan Sep 23, 2026
2b8bd90
test: cover compact region frontier
BestNathan Sep 23, 2026
ee828c2
research: record code locator run 35824075828
github-actions[bot] Sep 23, 2026
c366854
research: record code locator run 35824091019
github-actions[bot] Sep 23, 2026
77925b7
ci: run compact-region baseline once
BestNathan Sep 23, 2026
ae01461
research: record code locator run 35824169254
github-actions[bot] Sep 23, 2026
1b65a2e
ci: remove compact-region one-shot trigger
BestNathan Sep 23, 2026
8d6380e
fix: bound region frontier for TypeSafe budget
BestNathan Sep 23, 2026
464d583
research: record code locator run 35824262492
github-actions[bot] Sep 23, 2026
655a75a
research: record code locator run 35824273772
github-actions[bot] Sep 23, 2026
69d2c37
ci: run bounded-region baseline once
BestNathan Sep 23, 2026
df5099d
research: record code locator run 35824339286
github-actions[bot] Sep 23, 2026
7f7de97
ci: restore manual-only TypeSafe baseline runs
BestNathan Sep 23, 2026
c9ffd27
ci: persist code locator records only for main or manual runs
BestNathan Sep 23, 2026
ff53b41
research: record stable code locator baseline
BestNathan Sep 23, 2026
39bacb9
docs: align code locator README with stable region baseline
BestNathan Sep 23, 2026
86c5e0c
docs: establish stable code locator research baseline
BestNathan Sep 23, 2026
635aae4
chore: remove generated branch run snapshots
BestNathan Sep 23, 2026
a86779c
refactor: use semantic symbol frontier for source localization
BestNathan Sep 23, 2026
687bf91
fix: sanitize symbol outlines before model disclosure
BestNathan Sep 23, 2026
f9640a0
test: lock symbol frontier semantics
BestNathan Sep 23, 2026
86ad372
ci: rename source threshold to symbol threshold
BestNathan Sep 23, 2026
017950f
refactor: name third-stage threshold after symbols
BestNathan Sep 23, 2026
b1ed88b
fix: correct symbol declaration regexes
BestNathan Sep 23, 2026
f0a30c6
ci: run symbol-frontier baseline once
BestNathan Sep 23, 2026
37833b0
ci: remove symbol-frontier one-shot trigger
BestNathan Sep 23, 2026
7056d66
refactor: add file outline decision stage
BestNathan Sep 23, 2026
c5cfa31
test: enforce outline-gated symbol disclosure
BestNathan Sep 23, 2026
b46aeca
ci: wire outline stage into code locator workflow
BestNathan Sep 23, 2026
9f6d605
ci: run outline-symbol baseline once
BestNathan Sep 23, 2026
7028998
ci: run four-stage code locator baseline once
BestNathan Sep 23, 2026
65e0ef0
ci: remove outline-symbol one-shot trigger
BestNathan Sep 23, 2026
19bd66a
refactor: make outlines semantic symbol scopes
BestNathan Sep 23, 2026
523685b
test: validate semantic scope outlines
BestNathan Sep 23, 2026
2746da4
ci: run semantic-scope baseline once
BestNathan Sep 23, 2026
d25d4f1
ci: remove semantic-scope one-shot trigger
BestNathan Sep 23, 2026
a91a269
fix: group file-level symbols into module scopes
BestNathan Sep 23, 2026
95a8125
test: keep file-level symbols in one module scope
BestNathan Sep 23, 2026
be7b07e
ci: run module-scope baseline once
BestNathan Sep 23, 2026
df2e570
ci: remove module-scope one-shot trigger
BestNathan Sep 23, 2026
a8021b5
perf: share candidate state across Noul questions
BestNathan Sep 23, 2026
15563cd
test: enforce shared-state Noul request shape
BestNathan Sep 23, 2026
245ad2f
ci: run shared-state baseline once
BestNathan Sep 23, 2026
807830e
ci: remove shared-state one-shot trigger
BestNathan Sep 23, 2026
f424657
perf: isolate candidate context while sharing policy
BestNathan Sep 23, 2026
0363c13
test: enforce isolated candidate request context
BestNathan Sep 23, 2026
7e9b3ff
ci: run isolated-context baseline once
BestNathan Sep 23, 2026
0dd361c
ci: remove isolated-context one-shot trigger
BestNathan Sep 23, 2026
bf2fcdd
fix: restore isolated Noul decision semantics
BestNathan Sep 23, 2026
5b49443
refactor: add file outline before semantic scopes
BestNathan Sep 23, 2026
f5159d1
test: enforce five-stage semantic disclosure
BestNathan Sep 23, 2026
82e66d9
ci: expose file-outline and scope thresholds
BestNathan Sep 23, 2026
faea4d8
ci: run five-stage baseline once
BestNathan Sep 23, 2026
a6b4332
refactor: turn code locator into two-phase progressive reader
BestNathan Sep 23, 2026
f480bb5
test: cover two-phase progressive reader semantics
BestNathan Sep 23, 2026
0fb6d1d
ci: configure two-phase progressive reader experiment
BestNathan Sep 23, 2026
db6a99d
docs: describe two-phase progressive code reader
BestNathan Sep 23, 2026
c2e1360
docs: define progressive reader research model
BestNathan Sep 23, 2026
657f479
ci: run two-phase progressive reader baseline once
BestNathan Sep 23, 2026
6d52953
ci: remove two-phase one-shot trigger
BestNathan Sep 23, 2026
f6b710b
refactor: let observation scores shape read actions
BestNathan Sep 23, 2026
3c676dd
ci: run scored progressive reader baseline once
BestNathan Sep 23, 2026
615d476
ci: remove scored reader one-shot trigger
BestNathan Sep 23, 2026
431b318
research: record two-phase reader baseline
BestNathan Sep 23, 2026
4a1bfb7
research: record two-phase reader trace findings
BestNathan Sep 23, 2026
8faa36b
refactor: preserve multiple reader relevance hotspots
BestNathan Sep 23, 2026
c7016c3
test: cover multi-hotspot reader action frontiers
BestNathan Sep 23, 2026
fdc98e7
docs: capture multi-hotspot reader design
BestNathan Sep 23, 2026
cdbb04a
docs: link baseline to trace analysis
BestNathan Sep 23, 2026
5a4835e
ci: run multi-hotspot reader baseline once
BestNathan Sep 23, 2026
c6f72cd
ci: remove multi-hotspot one-shot trigger
BestNathan Sep 23, 2026
b36cd40
research: record multi-hotspot reader baseline
BestNathan Sep 23, 2026
a20476a
research: close multi-hotspot trace finding
BestNathan Sep 23, 2026
8cb287f
refactor: make reader round budget evidence-driven
BestNathan Sep 23, 2026
daa00f2
fix: preserve soft-round default with legacy flag
BestNathan Sep 23, 2026
d2fa048
test: cover soft and hard reader budgets
BestNathan Sep 23, 2026
bd7d2ed
ci: expose evidence-driven reader budgets
BestNathan Sep 23, 2026
b847a97
docs: explain soft and hard reader budgets
BestNathan Sep 23, 2026
3037c20
docs: define evidence-driven round continuation
BestNathan Sep 23, 2026
e434eab
ci: run evidence-driven round budget once
BestNathan Sep 23, 2026
8112b57
research: record evidence-driven round budget run
BestNathan Sep 23, 2026
e0484d7
refactor: make soft round extensions file-scoped
BestNathan Sep 23, 2026
2b6de2b
test: ensure soft budget continuation is file-scoped
BestNathan Sep 23, 2026
bca1f03
ci: remove round-budget one-shot trigger
BestNathan Sep 23, 2026
21eb9f7
test: account for eventual soft-budget stop
BestNathan Sep 23, 2026
5128124
research: record file-scoped reader budget baseline
BestNathan Sep 23, 2026
7b89135
docs: clarify file-scoped reader continuation
BestNathan Sep 23, 2026
0bb271b
docs: make continuation ownership file-scoped
BestNathan Sep 23, 2026
b70dcd4
research: add Claude reference comparison
BestNathan Sep 23, 2026
7a280b8
test: cover Claude reference comparison metrics
BestNathan Sep 23, 2026
df993f6
research: add System One vs Claude Code accuracy workflow
BestNathan Sep 23, 2026
03ed4f2
ci: keep accuracy reference workflow manual
BestNathan Sep 23, 2026
60573a2
research: measure Claude evidence-region coverage
BestNathan Sep 23, 2026
75e6d80
test: cover evidence-region accuracy metrics
BestNathan Sep 23, 2026
b0381d3
research: record Claude reference localization accuracy
BestNathan Sep 23, 2026
94003e5
docs: add Claude localization accuracy reference
BestNathan Sep 23, 2026
d984b97
docs: record static-frontier accuracy finding
BestNathan Sep 23, 2026
4d497af
refactor: let System One globally schedule reader files
BestNathan Sep 23, 2026
6399f4a
fix: report global reader activation threshold
BestNathan Sep 23, 2026
804ddad
test: cover global file scheduling without fixed batches
BestNathan Sep 23, 2026
c179c33
research: capture Claude Code localization execution path
BestNathan Sep 23, 2026
1936a1a
ci: remove fixed Phase-2 file batching
BestNathan Sep 23, 2026
38112ee
research: preserve Claude Code tool trajectory
BestNathan Sep 23, 2026
79b4884
test: distinguish scheduler epochs from file reads
BestNathan Sep 23, 2026
f558efd
test: cover Claude Code execution-path extraction
BestNathan Sep 23, 2026
fc84f5e
research: treat Claude comparison as observation only
BestNathan Sep 23, 2026
b042bac
ci: run global-scheduler cross-trace once
BestNathan Sep 23, 2026
2997475
ci: keep cross-trace workflow manual
BestNathan Sep 23, 2026
7cacc97
research: record global scheduler and Claude cross-trace
BestNathan Sep 23, 2026
f6b53af
docs: describe global System One file scheduler
BestNathan Sep 23, 2026
6a32d07
docs: replace batch reader with global file scheduler
BestNathan Sep 23, 2026
745490c
research: clarify Claude trace is not ground truth
BestNathan Sep 23, 2026
2137664
research: define canonical localization result schema
BestNathan Sep 23, 2026
01f93a3
research: emit structured System One localization result
BestNathan Sep 23, 2026
4dd647d
research: normalize Claude output to shared localization schema
BestNathan Sep 23, 2026
47b43f3
fix: normalize Claude overall confidence
BestNathan Sep 23, 2026
d4f244c
research: add symmetric localization result comparator
BestNathan Sep 23, 2026
6a42827
research: compare canonical localization results
BestNathan Sep 23, 2026
ee289bb
test: load shared localization result module
BestNathan Sep 23, 2026
7138b99
test: load shared localization result module
BestNathan Sep 23, 2026
0b4c5f8
test: validate canonical Claude localization result
BestNathan Sep 23, 2026
4d3b564
test: cover canonical localization result structure
BestNathan Sep 23, 2026
afa5dcd
test: cover symmetric localization result comparison
BestNathan Sep 23, 2026
29f51fc
ci: validate canonical localization artifacts once
BestNathan Sep 23, 2026
3f5f46a
docs: define canonical localization result contract
BestNathan Sep 23, 2026
bd92909
docs: link canonical localization output
BestNathan Sep 23, 2026
1f6832b
docs: link canonical localization output
BestNathan Sep 23, 2026
0fe9da5
ci: finalize canonical cross-trace workflow
BestNathan Sep 23, 2026
ec0a967
docs: record canonical result validation run
BestNathan Sep 23, 2026
d33d4f2
research: separate Claude localization and confidence stages
BestNathan Sep 23, 2026
6ceb592
research: make first Claude session localization-only
BestNathan Sep 23, 2026
9e782a8
research: finalize Claude confidence in a fresh session
BestNathan Sep 23, 2026
5d143b4
research: split Claude localization and confidence sessions
BestNathan Sep 23, 2026
4a36b9f
test: cover two-session confidence and cost contract
BestNathan Sep 23, 2026
8dea657
test: make Claude trace localization-only
BestNathan Sep 23, 2026
370973a
research: include cost records in localization comparison
BestNathan Sep 23, 2026
152ed48
test: compare localization cost records
BestNathan Sep 23, 2026
147886e
docs: define two-session confidence and cost schema
BestNathan Sep 23, 2026
c03c415
docs: explain two-session Claude result pipeline
BestNathan Sep 23, 2026
d694a90
docs: separate Claude localization from confidence
BestNathan Sep 23, 2026
b7a5a54
ci: validate two-session Claude localization once
BestNathan Sep 23, 2026
0a97d5f
ci: finalize two-session cross-trace workflow
BestNathan Sep 23, 2026
a9f3537
research: record two-session cost validation
BestNathan Sep 23, 2026
cfec3e6
research: add range-only System One runtime v0
BestNathan Sep 23, 2026
666d18d
test: cover range runtime action generation and stop policy
BestNathan Sep 23, 2026
658e3c4
research: add one-shot range runtime v0 experiment
BestNathan Sep 23, 2026
0346d67
ci: keep range runtime experiment manual
BestNathan Sep 23, 2026
310c980
fix: transport-batch range action scoring
BestNathan Sep 23, 2026
ad2f3c2
ci: rerun range runtime v0 after transport batching
BestNathan Sep 23, 2026
de222aa
ci: restore manual range runtime trigger
BestNathan Sep 23, 2026
5a2b58a
fix: make range action scoring token-adaptive
BestNathan Sep 23, 2026
ec12563
ci: rerun token-adaptive range runtime v0
BestNathan Sep 23, 2026
b519510
ci: keep range runtime v0 manual
BestNathan Sep 23, 2026
86a5655
refactor: shrink System One range decision view
BestNathan Sep 23, 2026
436c5ec
ci: rerun range runtime with compact decision view
BestNathan Sep 23, 2026
244d931
ci: restore manual range runtime trigger
BestNathan Sep 23, 2026
a9b3262
refactor: bound range runtime decision context mechanically
BestNathan Sep 23, 2026
995b891
ci: rerun bounded-context range runtime v0
BestNathan Sep 23, 2026
22fe171
ci: restore manual range runtime trigger
BestNathan Sep 23, 2026
d04cc80
research: record range-only runtime v0 experiment
BestNathan Sep 23, 2026
169176d
refactor: make range runtime file-local
BestNathan Sep 23, 2026
39d533b
test: align range runtime with per-file state machines
BestNathan Sep 23, 2026
5a80162
ci: run per-file range runtime prototype
BestNathan Sep 23, 2026
eae87ae
test: fix per-file range runtime expectations
BestNathan Sep 23, 2026
63c81f0
ci: run per-file range runtime v0 once
BestNathan Sep 23, 2026
2c418be
ci: restore manual per-file range runtime trigger
BestNathan Sep 23, 2026
7db6aab
fix: batch evidence scoring and avoid overlapping reads
BestNathan Sep 23, 2026
f6f19e6
test: reject overlapping parallel range effects
BestNathan Sep 23, 2026
3aa28d0
test: give parallel selector real ranges
BestNathan Sep 23, 2026
20e904f
ci: validate per-file range runtime v0
BestNathan Sep 23, 2026
c95bd10
ci: restore manual per-file range runtime trigger
BestNathan Sep 23, 2026
979f741
research: record per-file range runtime v0
BestNathan Sep 23, 2026
4ebccf1
docs: make per-file range runtime the Phase-2 baseline
BestNathan Sep 23, 2026
2e1e0e8
docs: redesign Phase 2 as independent file runtimes
BestNathan Sep 23, 2026
916f145
research: define StopFile as localization-result sufficiency
BestNathan Sep 23, 2026
fd02edc
ci: run StopFile sufficiency experiment once
BestNathan Sep 23, 2026
d392f3e
ci: restore manual StopFile experiment trigger
BestNathan Sep 23, 2026
ce7032a
research: make StopFile a marginal-value decision
BestNathan Sep 23, 2026
a98e04b
ci: run StopFile marginal-value experiment once
BestNathan Sep 23, 2026
8ce9c32
ci: restore manual StopFile experiment trigger
BestNathan Sep 23, 2026
a7b8f7f
research: score StopFile on the same action-utility scale
BestNathan Sep 23, 2026
58a1749
ci: run unified StopFile utility experiment once
BestNathan Sep 23, 2026
ba18fd1
ci: restore manual unified-stop experiment trigger
BestNathan Sep 23, 2026
134aa35
refactor: separate StopFile control from read utility scoring
BestNathan Sep 23, 2026
04bf43a
test: cover StopFile control choice semantics
BestNathan Sep 23, 2026
99ffc09
ci: run StopFile control-choice experiment once
BestNathan Sep 23, 2026
2aef300
ci: restore manual StopFile control-choice trigger
BestNathan Sep 23, 2026
61f26ad
research: let StopFile Choice terminate without threshold gate
BestNathan Sep 23, 2026
bb71203
test: make StopFile Choice itself terminal
BestNathan Sep 23, 2026
cadb6e5
ci: run representative StopFile choice experiment once
BestNathan Sep 23, 2026
dbc9709
ci: restore manual representative-stop trigger
BestNathan Sep 23, 2026
a0283f6
fix: exhaust file before asking StopFile control
BestNathan Sep 23, 2026
b7c89eb
ci: validate clean representative StopFile semantics once
BestNathan Sep 23, 2026
434b3ae
ci: restore manual clean-stop trigger
BestNathan Sep 23, 2026
8feb29d
research: record StopFile control-choice experiment
BestNathan Sep 23, 2026
63c1631
docs: make StopFile Choice the current baseline
BestNathan Sep 23, 2026
92c9dad
docs: separate stop control from read utility
BestNathan Sep 23, 2026
b81b9fe
research: add blind localization quality evaluator
BestNathan Sep 23, 2026
3b7fa72
test: cover localization quality evaluator contract
BestNathan Sep 23, 2026
339c588
research: reveal candidate mapping only in final report
BestNathan Sep 23, 2026
6162ddb
research: add blind Claude quality scoring workflow
BestNathan Sep 23, 2026
649782d
ci: validate blind localization quality evaluation once
BestNathan Sep 23, 2026
949cc8a
ci: restore manual quality evaluation workflow trigger
BestNathan Sep 23, 2026
dca4125
fix: robustly extract evaluator JSON from final text
BestNathan Sep 23, 2026
f8b546f
ci: preserve raw quality evaluation on failure
BestNathan Sep 23, 2026
3f5e802
research: add replayable localization quality workflow
BestNathan Sep 23, 2026
84cd7d3
ci: keep quality evaluation replay manual
BestNathan Sep 23, 2026
5d72034
research: implement explore-guided evidence filtering
BestNathan Sep 23, 2026
35ad5fd
research: implement adaptive semantic zoom search
BestNathan Sep 23, 2026
00d43f7
test: cover evidence filtering and adaptive zoom algorithms
BestNathan Sep 23, 2026
0b6052e
fix: make evidence redundancy explicit
BestNathan Sep 23, 2026
d8f6172
test: provide zoom fixture extension metadata
BestNathan Sep 23, 2026
627c1ad
ci: add manual A/B localization algorithm experiments
BestNathan Sep 23, 2026
cb8b4f2
docs: describe two System One localization algorithms
BestNathan Sep 23, 2026
b8393fc
fix: align evidence prompt wording with tests
BestNathan Sep 23, 2026
1dae505
docs: link experimental localization algorithms
BestNathan Sep 23, 2026
fbea38c
ci: validate localization algorithms once
BestNathan Sep 23, 2026
b5f7a58
ci: restore manual localization algorithm trigger
BestNathan Sep 23, 2026
446f818
fix: default A/B experiment subject on push
BestNathan Sep 23, 2026
68a6b55
ci: rerun A/B algorithms on nession once
BestNathan Sep 23, 2026
4922976
ci: restore manual A/B experiment trigger
BestNathan Sep 23, 2026
975fcd6
fix: bound evidence-guided redundancy context
BestNathan Sep 23, 2026
980c1a1
ci: validate evidence-guided algorithm once
BestNathan Sep 23, 2026
7720a6a
ci: restore manual A/B localization workflow
BestNathan Sep 23, 2026
5708c9b
research: record real A/B localization runs
BestNathan Sep 23, 2026
9136815
fix: reconcile StopFile with concrete read utility
BestNathan Sep 23, 2026
64515ce
test: cover StopFile conflict reconciliation
BestNathan Sep 23, 2026
71586e9
feat: canonicalize per-file range runtime results
BestNathan Sep 23, 2026
2847a55
feat: emit canonical result from range runtime
BestNathan Sep 23, 2026
190dd84
ci: compare Claude against current range runtime
BestNathan Sep 23, 2026
0b5aa18
feat: bind canonical localization to subject revision
BestNathan Sep 23, 2026
53d9972
feat: record frozen subject in range runtime output
BestNathan Sep 23, 2026
58bf156
feat: bind Claude canonical result to frozen subject
BestNathan Sep 23, 2026
d908152
ci: pin canonical comparison to one subject revision
BestNathan Sep 23, 2026
2c7569f
fix: reject cross-revision localization comparisons
BestNathan Sep 23, 2026
61def42
test: enforce canonical subject comparison gate
BestNathan Sep 23, 2026
46c4919
test: cover canonical range-runtime result
BestNathan Sep 23, 2026
f3eabd8
test: regress evaluator non-bare JSON output
BestNathan Sep 23, 2026
8f95e2b
docs: describe reconciled range-runtime semantics
BestNathan Sep 23, 2026
e930234
docs: define canonical subject identity gate
BestNathan Sep 23, 2026
c1d4a69
ci: trigger one cross-trace run from PR branch
BestNathan Sep 23, 2026
c10f6d1
ci: restore manual cross-trace trigger
BestNathan Sep 23, 2026
b79d73c
fix: expose malformed confidence assessment diagnostics
BestNathan Sep 23, 2026
6d7ee9b
ci: preserve malformed Claude confidence artifacts
BestNathan Sep 23, 2026
c442e84
ci: trigger corrected cross-trace run
BestNathan Sep 23, 2026
0888d21
ci: restore manual cross-trace after corrected run
BestNathan Sep 23, 2026
0ffdad6
fix: normalize nested Claude confidence files
BestNathan Sep 23, 2026
bf6f95d
test: cover nested confidence schema repair
BestNathan Sep 23, 2026
4753552
ci: trigger cross-trace after confidence schema repair
BestNathan Sep 23, 2026
50809f9
ci: restore manual cross-trace after schema-repair run
BestNathan Sep 23, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
223 changes: 223 additions & 0 deletions .github/workflows/localization-quality-evaluation.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,223 @@
name: Localization Quality Evaluation

on:
workflow_dispatch:
inputs:
source_run_id:
description: Cross-trace run ID whose localization artifacts should be evaluated
required: true
default: "35869966400"
type: string

permissions:
contents: read
actions: read

env:
SOURCE_RUN_ID: ${{ inputs.source_run_id || '35869966400' }}

jobs:
evaluate:
name: Blind Claude localization quality review
runs-on: ubuntu-24.04
environment: ds
timeout-minutes: 120
env:
CLAUDE_MODEL: ${{ vars.CLAUDE_MODEL }}
ANTHROPIC_BASE_URL: ${{ vars.ANTHROPIC_BASE_URL || 'https://api.anthropic.com' }}
steps:
- name: Checkout evaluator code
uses: actions/checkout@v4
with:
path: narness
fetch-depth: 1
persist-credentials: false

- uses: actions/setup-node@v4
with:
node-version: "22.14.0"

- uses: actions/setup-python@v5
with:
python-version: "3.13"

- name: Install pinned Claude Code
run: |
set -euo pipefail
npm install --global --no-fund --no-audit @anthropic-ai/claude-code@2.1.278
claude --version

- name: Download frozen localization artifacts
env:
GH_TOKEN: ${{ github.token }}
run: |
set -euo pipefail
mkdir -p artifacts/system-one artifacts/claude-reference
gh run download "$SOURCE_RUN_ID" \
--repo "$GITHUB_REPOSITORY" \
--name "code-locator-cross-trace-system-one-$SOURCE_RUN_ID" \
--dir artifacts/system-one
gh run download "$SOURCE_RUN_ID" \
--repo "$GITHUB_REPOSITORY" \
--name "code-locator-cross-trace-claude-$SOURCE_RUN_ID" \
--dir artifacts/claude-reference

- name: Resolve frozen subject revision
id: source
run: |
set -euo pipefail
python3 - <<'PY' >> "$GITHUB_OUTPUT"
import json
from pathlib import Path
manifest=json.loads(Path("artifacts/system-one/manifest.json").read_text())
print(f"repo={manifest['subject_repository']}")
print(f"sha={manifest['subject_sha']}")
print(f"query={manifest['query']}")
PY

- name: Checkout exact subject revision
uses: actions/checkout@v4
with:
repository: ${{ steps.source.outputs.repo }}
ref: ${{ steps.source.outputs.sha }}
path: subject
fetch-depth: 1
persist-credentials: false

- name: Require evaluator configuration
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
run: |
set -euo pipefail
test -n "$ANTHROPIC_API_KEY"
test -n "$ANTHROPIC_BASE_URL"
test -n "$CLAUDE_MODEL"

- name: Prepare anonymous candidates
run: |
set -euo pipefail
mkdir -p artifacts/quality-evaluation/candidate-a
mkdir -p artifacts/quality-evaluation/candidate-b

python3 narness/research/code-locator/src/localization_quality_evaluation.py prepare \
--input artifacts/system-one/localization-result.json \
--candidate-id candidate-a \
--output-candidate artifacts/quality-evaluation/candidate-a/candidate.json \
--output-prompt artifacts/quality-evaluation/candidate-a/prompt.txt

python3 narness/research/code-locator/src/localization_quality_evaluation.py prepare \
--input artifacts/claude-reference/localization-result.json \
--candidate-id candidate-b \
--output-candidate artifacts/quality-evaluation/candidate-b/candidate.json \
--output-prompt artifacts/quality-evaluation/candidate-b/prompt.txt

- name: Evaluate anonymous candidate A
working-directory: subject
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
run: |
set -euo pipefail
claude -p \
--output-format stream-json \
--verbose \
--model "$CLAUDE_MODEL" \
--effort high \
--max-turns 80 \
--permission-mode dontAsk \
--no-session-persistence \
--setting-sources '' \
--strict-mcp-config \
--mcp-config '{"mcpServers":{}}' \
--disable-slash-commands \
--tools 'Bash,Read,Glob,Grep' \
--allowedTools 'Bash(*)' 'Read' 'Glob' 'Grep' \
--disallowedTools 'Edit' 'Write' 'Agent' 'Task' 'WebFetch' 'WebSearch' 'EndConversation' 'mcp__*' \
--settings '{"disableAllHooks":true}' \
< ../artifacts/quality-evaluation/candidate-a/prompt.txt \
> ../artifacts/quality-evaluation/candidate-a/evaluator.raw.jsonl

- name: Evaluate anonymous candidate B
working-directory: subject
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
run: |
set -euo pipefail
claude -p \
--output-format stream-json \
--verbose \
--model "$CLAUDE_MODEL" \
--effort high \
--max-turns 80 \
--permission-mode dontAsk \
--no-session-persistence \
--setting-sources '' \
--strict-mcp-config \
--mcp-config '{"mcpServers":{}}' \
--disable-slash-commands \
--tools 'Bash,Read,Glob,Grep' \
--allowedTools 'Bash(*)' 'Read' 'Glob' 'Grep' \
--disallowedTools 'Edit' 'Write' 'Agent' 'Task' 'WebFetch' 'WebSearch' 'EndConversation' 'mcp__*' \
--settings '{"disableAllHooks":true}' \
< ../artifacts/quality-evaluation/candidate-b/prompt.txt \
> ../artifacts/quality-evaluation/candidate-b/evaluator.raw.jsonl

- name: Build scorecards and report
run: |
set -euo pipefail
python3 narness/research/code-locator/src/localization_quality_evaluation.py finalize \
--candidate artifacts/quality-evaluation/candidate-a/candidate.json \
--raw-jsonl artifacts/quality-evaluation/candidate-a/evaluator.raw.jsonl \
--subject-root subject \
--model "$CLAUDE_MODEL" \
--output-json artifacts/quality-evaluation/candidate-a/scorecard.json \
--output-markdown artifacts/quality-evaluation/candidate-a/report.md

python3 narness/research/code-locator/src/localization_quality_evaluation.py finalize \
--candidate artifacts/quality-evaluation/candidate-b/candidate.json \
--raw-jsonl artifacts/quality-evaluation/candidate-b/evaluator.raw.jsonl \
--subject-root subject \
--model "$CLAUDE_MODEL" \
--output-json artifacts/quality-evaluation/candidate-b/scorecard.json \
--output-markdown artifacts/quality-evaluation/candidate-b/report.md

python3 narness/research/code-locator/src/localization_quality_evaluation.py report \
--candidate-a artifacts/quality-evaluation/candidate-a/scorecard.json \
--candidate-b artifacts/quality-evaluation/candidate-b/scorecard.json \
--candidate-a-label "System One" \
--candidate-b-label "Claude Code" \
--output-json artifacts/quality-evaluation/evaluation-report.json \
--output-markdown artifacts/quality-evaluation/evaluation-report.md

python3 - <<'PY'
import json
import os
from pathlib import Path
root=Path("artifacts/quality-evaluation")
payload={
"schema_version": 1,
"kind": "code-localization-quality-evaluation-run",
"source_run_id": os.environ["SOURCE_RUN_ID"],
"subject_repository": "${{ steps.source.outputs.repo }}",
"subject_sha": "${{ steps.source.outputs.sha }}",
"query": "${{ steps.source.outputs.query }}",
"evaluator_model": os.environ["CLAUDE_MODEL"],
"candidate_mapping": {
"candidate-a": "System One",
"candidate-b": "Claude Code",
},
}
(root/"manifest.json").write_text(
json.dumps(payload, indent=2, ensure_ascii=False)+"\n"
)
PY

cat artifacts/quality-evaluation/evaluation-report.md >> "$GITHUB_STEP_SUMMARY"

- name: Upload quality evaluation
if: always()
uses: actions/upload-artifact@v4
with:
name: code-locator-quality-evaluation-${{ github.run_id }}
path: artifacts/quality-evaluation/
if-no-files-found: warn
retention-days: 90
Loading
Loading