Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
46 commits
Select commit Hold shift + click to select a range
722f019
research: add code locator fixture
BestNathan Sep 23, 2026
1fe0b92
research: add code locator fixture
BestNathan Sep 23, 2026
44db245
research: add code locator fixture
BestNathan Sep 23, 2026
fea07df
research: add System One code locator demo
BestNathan Sep 23, 2026
0f35379
test: cover System One code locator
BestNathan Sep 23, 2026
700543f
docs: document System One code locator
BestNathan Sep 23, 2026
5d919bc
research: record full code locator candidate evidence
BestNathan Sep 23, 2026
37f3ffd
docs: add System One code localization research note
BestNathan Sep 23, 2026
3394a44
ci: collect System One research evidence
BestNathan Sep 23, 2026
3f7b269
docs: extend System One topic with code localization
BestNathan Sep 23, 2026
a1d620a
docs: index System One code locator
BestNathan Sep 23, 2026
b9f3ce5
ci: harden System One evidence collection
BestNathan Sep 23, 2026
af15b39
ci: initialize evidence roots before fallible steps
BestNathan Sep 23, 2026
173ad90
ci: use typesave environment for System One [system-one-e2e]
BestNathan Sep 23, 2026
552013a
ci: read TypeSafe key from typesave env [system-one-e2e]
BestNathan Sep 23, 2026
f276455
ci: bind System One research to typesafe env [system-one-e2e]
BestNathan Sep 23, 2026
fc44a79
ci: keep System One collection manual on typesave env
BestNathan Sep 23, 2026
2c51ccf
ci: use correct typesafe environment
BestNathan Sep 23, 2026
138dac2
ci: trigger System One e2e on typesafe [system-one-e2e]
BestNathan Sep 23, 2026
367121a
ci: finalize typesafe environment wiring
BestNathan Sep 23, 2026
7f8f3e1
ci: rerun System One e2e [system-one-e2e]
BestNathan Sep 23, 2026
0799b2c
ci: default code locator inputs during e2e [system-one-e2e]
BestNathan Sep 23, 2026
13d1221
fix: project code literals before System One transport [system-one-e2e]
BestNathan Sep 23, 2026
852abe8
ci: restore manual System One collection
BestNathan Sep 23, 2026
86e6503
docs: record first real System One localization pilot
BestNathan Sep 23, 2026
a8e3e27
docs: add stage cost breakdown to System One pilot
BestNathan Sep 23, 2026
0ce6090
docs: link real System One localization evidence
BestNathan Sep 23, 2026
aeed884
ci: split code locator into dedicated workflow
BestNathan Sep 23, 2026
9d35fd2
ci: finalize dedicated code locator workflow
BestNathan Sep 23, 2026
f388d6d
ci: remove mixed System One research workflow
BestNathan Sep 23, 2026
50aaa21
docs: separate System One experiment workflows
BestNathan Sep 23, 2026
ed5ae0e
docs: scope evidence contract to code locator workflow
BestNathan Sep 23, 2026
eb9ef97
ci: isolate Kubernetes experiment triggers
BestNathan Sep 23, 2026
9f02184
ci: isolate Kubernetes workflow from code locator changes
BestNathan Sep 23, 2026
e313750
docs: document dedicated code locator workflow
BestNathan Sep 23, 2026
629bba4
ci: keep Kubernetes workflow out of code locator PR
BestNathan Sep 23, 2026
805e169
ci: persist code locator experiment records
BestNathan Sep 23, 2026
fd71e36
research: record code locator run 35819260887
github-actions[bot] Sep 23, 2026
63aee26
fix: render persisted experiment metadata safely
BestNathan Sep 23, 2026
85c049a
research: repair metadata for code locator run 35819260887
BestNathan Sep 23, 2026
5620306
research: record code locator run 35819390521
github-actions[bot] Sep 23, 2026
5112470
docs: define code locator experiment record layout
BestNathan Sep 23, 2026
4b25c92
research: move code locator records to root research directory
BestNathan Sep 23, 2026
84b338e
research: record code locator run 35819796535
github-actions[bot] Sep 23, 2026
fbf5c0f
research: make code locator a self-contained research unit
BestNathan Sep 23, 2026
c07e696
research: record code locator run 35820001391
github-actions[bot] Sep 23, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
435 changes: 435 additions & 0 deletions .github/workflows/system-one-code-locator.yml

Large diffs are not rendered by default.

3 changes: 2 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -224,7 +224,7 @@ bash plugins/narness/scripts/narness-rust-test-integration.sh examples/rust-work

The example contains a root workspace contract, a task Skill with progressive disclosure, unit and integration evidence, an environment declaration, and a stable CI required gate.

The repository also contains research prototypes that are intentionally outside the canonical Narness runtime scope, including the [System One Kubernetes command generator](examples/system-one-k8s/README.md) used by the progressive-action-space topic.
The repository also contains research prototypes that are intentionally outside the canonical Narness runtime scope, including the [System One Kubernetes command generator](examples/system-one-k8s/README.md) and [System One code locator](research/code-locator/README.md) used by the progressive-action-space topic.

## Adopt Narness

Expand Down Expand Up @@ -259,6 +259,7 @@ Start with [docs/adoption.md](docs/adoption.md):
- [Change-to-Evidence Planning topic](docs/topics/change-to-evidence-planning/README.md)
- [System One Progressive Action Spaces topic](docs/topics/system-one-progressive-action-space/README.md)
- [System One Kubernetes command generator](examples/system-one-k8s/README.md)
- [System One code locator](research/code-locator/README.md)
- [AI Workspace architecture](docs/architecture.md)
- [Adoption guide](docs/adoption.md)
- [Runnable Rust example](examples/rust-workspace/README.md)
Expand Down
60 changes: 57 additions & 3 deletions docs/topics/system-one-progressive-action-space/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -587,7 +587,57 @@ The result should distinguish between tasks that are naturally frontier-driven a
- Can the same runtime interface serve Kubernetes, coding, and browser environments?
- How much state should be sent to System One at each branch?

## 18. Relationship to Narness
## 18. Code localization experiment

The second reference experiment applies the same progressive-disclosure idea to source-code localization.

```text
repository
-> directory candidates
-> relevant directories
-> file candidates
-> relevant files
-> source-line candidates
-> grounded code snippets
```

Unlike Kubernetes action selection, this stage is not a single-winner decision. Multiple directories, files, and source ranges can all be relevant, so the prototype uses independent Noul judgments rather than Choice.

The harness owns traversal, filesystem IO, batching, thresholds, provenance, and range merging. System One only estimates candidate relevance.

This broadens the working runtime abstraction:

```text
StateSpaceGenerator
-> CandidateSet
-> DecisionPrimitive
-> TransitionPolicy
-> EvidenceRecorder
```

The decision primitive can therefore vary by state:

- `Choice` for selecting one grounded action from a local frontier;
- `Noul` for independently retaining multiple relevant candidates;
- deterministic short-circuiting where model judgment is unnecessary.

See [Hierarchical Code Localization with System One](../../../research/code-locator/docs/design.md) for the hypotheses, threshold/recall analysis, evidence contract, and experiment plan.

## 19. Research evidence workflows

The experiments use separate workflows because they exercise different decision primitives, inputs, costs, and failure modes.

- [System One Kubernetes experiment](../../../.github/workflows/system-one-k8s-experiment.yml) owns only the Kubernetes state-machine demo.
- [System One Code Locator experiment](../../../.github/workflows/system-one-code-locator.yml) owns only hierarchical code localization.

The Code Locator workflow has two layers:

1. deterministic fixture validation on Code Locator changes, requiring no external model service;
2. manually dispatched real-System-One localization against a configurable subject repository.

Its artifact contains only Code Locator evidence: run metadata, execution log, `trace.jsonl`, `result.json`, and `summary.md`. This keeps repository-localization research independent from Kubernetes results and makes failures, costs, and parameter sweeps attributable to one experiment.

## 20. Relationship to Narness

Narness currently defines itself as AI Workspace Engineering and explicitly does not aim to become a general Agent Runtime.

Expand All @@ -599,9 +649,13 @@ It is still relevant because it explores a neighboring harness-engineering quest

If the conclusions stabilize, some of them may graduate into Narness concepts around progressive disclosure, capability representation, deterministic execution, policy, evidence, and agent observability without requiring Narness itself to own a runtime.

## Reference implementation
## Reference implementations

See [System One Kubernetes Command Generator](../../../examples/system-one-k8s/README.md).
- [System One Kubernetes Command Generator](../../../examples/system-one-k8s/README.md)
- [System One Code Locator](../../../research/code-locator/README.md)
- [Code localization research note](../../../research/code-locator/docs/design.md)
- [System One Kubernetes experiment workflow](../../../.github/workflows/system-one-k8s-experiment.yml)
- [System One Code Locator experiment workflow](../../../.github/workflows/system-one-code-locator.yml)

## Related Narness topics

Expand Down
219 changes: 219 additions & 0 deletions research/code-locator/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,219 @@
# System One Code Locator

> Experimental research project. This is a localization harness, not a code-changing agent and not part of the canonical Narness runtime.

This directory is the complete Code Locator research unit: implementation, fixtures, tests, design notes, pilot analyses, and immutable run records live together here.

## Project layout

```text
research/code-locator/
├── README.md
├── src/
│ └── system_one_code_locator.py
├── tests/
│ └── test_system_one_code_locator.py
├── fixtures/
│ └── repository/
├── docs/
│ ├── design.md
│ └── pilots/
│ └── nession-websocket-2026-09-23.md
└── runs/
└── <time>-run-<run_id>-attempt-<n>/
```

The `src/`, `tests/`, and `fixtures/` directories define the experiment. `docs/` contains evolving research reasoning. `runs/` is append-only evidence produced by GitHub Actions.

See [design.md](docs/design.md) for the hypotheses and experiment design, and [the first Nession pilot](docs/pilots/nession-websocket-2026-09-23.md) for the first real TypeSafe baseline.

This project explores whether a fast System One model can localize code by repeatedly judging a progressively disclosed search space rather than driving repository browsing through a ReAct loop.

Given a request such as:

> Help me optimize the websocket connection implementation.

and a repository root, the harness performs:

```text
repository
-> enumerate directory candidates
-> System One relevance scores
-> keep directories above threshold
-> enumerate files under retained directories
-> System One relevance scores
-> keep files above threshold
-> split retained files into source lines with small local context
-> System One relevance scores
-> merge adjacent relevant lines into snippets
```

The model never chooses filesystem tools, constructs paths, or controls traversal. The harness owns state expansion, thresholds, IO, provenance, batching, and result assembly.

## Why Noul instead of Choice

Directory, file, and line relevance are independent judgments: several candidates may all be relevant. The TypeSafe System One API's `noul` primitive returns a yes-probability for each question, so the runtime can retain every candidate whose score crosses the stage threshold.

`choice` would instead normalize probability across candidates and force one winner, which is appropriate for the Kubernetes action-frontier demo but not for multi-hit code localization.

## State progression

```text
RepositoryState
|
| deterministic enumeration
v
DirectoryCandidate[]
|
| Noul relevance + low threshold
v
RelevantDirectory[]
|
| deterministic enumeration
v
FileCandidate[]
|
| Noul relevance + medium threshold
v
RelevantFile[]
|
| read source + split on newlines
v
LineCandidate[]
|
| Noul relevance + higher threshold
v
RelevantLine[]
|
| deterministic range merge
v
CodeSnippet[]
```

The defaults intentionally become stricter as the search gets deeper:

```text
directory >= 0.35
file >= 0.50
line >= 0.70
```

These are experiment parameters, not claimed optimal values. Early false negatives are especially expensive because pruning a directory removes its entire subtree. A future experiment should compare simple thresholds against threshold-plus-beam retention.

## Run the deterministic fixture

No API key is required:

```bash
python3 research/code-locator/src/system_one_code_locator.py \
research/code-locator/fixtures/repository \
"Help me optimize the websocket connection implementation" \
--offline-decider \
--line-threshold 0.60
```

The offline scorer is only a deterministic fixture implementation. It is not intended to emulate System One quality.

## Run with TypeSafe System One

```bash
export TYPESAFE_API_KEY=...

python3 research/code-locator/src/system_one_code_locator.py \
/path/to/repository \
"Help me optimize the websocket connection implementation"
```

Optional environment variables:

```bash
export TYPESAFE_API_URL=https://api.typesafe.ai/v1/systemone
export TYPESAFE_MODEL=jev-latest
```

## Research trace

Use `--trace-file` to create an append-only JSONL trace and `--output-json` to persist the final result:

```bash
python3 research/code-locator/src/system_one_code_locator.py \
/path/to/repository \
"Help me optimize the websocket connection implementation" \
--trace-file evidence/code-locator.trace.jsonl \
--output-json evidence/code-locator.result.json
```

The trace records:

- the query, model, root, and thresholds;
- every directory/file/line candidate disclosed by the harness;
- every TypeSafe request body without credentials;
- model responses reduced to candidate scores, latency, and token usage;
- every threshold application and retained candidate set;
- the final result and aggregate metrics.

This makes a workflow run useful as research evidence rather than only as a pass/fail CI event.

## Output contract

The final result contains three provenance-preserving layers:

```text
relevant directories
relevant files
relevant snippets { path, start_line, end_line, score, content }
```

It also reports exposed/selected counts, model calls, token usage, and elapsed time.

## Tests

```bash
python3 -m unittest discover \
-s research/code-locator/tests \
-p 'test_*.py' \
-v
```

The tests verify that the deterministic fixture finds the websocket client and that real API requests use independent Noul questions.

## Dedicated research workflow

Code Locator has its own GitHub Actions workflow:

```text
.github/workflows/system-one-code-locator.yml
```

It does not run the Kubernetes experiment.

On relevant pushes and pull requests it runs only the deterministic Code Locator fixture and uploads:

```text
system-one-code-locator-offline-<run>-<attempt>/
run-manifest.json
tests.log
run.log
result.json
trace.jsonl
summary.md
```

A manual `workflow_dispatch` can additionally run the real TypeSafe experiment against a configurable subject repository. That job uses the `typesafe` environment and uploads a separate `system-one-code-locator-typesafe-*` artifact.

After push or manual runs complete, a persist job writes the evidence permanently under `runs/<time>-run-<run_id>-attempt-<n>/`. Pull-request validation never writes back to the repository.

## Research limitations

The current prototype intentionally leaves several questions open:

- whether whole-tree directory enumeration is better than recursive frontier expansion;
- threshold calibration and compounding false-negative risk;
- threshold-only pruning versus keeping a minimum semantic beam;
- line-level scoring versus symbol/block/chunk-level scoring;
- caching repeated judgments across nearby queries;
- comparison against ripgrep, embeddings, language-server indexes, and System Two browsing;
- evaluation against gold relevant-file and relevant-range labels;
- escalation when no candidate survives a stage.

See the [System One Progressive Action Spaces topic](../../docs/topics/system-one-progressive-action-space/README.md) for the shared model behind this demo and the Kubernetes command generator.
Loading