Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions .vulture_whitelist.py
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,9 @@
_.VALIDATION_TARGETS
_.HOLDOUT
_.DEFAULTS
# research_context reads tags through getattr; plugin integrity tests check capabilities.
_.PATTERN_TAGS
_.RELEASE_VALIDATION_SUPPORTED
_.MAXIMIZE
_.FAIL_SCORE
_.TOTAL_DESC
Expand Down Expand Up @@ -47,6 +50,8 @@
_.retro_slot
_.publish_slot
_.official_solution_path
# Isolation tests inspect the bounded worker mount through this helper.
_._mount_source
# pytest invokes this autouse fixture by registration.
_.subscription_auth
# Canonical provider API callers may be outside an incremental staged-file scan.
Expand Down
12 changes: 12 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,17 @@
# Changelog

## Matrix multiplication joins the governed night, 2026-09-14

- `night.json` gains a fixed-provider `matrix_multiplication` research slot (90 min, 12+3 units) that runs after the counterbalanced trial pair in information-gain order; night allowance is now 105 units over 540 minutes.
- `load_schedule` accepts extra research slots beyond the trial pair: problems must be unique, required trial problems present, configured providers limited to fable/astra/paired.
- `evaluation.build_manifest` supports `CONFIRMATION_ON_DEVELOPMENT`: confirmation re-runs the development targets under fresh seeds (classification `same_target_fresh_seed_replication`), and the manifest now reports `concealed` targets separately from `confirmation`.
- `matrix_multiplication.prompt_for_targets` keeps the interface contract, strategy notes and honest framing while dropping lines that name withheld targets.
- Dashboard allowance validation matches the scheduler's 130-unit ceiling, so the new 105-unit default can be saved without reducing existing slot allowances.
- The task installer prepares a 21:00 start and 9h15m scheduler limit, with the runner still stopping at 06:00. Existing Windows task registrations require separate activation; morning jobs keep their times.
- Matrix-multiplication workers receive the allowlisted exact verifier, so the incumbent and generated solvers can execute through Docker isolation.
- Matrix multiplication accepts a replicated gain on one target when no matched evaluation regresses. Other plugins retain their median gate; evidence keeps the overall median and reports target gains separately. Local incumbent advancement remains separate from publication.
- Partial research runs no longer count as successful runs in advisory scheduling history.

## Cross-problem prompt context and schedule advice, 2026-09-14

- Generation prompts now carry the repo-wide dead-ends ledger (problem-scoped, sanitized) and a ranked cross-problem pattern digest aggregated from `problems/*/patterns/` and `nightly/patterns/`; injected ids are recorded in `evidence.json` under `prompt_context`.
Expand Down
15 changes: 8 additions & 7 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -87,7 +87,7 @@ Use the activated virtual environment for the commands below. If PowerShell bloc

Runtime and worker dependencies are pinned. No API key is required. Provider preflight rejects API-key authentication and does not silently fall back to API billing. The registry maps Fable and Opus to the Anthropic subscription CLI, and Astra and Sol to the OpenAI subscription CLI; tools and external integrations are disabled for those calls.

The default nightly **research allowance is 90 accounting units**, shared across generation, reviews and retrospectives. This is not a cash budget. Claude's reported API-equivalent cost is an estimate of usage; Codex calls without a dollar estimate conservatively consume their reservation. Calls and token usage are retained. Subscription rate limits still apply.
The default nightly **research allowance is 105 accounting units**, shared across generation, reviews and retrospectives. This is not a cash budget. Claude's reported API-equivalent cost is an estimate of usage; Codex calls without a dollar estimate conservatively consume their reservation. Calls and token usage are retained. Subscription rate limits still apply.

## Running research

Expand Down Expand Up @@ -172,7 +172,7 @@ flowchart TD
6. For general MIP heuristics, compare against a freshly executed HiGHS baseline in the same worker environment before making a baseline-superiority claim.
7. Preserve evidence and confirmed lineage for subsequent nights. Feed only development observations and sanitized lessons into future generation.

Known benchmark targets remain labeled previously exposed. The existing MIP heuristic holdout is reusable confirmation data, not a sealed generalization test. There is currently no sealed release dataset.
Known benchmark targets remain labeled previously exposed. The existing MIP heuristic holdout is reusable confirmation data, not a sealed generalization test. Matrix multiplication confirms on the same development targets under fresh seeds, which measures solver repeatability rather than unseen generalization. There is currently no sealed release dataset.

## Problems

Expand All @@ -183,6 +183,7 @@ Known benchmark targets remain labeled previously exposed. The existing MIP heur
| miplib_open | Open mixed-integer programs | Original bounds, integrality, row activities and objective |
| miplib | Legacy open-instance experiments | Original MPS and uncertainty-aware record comparison |
| pglib_opf | AC power-flow validation | Original-case residuals at 1e-8, baseline rounding uncertainty and reference polishing |
| matrix_multiplication | Exact bilinear rank search | Exact tensor-identity verification; a verified rank below the best known is a benchmark record |
| circle_packing | Geometric optimization | Finite values, containment, separation and an explicit improvement margin |

The default nightly trial focuses on routing and general optimization, with a validation-only power-grid stage. See [research portfolio](docs/RESEARCH-PORTFOLIO.md) for intended beneficiaries, success measures, and evidence needed before claiming practical benefit.
Expand Down Expand Up @@ -218,25 +219,25 @@ Supporting scripts:

## Nightly integration

`night.json` controls an eight-hour window, per-slot and per-call limits, a local ARC snapshot refresh, and a 14-night counterbalanced Fable/Astra/paired trial. The runner uses an exclusive lock, checkpoints, heartbeat, pause handling, process-tree timeouts and explicit zero-work/partial/failure statuses. Installed `--scheduled` runs use `<local-evening-date>-scheduled`; this prevents a completed manual date-named run from suppressing the scheduled night while retaining the logical date for trial assignment and morning reporting.
`night.json` controls a nine-hour window, per-slot and per-call limits, a local ARC snapshot refresh, and a 14-night counterbalanced Fable/Astra/paired trial. The runner uses an exclusive lock, checkpoints, heartbeat, pause handling, process-tree timeouts and explicit zero-work/partial/failure statuses. Installed `--scheduled` runs use `<local-evening-date>-scheduled`; this prevents a completed manual date-named run from suppressing the scheduled night while retaining the logical date for trial assignment and morning reporting.

| Stage | Maximum time | Research allowance | Retrospective allowance |
| --- | ---: | ---: | ---: |
| Routing research | 180 min + 30 min retrospective | 40 | 5 |
| General MIP heuristic research | 180 min + 30 min retrospective | 40 | 5 |
| Matrix-multiplication rank search | 70 min + 20 min retrospective | 12 | 3 |
| Power-grid validation only | 30 min | 0 | 0 |
| Unallocated time buffer | 30 min | 0 | 0 |
| **Night limit** | **480 min** | **90 units total across all calls** | **Included** |
| **Night limit** | **540 min** | **105 units total across all calls** | **Included** |

Research order alternates. Each track receives five Fable, five Astra and four paired requested arms per cycle. Equal configured allowances do not imply equal tokens or equivalent subscription consumption; the trial is exploratory. A fallback or routing override is useful operational evidence but is excluded from clean formal-trial comparisons.
Research order alternates. Each track receives five Fable, five Astra and four paired requested arms per cycle. The matrix-multiplication slot is not part of the counterbalanced trial: it keeps its configured paired provider and runs after the trial pair in information-gain order. Equal configured allowances do not imply equal tokens or equivalent subscription consumption; the trial is exploratory. A fallback or routing override is useful operational evidence but is excluded from clean formal-trial comparisons.

On Windows, preview the scheduled-task changes first:

```powershell
powershell -NoProfile -ExecutionPolicy Bypass -File scripts/install-night-tasks.ps1
```

The installer exports existing XML before any change. Its `-Apply` switch updates the 22:00 research task, connects the 06:40 meditation and 06:57 briefing to fresh evidence, and installs the localhost dashboard at logon. Rollback commands are printed with the backup paths. Scheduled catch-up is restricted to the overnight window. The original meditation runner is reused with sanitized research context injected in memory; no harness source file is modified.
The installer exports existing XML before any change. Its `-Apply` switch configures research for 21:00–06:00 with a 9h15m scheduler limit, connects the 06:40 meditation and 06:57 briefing to fresh evidence, and installs the localhost dashboard at logon. Existing 22:00 installations need a separately approved task update to provide the full nine-hour window; merging or pulling code does not change Windows task registration. Rollback commands are printed with the backup paths. Scheduled catch-up is restricted to the overnight window. The original meditation runner is reused with sanitized research context injected in memory; no harness source file is modified.

A missing or partial research run is explicitly reported to meditation. The briefing requires a current meditation artifact rather than silently reusing yesterday's. The existing briefing's external delivery behavior is unchanged; installation does not send a message.

Expand Down
2 changes: 1 addition & 1 deletion dashboard.py
Original file line number Diff line number Diff line change
Expand Up @@ -393,7 +393,7 @@ def update_schedule(self, payload: dict[str, Any]) -> dict[str, Any]:
duration = payload["duration_minutes"]
if isinstance(duration, bool) or not isinstance(duration, int) or not 60 <= duration <= 720:
raise ApiError(HTTPStatus.BAD_REQUEST, "invalid_payload", "Duration must be a whole number from 60 to 720.")
budget = _number(payload["nightly_budget_usd"], "Nightly research allowance", 0, 90)
budget = _number(payload["nightly_budget_usd"], "Nightly research allowance", 0, 130)
if budget <= 0:
raise ApiError(
HTTPStatus.BAD_REQUEST, "invalid_payload", "Nightly research allowance must be greater than zero."
Expand Down
8 changes: 8 additions & 0 deletions docs/DECISIONS.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,13 @@
# Research decisions

## 2026-09-14: Governed slot for the matmul frontier

Promote `matrix_multiplication` into `night.json` as a fixed-provider research slot outside the counterbalanced trial: it keeps its configured provider, is ordered by the information-gain heuristic after the trial pair, and runs before pglib validation. The nightly allowance rises from 90 to 105 accounting units and the deadline from 480 to 540 minutes to fund it; trial slot allowances are untouched so the 14-night comparison stays clean.

Confirmation semantics differ from the benchmark plugins: the solver is stochastic and the incumbent already emits verified decompositions for every target, so confirmation re-runs the same development targets under fresh seeds (`CONFIRMATION_ON_DEVELOPMENT`, classification `same_target_fresh_seed_replication`). That measures repeatability, not unseen generalization, and nothing is withheld from prompts; the exact tensor-identity verifier and the incumbent gate remain the primary evidence. Matrix multiplication uses an explicit per-target Pareto gate: at least one target must improve by the minimum effect, no matched target/seed evaluation may regress, candidate evaluations must all succeed, and the required fresh seeds must complete. This lets a real improvement on one open target advance the incumbent without allowing the proven-optimal n=2 calibration to deteriorate.

Incumbent advancement remains separate from publication. The current all-target release gate cannot mark a matrix-multiplication run publishable because n=2 is already proven optimal and therefore cannot beat its record. A confirmed solver can still advance local research, but publication of an individual record-breaking target needs a separately reviewed, target-scoped release path; this change does not broaden publication.

## 2026-09-07: Route execution separately from trial assignment

Keep the historical Fable/Astra/paired arm as the scheduled experiment identity, while recording the exact configured model and provider family that executed each physical call. The registry is Fable/Opus in Anthropic and Astra/Sol in OpenAI. The default route is Fable, Opus, Astra, Sol, with requested-model then same-family then other-family fallback. Policies may restrict the route or preserve the scheduled arm; disabled families are explicit.
Expand Down
12 changes: 12 additions & 0 deletions docs/ERRORS.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,17 @@
# Errors and lessons

## 2026-09-14: Nightly allowance exceeded dashboard validation

The matrix-multiplication slot raised the default allowance to 105 units while the dashboard form and API still capped it at 90. Saving the unchanged schedule failed despite the scheduler accepting it. Aligning both dashboard limits with the scheduler's 130-unit ceiling restores settings saves. Schedule changes now receive an unchanged-default save check through the human control surface as well as the runner's dry run.

The same review found that a 540-minute plan was still constrained by a 22:00 task start and the runner's 06:00 cutoff. The prepared installer now starts at 21:00 with a 9h15m task limit, preserving the morning jobs. Verification covers the catch-up boundary and actual installer values; existing task activation remains separate from code delivery. Historical activation evidence retains the values actually checked at that time.

A real worker probe rejected `matrix_multiplication` before execution because the Docker input allowlist omitted the plugin. Adding its trusted `verify.py` helper enables the existing worker path without mounting other repository files. New governed slots receive one real isolated incumbent evaluation before their schedule is considered runnable.

The default median across all three matrix sizes also rejected a single-target improvement when the other sizes tied. Matrix multiplication now opts into a target-aware gate requiring a replicated improvement and no matched-case regressions. Synthetic tests cover both acceptance and rejection, and keep local promotion distinct from the existing all-target release gate. Future plugin admission checks include an achievable improvement case and a regression case, alongside real worker execution.

The incremental commit hook omitted callers in unstaged files and flagged the worker mount test helper, dynamic pattern tags, and plugin capability marker. Their callers were verified before adding these names to the existing Vulture whitelist; the hook remains enabled.

## 2026-09-08: Real catalogue and scheduler probes caught fixture-shaped assumptions

The first ARC import rejected a valid 40-character Git commit because its validator incorrectly reused the 64-character SHA-256 pattern. Splitting Git object validation from content-hash validation fixed the real import. The same review found that a cached snapshot could have been edited after validation and that its normalized hash changed with every import timestamp. Cache loads now recompute a timestamp-independent normalized hash, compare executable admissions to the local reviewed bindings, and retain a raw hash over every source byte. The live checkout also records `worktree_dirty` because its two integration cards were local additions beyond the cited commit.
Expand Down
Loading
Loading