Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,11 @@
# Changelog

## Cross-problem prompt context and schedule advice, 2026-09-14

- Generation prompts now carry the repo-wide dead-ends ledger (problem-scoped, sanitized) and a ranked cross-problem pattern digest aggregated from `problems/*/patterns/` and `nightly/patterns/`; injected ids are recorded in `evidence.json` under `prompt_context`.
- Plugins declare structural `PATTERN_TAGS` used to rank transferable patterns.
- `scripts/schedule_night.py` scores governed `runs/research/*/run.json` history alongside legacy loop reports; nightly planning orders non-trial research slots by its heuristic and records the advisory allocation as `schedule_plan` in the night status, without touching the counterbalanced trial assignments.

## ARC-AGI-N local companion, 2026-09-08

- Import a bounded, content-hashed local snapshot of reviewed ARC-AGI-N catalogue records without pulling, executing upstream code, or passing upstream prose into model prompts.
Expand Down
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -212,8 +212,8 @@ Supporting scripts:

- `scripts/loop_report.py` — build the dashboard; `--dir` regenerates one for an old run
- `scripts/publish_draft.py` — draft generation for verified record breaks
- `scripts/dead_ends.py` — repo-wide ledger of failed approaches (`list`/`check`/`record`), consulted before each night's candidate design
- `scripts/schedule_night.py` — split the night's compute budget across problems by expected information gain
- `scripts/dead_ends.py` — repo-wide ledger of failed approaches (`list`/`check`/`record`); matching entries are injected into every night's candidate prompt automatically
- `scripts/schedule_night.py` — scores problems by expected information gain; orders non-trial night slots and records the advisory allocation
- `scripts/detached.py` — `status` prints the dashboard path when one exists

## Nightly integration
Expand Down
4 changes: 4 additions & 0 deletions docs/RESEARCH-IMPLEMENTATION.md
Original file line number Diff line number Diff line change
Expand Up @@ -68,6 +68,10 @@ The prompt projection is bounded and sanitized: development provider, actual mod

Auto allocation prioritizes enabled model-role choices with fewer than three attributable development attempts; only after each enabled choice reaches that threshold can it rank a mature primary. It does not change the fixed routing fallback chain and receives no holdout, confirmation, promotion, or reward feedback. This is capacity and attribution work, not evidence that a model is better.

`research_context.py` injects two bounded, sanitized blocks into every generation prompt: the repo-wide dead-ends ledger (`problems/_dead_ends.json`, entries for the active problem plus `general`) and the cross-problem pattern digest (`problems/*/patterns/` and `nightly/patterns/`, ranked by `PATTERN_TAGS` overlap then recorded outcomes). Withheld target names and local paths are stripped before injection, and `evidence.json` records exactly which dead-end ids and pattern names the run saw under `prompt_context`. Plugins declare their structural features via `PATTERN_TAGS`; a problem with no tags still receives the pattern digest ranked by outcomes.

Nightly slot planning consults `scripts/schedule_night.py`: non-trial research slots are ordered by its information-gain heuristic after the counterbalanced trial pair, and the advisory allocation is recorded in the night's status as `schedule_plan`. The trial pair's order and allowances are unchanged; the heuristic scores governed `runs/research/*/<problem>/run.json` history alongside legacy `loop_report.json` runs and can never alter the fixed trial assignments.

The local evidence scan covered 31 solver candidates across `runs-cvrp`, `runs-miplib_heur`, and `runs`: 0 syntax failures and 0 exact AST duplicates. It found seven structured CVRP development-history records, but none had actual-model, model, or role fields. The immediate rationale is therefore prospective: cheap duplicate prevention and reliable attribution before allocation. It has not demonstrated saved compute or a discovery gain. Consciously deferred: near-similarity suppression, bandit allocation, additional model calls, and any holdout-fed reward.

## Dashboard contract
Expand Down
42 changes: 36 additions & 6 deletions loop.py
Original file line number Diff line number Diff line change
Expand Up @@ -29,6 +29,7 @@
from pathlib import Path

import evaluation
import research_context
from model_registry import (
DEFAULT_CHAIN,
MODEL_REGISTRY,
Expand All @@ -40,8 +41,11 @@
from research_memory import (
analyze_candidate,
is_development_observation,
mentions_target,
operational_stats,
rank_auto_allocation,
redact_targets,
strip_local_paths,
summarize_development,
)
from routing import RoutingJournal, route_call, routing_summary
Expand Down Expand Up @@ -339,8 +343,16 @@ def build_research_prompt(
retro_memory=None,
history_total=None,
mission=None,
context_blocks=None,
):
"""Build a prompt from development data only."""
if context_blocks is None:
name = getattr(self, "name", None)
context_blocks = (
research_context.blocks(name, self.P, getattr(self, "root", HERE), hidden_targets)
if name
else {"text": "", "dead_ends": [], "patterns": []}
)
if hasattr(self.P, "prompt_for_targets"):
context = self.P.prompt_for_targets(list(targets))
else:
Expand All @@ -364,11 +376,7 @@ def build_research_prompt(
prior = json.dumps(memory, sort_keys=True, separators=(",", ":")) if memory["entries"] else "(none yet)"
retro = {key: str((retro_memory or {}).get(key, ""))[:1000] for key in ("lessons", "next_experiment")}
for key, value in retro.items():
for target in hidden_targets:
value = value.replace(str(target), "[withheld reference removed]")
retro[key] = re.sub(
r"(?<![:\w])(?:[A-Za-z]:[\\/]|/(?!/))[A-Za-z0-9_.~\\/-]+", "[local path removed]", value
)
retro[key] = strip_local_paths(redact_targets(value, hidden_targets))
retro_text = json.dumps(retro, sort_keys=True, separators=(",", ":")) if any(retro.values()) else "(none yet)"
mission_text = "(no reviewed ARC mission is bound to this run)"
if mission:
Expand Down Expand Up @@ -415,6 +423,8 @@ def build_research_prompt(
PRIOR RETROSPECTIVE NOTES (the next experiment is an untested hypothesis, not evidence):
{retro_text}

{context_blocks['text']}

{self.P.TASK}

Begin the idea with an algorithm-family tag: "IDEA: [kind: <algorithm family>] <one sentence>".
Expand All @@ -424,7 +434,7 @@ def build_research_prompt(
correctness.

OUTPUT FORMAT: the tagged IDEA line, then exactly one ```python block with the full file. Nothing else."""
leaked = [str(target) for target in hidden_targets if str(target) in prompt]
leaked = [str(target) for target in hidden_targets if mentions_target(prompt, target)]
if leaked:
raise ValueError(f"generation prompt exposes withheld targets: {leaked}")
return prompt
Expand Down Expand Up @@ -741,6 +751,23 @@ def _load_problem_for_research(name, root):
return safe_load_problem(name)


def _prompt_context(prior_evidence, problem, plugin, root, hidden_targets):
"""Prompt context for a run, restored verbatim on resume.

The rendered context is recorded in evidence, so a resumed run reuses the
exact text its earlier generations saw instead of rebuilding from ledgers
that may have changed since.
"""
recorded = prior_evidence.get("prompt_context") if isinstance(prior_evidence, dict) else None
if isinstance(recorded, dict) and isinstance(recorded.get("text"), str):
return {
"text": recorded["text"],
"dead_ends": list(recorded.get("dead_ends") or []),
"patterns": list(recorded.get("patterns") or []),
}
return research_context.blocks(problem, plugin, root, hidden_targets)


def run_research(
problem,
provider="paired",
Expand Down Expand Up @@ -980,6 +1007,7 @@ def run_research(
retro_memory = read_json(os.path.join(evidence_base, "development-history", f"{problem}-retro.json"), {}) or {}
if not isinstance(retro_memory, dict) or retro_memory.get("schema_version") not in (None, 1):
retro_memory = {}
prompt_context = _prompt_context(prior_evidence, problem, plugin, root, hidden_targets)
effective_routing_chain = routing_chain
auto_allocation = None
if routing_policy == "auto":
Expand Down Expand Up @@ -1096,6 +1124,7 @@ def run_research(
"usage": usage,
"limitations": list(manifest["limitations"]),
"mission": mission,
"prompt_context": prompt_context,
"legacy_incumbent": {
"path": _repo_relative(incumbent_snapshot, root),
"sha256": _sha256(incumbent_snapshot),
Expand Down Expand Up @@ -1182,6 +1211,7 @@ def append_development_record(record):
retro_memory,
development_memory["total_observations"],
mission,
prompt_context,
)
responses = []
deferred_stop = None
Expand Down
54 changes: 50 additions & 4 deletions night.py
Original file line number Diff line number Diff line change
Expand Up @@ -118,8 +118,8 @@ def load_schedule(path=SCHEDULE):
# scheduled trial policy over the canonical default chain, so normalize in
# memory without requiring a schedule migration.
night["routing"] = routing_config(config)
if not 0 < float(night.get("budget_usd", 0)) <= 90:
raise ValueError("night API-equivalent allowance must be in (0, 90]")
if not 0 < float(night.get("budget_usd", 0)) <= 130:
raise ValueError("night API-equivalent allowance must be in (0, 130]")
if not 1 <= int(night.get("deadline_minutes", 0)) <= 720:
raise ValueError("night deadline_minutes must be in [1, 720]")
modes = {"fable", "astra", "paired"}
Expand All @@ -135,9 +135,17 @@ def load_schedule(path=SCHEDULE):
if sorted(entry.get("order", [])) != ["cvrp", "miplib_heur"]:
raise ValueError("each trial night must order cvrp and miplib_heur once")
slots = config.get("slots", [])
problems = [slot.get("problem") for slot in slots]
if len(set(problems)) != len(problems):
raise ValueError("each problem may appear in at most one slot")
research = {slot.get("problem") for slot in slots if slot.get("kind") == "research"}
if research != {"cvrp", "miplib_heur"}:
raise ValueError("research slots must be exactly cvrp and miplib_heur")
if not {"cvrp", "miplib_heur"} <= research:
raise ValueError("research slots must include cvrp and miplib_heur")
if any(
slot.get("kind") == "research" and slot.get("provider") is not None and slot["provider"] not in modes
for slot in slots
):
raise ValueError("configured slot providers must be fable, astra, or paired")
validation = [slot for slot in slots if slot.get("problem") == "pglib_opf"]
if len(validation) != 1 or validation[0].get("kind") != "validation":
raise ValueError("pglib_opf must appear exactly once and validation-only")
Expand Down Expand Up @@ -185,11 +193,46 @@ def planned_slots(config, run_id):
float(config["night"]["provider_caps_usd"][slot["provider"]]),
)
ordered.append(slot)
# Research slots outside the trial keep their configured provider and run
# in information-gain order before the validation tail.
extras = [
slot
for problem, slot in by_problem.items()
if problem not in assignment["order"] and problem != "pglib_opf" and slot.get("kind") == "research"
]
Comment on lines +198 to +202

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Extra-slot heuristic never receives slots

For validated schedules, extras is always empty. load_schedule rejects every research problem outside the trial pair, making heuristic ordering unreachable.

Learn more

planned_slots is called with configurations produced by load_schedule. That validator requires the set of research problems to equal exactly cvrp and miplib_heur, and the trial assignment always contains both. The new comprehension excludes those two, so no validated configuration can reach sorting or append an extra research slot.

Example: Adding a matrix_multiplication research slot to night.json makes load_schedule raise research slots must be exactly cvrp and miplib_heur. Without that slot, extras remains empty.

Recommended fix: Extend schedule validation to admit bounded non-trial research slots while retaining the required trial pair and validation tail. Validate their providers, budgets, IDs, and total deadline before relying on this ordering path.

Devin Review

Was this helpful? React with 👍 or 👎 to provide feedback.

try:
from scripts.schedule_night import score_problem

extras.sort(key=lambda slot: score_problem(slot["problem"])["score"], reverse=True)
except Exception:
pass # ordering is advisory; never break the night's plan
for slot in extras:
slot["provider"] = slot.get("provider") or "paired"
slot["effective_slot_budget_usd"] = min(
float(slot["slot_budget_usd"]),
float(config["night"]["provider_caps_usd"][slot["provider"]]),
)
ordered.append(slot)
# PGLib is confirmation-only and deliberately has no generation provider.
ordered.append(by_problem["pglib_opf"])
return ordered


def schedule_advice(config, slots):
"""Advisory information-gain allocation for the night's status and morning report."""
try:
from scripts.schedule_night import allocate
except ImportError:
return None
try:
return allocate(
float(config["night"]["deadline_minutes"]) * 60.0,
[slot["problem"] for slot in slots],
)
except Exception:
return None


def _routing_families(config, slots):
"""Return only families that can be selected by the configured night."""
routing = config["night"]["routing"]
Expand Down Expand Up @@ -510,6 +553,7 @@ def run_night(
ledger_path = run_root / "budget.json"
slots = planned_slots(config, run_id)
planned_order = [slot["id"] for slot in slots]
schedule_plan = schedule_advice(config, slots)
arc_plan, arc_summary = _prepare_arc(config, slots, run_id)
requested_next = arc_summary.get("requested_next")
chosen_slot = next(
Expand All @@ -530,6 +574,7 @@ def run_night(
"budget_accounting": config["night"].get("budget_accounting"),
"deadline_minutes": int(config["night"]["deadline_minutes"]),
"slots": slots,
"schedule_plan": schedule_plan,
"routing": {**routing, "override": bool(run_routing_override)},
"arc": arc_summary,
}
Expand Down Expand Up @@ -585,6 +630,7 @@ def run_night(
scheduled_run_id=scheduled_run_id,
)
status["arc"] = arc_summary
status["schedule_plan"] = schedule_plan
disabled_slot_ids = set(arc_summary.get("disabled_slot_ids", []))
for slot_id, mission in arc_plan.items():
plugin = mission["plugin"]
Expand Down
1 change: 1 addition & 0 deletions problems/circle_packing/problem.py
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,7 @@
VALIDATION = []
RELEASE_HOLDOUT = []
DEFAULTS = {"time": 120, "workers": 3}
PATTERN_TAGS = ["continuous", "fast-verifier", "local-search", "generatable-test-cases"]
MAXIMIZE = True
FAIL_SCORE = 0.0
WIN_MARGIN = 1e-10
Expand Down
1 change: 1 addition & 0 deletions problems/cvrp/problem.py
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,7 @@
VALIDATION = []
RELEASE_HOLDOUT = []
DEFAULTS = {"time": 120, "workers": 3}
PATTERN_TAGS = ["local-search", "fast-verifier", "combinatorial", "route-structure", "generatable-test-cases", "huge-raw-search-space"]
MAXIMIZE = False
FAIL_SCORE = -1.0 # a crash / timeout / infeasible output; strictly worse than any feasible run (gap clipped at 0.5)
GAP_CLIP = 0.5
Expand Down
1 change: 1 addition & 0 deletions problems/matrix_multiplication/problem.py
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,7 @@
VALIDATION = []
RELEASE_HOLDOUT = []
DEFAULTS = {"time": 300, "workers": 1}
PATTERN_TAGS = ["block-structure", "disjoint-outputs", "subproblems", "recursive-structure", "multiplicative-cost", "technique-library", "composition-operators", "huge-raw-search-space", "fast-verifier", "generatable-test-cases"]
MAXIMIZE = False
FAIL_SCORE = -1.0 # crash / timeout / infeasible output; worse than any feasible run
GAP_CLIP = 0.5
Expand Down
1 change: 1 addition & 0 deletions problems/miplib/problem.py
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,7 @@
"neos-5045105-creuse": "3848 vars / 252 rows, integer knapsacks, general integers",
}
DEFAULTS = {"time": 400, "workers": 3}
PATTERN_TAGS = ["local-search", "combinatorial", "huge-raw-search-space"]
MAXIMIZE = False
FAIL_SCORE = -10.0
REL_TOL = 1e-6
Expand Down
1 change: 1 addition & 0 deletions problems/miplib_heur/problem.py
Original file line number Diff line number Diff line change
Expand Up @@ -46,6 +46,7 @@ def _desc(name):

INFO = {t: _desc(t) for t in TARGETS}
DEFAULTS = {"time": 60, "workers": 3}
PATTERN_TAGS = ["local-search", "fast-verifier", "combinatorial", "huge-raw-search-space", "generatable-test-cases"]
MAXIMIZE = False
FAIL_SCORE = -1.0 # added to the champion total directly (score space), so a failed target costs a 100% gap
WIN_MARGIN = 1e-4 # gap must improve on HiGHS default by 0.01% of the objective to count (timing noise floor)
Expand Down
1 change: 1 addition & 0 deletions problems/miplib_open/problem.py
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,7 @@
VALIDATION = []
RELEASE_HOLDOUT = []
DEFAULTS = {"time": 600, "workers": 3} # justified from a measured seed run in BASELINE.md
PATTERN_TAGS = ["local-search", "combinatorial", "huge-raw-search-space"]
MAXIMIZE = False # value is min-sense (lower is better); the loop maximises total = minus the summed value
FAIL_SCORE = -1.0 # a crash / timeout / infeasible output, in score space (worse than any clipped feasible gap)
GAP_CLIP = 1.0 # one hopeless instance (100% above best-known) cannot dominate the champion total
Expand Down
1 change: 1 addition & 0 deletions problems/pglib_opf/problem.py
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,7 @@
VALIDATION = []
RELEASE_HOLDOUT = []
DEFAULTS = {"time": 90, "workers": 4}
PATTERN_TAGS = ["continuous", "fast-verifier", "local-search"]
MAXIMIZE = False
FAIL_SCORE = -1.0
WIN_MARGIN = 1e-4 # conservative preliminary screen; release validation uses each row's exact printed uncertainty
Expand Down
Loading
Loading