Wire dead-ends, pattern library and schedule heuristic into the governed loop - #7
Conversation
…ned loop - New research_context.py injects two sanitized blocks into every generation prompt: problem-scoped entries from problems/_dead_ends.json and a ranked cross-problem pattern digest from problems/*/patterns/ + nightly/patterns/. Injected ids/names are recorded in evidence.json under prompt_context. - Plugins declare PATTERN_TAGS used to rank transferable patterns. - Hidden-target sanitization and the prompt leak check are now token-boundary aware, so single-character targets (e.g. "4") no longer mangle text like de-004 or 400s or false-flag a prompt. - schedule_night scores governed runs/research history as well as legacy loop reports; night planning orders non-trial research slots by its information-gain heuristic and records the advisory allocation as schedule_plan in night status and dry-run output. Trial order, providers and allowances are unchanged. Co-Authored-By: Wes Sander <sandman.uc@gmail.com>
🤖 Devin AI EngineerI'll be helping with this pull request! Here's what you should know: ✅ I will automatically:
Note: I can only respond to comments from users who have write access to this repository. ⚙️ Control Options:
|
There was a problem hiding this comment.
Devin Review found 4 potential issues.
1 flag not posted on this PR by your GitHub settings — view it in Devin Review. (Configure)
| retro_memory = read_json(os.path.join(evidence_base, "development-history", f"{problem}-retro.json"), {}) or {} | ||
| if not isinstance(retro_memory, dict) or retro_memory.get("schema_version") not in (None, 1): | ||
| retro_memory = {} | ||
| prompt_context = research_context.blocks(problem, plugin, root, hidden_targets) |
There was a problem hiding this comment.
🟡 Resume changes recorded prompt context
When ledgers change after interruption, prompt_context is recomputed for resumed generations. Evidence then replaces the original context, breaking prompt lineage.
Learn more
A resumable run can have prior evidence with pending or completed candidate work. The prompt context is computed from mutable repository ledgers before that prior evidence is restored. If either ledger changed, subsequent generations use a different prompt context under the same run ID, and the newly constructed evidence records only the replacement IDs. This violates the resume lineage preserved for incumbents, targets, routing, and candidates elsewhere in run_research.
Example: A run starts with dead end de-7, then stops after one candidate. An operator adds de-8 before resuming. The resumed iteration sees both entries, and evidence.json says the run used both even though its first iteration never saw de-8.
Recommended fix: Persist the rendered, sanitized context or a hash-bound snapshot in evidence at initial creation. On resume, restore that snapshot and validate it instead of rebuilding from live ledgers.
Was this helpful? React with 👍 or 👎 to provide feedback.
| line = f"- {pattern['name']} (from {pattern['origin']}, {pattern['outcomes']} outcome(s)" | ||
| if overlap: | ||
| line += f", {overlap} matching tag(s)" | ||
| line += f"): {pattern['description']}" | ||
| if pattern["ref"]: | ||
| line += f" [{pattern['ref'][:100]}]" |
There was a problem hiding this comment.
🟡 Pattern prompts include forbidden feedback
Every selected pattern injects its outcomes count and optional transform_ref. These reward outcomes and paths violate the development-memory prompt contract.
Learn more
The repository contract prohibits feeding paths or reward outcomes back into development-memory prompts. Pattern records carry successes and failures, and several transform_ref values identify repository modules or files. The renderer exposes the aggregate outcome count and the reference directly to every generation call. Sanitizing absolute paths does not remove these relative repository paths.
Example: composition-search contributes its promotion-derived outcome count and problems/matrix_multiplication/smart_loop.py reference. A later model can use both the reward signal and repository location when proposing a candidate.
Recommended fix: Use outcome counts only for local ranking. Render neither counts nor transform_ref; include only sanitized names, descriptions, origins, and structural applicability.
Was this helpful? React with 👍 or 👎 to provide feedback.
| if not isinstance(record, dict) or not record.get("name") or record["name"] in seen: | ||
| continue | ||
| seen.add(record["name"]) | ||
| origin = record.get("origin_problem") or os.path.basename(os.path.dirname(directory)) | ||
| patterns.append(_normalize_pattern(record, origin, path)) |
There was a problem hiding this comment.
🟡 Duplicate patterns lose structural tags
For duplicate names, seen keeps the top-level record and drops the problem-specific record. Exact applies_when tags vanish, so PATTERN_TAGS ranks unrelated patterns first.
Learn more
The scan sorts directories lexically, so nightly/patterns precedes problems/.../patterns. Several names exist in both locations. The nightly schema derives loose word tokens from prose, while the problem schema has exact structural tags and richer current counts. Keeping the first name therefore discards the metadata needed by patterns_for.
Example: For matrix multiplication, the nightly block-decomposition record yields tokens such as block and subproblem, not the plugin tag block-structure. The problem-specific record containing block-structure is skipped, while circle-packing patterns with another matching tag can rank ahead.
Recommended fix: Merge duplicate records deterministically or prefer the problem-specific schema for tags and counts. Preserve a canonical record per name only after combining exact applies_when metadata.
Was this helpful? React with 👍 or 👎 to provide feedback.
| extras = [ | ||
| slot | ||
| for problem, slot in by_problem.items() | ||
| if problem not in assignment["order"] and problem != "pglib_opf" and slot.get("kind") == "research" | ||
| ] |
There was a problem hiding this comment.
🟡 Extra-slot heuristic never receives slots
For validated schedules, extras is always empty. load_schedule rejects every research problem outside the trial pair, making heuristic ordering unreachable.
Learn more
planned_slots is called with configurations produced by load_schedule. That validator requires the set of research problems to equal exactly cvrp and miplib_heur, and the trial assignment always contains both. The new comprehension excludes those two, so no validated configuration can reach sorting or append an extra research slot.
Example: Adding a matrix_multiplication research slot to night.json makes load_schedule raise research slots must be exactly cvrp and miplib_heur. Without that slot, extras remains empty.
Recommended fix: Extend schedule validation to admit bounded non-trial research slots while retaining the required trial pair and validation tail. Validate their providers, budgets, IDs, and total deadline before relying on this ordering path.
Was this helpful? React with 👍 or 👎 to provide feedback.
…pts, restore prompt context on resume, admit bounded non-trial research slots - load_patterns merges same-named records (tag union, summed outcomes) so the problem-schema applies_when tags survive a looser nightly record. - Pattern lines render name, origin, matching-tag count and description only: outcome counts and transform_ref paths stay local and never reach a prompt. - run_research records the rendered prompt_context in evidence and restores it verbatim on resume, so ledger changes mid-run cannot rewrite prompt lineage. - load_schedule admits extra research slots beyond the trial pair: problems unique, required trial problems present, configured providers in fable/astra/paired. This makes the information-gain ordering path reachable. Co-Authored-By: Wes Sander <sandman.uc@gmail.com>
Summary
Three pieces of meta-loop infrastructure existed but were never consulted by the governed pipeline: the repo-wide dead-ends ledger (
problems/_dead_ends.json), the cross-problem pattern library (problems/*/patterns/,nightly/patterns/), and the information-gain scheduler (scripts/schedule_night.py). This wires all three in, so a night's candidate design actually sees proven failures, transferable techniques, and ordering advice — closing the learning loop the README already described.What changed
research_context.py(new): builds a bounded prompt block per run —dead_ends_for(problem, ...): newest entries scoped to the active problem plusgeneral, rendered as[id] approach -- failed: why (tags).patterns_for(problem, plugin, ...): normalizes both pattern schemas (applies_when/successesandapplicability/success_count), ranks by overlap with each plugin's newPATTERN_TAGSthen recorded outcomes.blocks()returns the rendered text plus the injected ids, whichrun_researchrecords inevidence.jsonasprompt_context(audit trail of exactly what the model saw).loop.py:build_research_prompt(..., context_blocks=None)injectsKNOWN DEAD ENDS+TRANSFERABLE PATTERNSsections between the retro notes andTASK; computed once per run and threaded through so evidence, prompt and records agree.research_memory.py: target redaction and the prompt leak check are now token-boundary aware (redact_targets/mentions_target/strip_local_paths). Previously a single-character withheld target like"4"would have corrupted text (de-004→de-00[withheld],400s→[withheld]00s) and false-flagged every prompt — a latent blocker for running matrix_multiplication under the governed loop.scripts/schedule_night.py:_historynow also scores governedruns/research/*/<problem>/run.jsonruns, not just legacyloop_report.jsonruns.night.py:planned_slotsorders non-trial research slots (none configured yet) byscore_problemafter the counterbalanced pair and before the pglib tail;run_nightrecordsschedule_planin status and dry-run output. Trial providers, order, and allowances are untouched — scoring is advisory and can never break the plan (guarded import, try/except).problem.pygainsPATTERN_TAGSdescribing its structure (e.g. cvrp →local-search,route-structure; matmul →block-structure,disjoint-outputs,recursive-structure).docs/RESEARCH-IMPLEMENTATION.mdmemory section, CHANGELOG entry, README script descriptions.Verification
python scripts/check.py: 259 tests pass (8 new intests/test_research_context.py), Ruff clean, 76 files compile. The new tests cover scoped/sanitized dead-end injection, both pattern schemas,prompt_contextin evidence, non-trial slot ordering, governed-run scoring, and the dry-runschedule_plan.Link to Devin session: https://app.devin.ai/sessions/a2d291aa33784f92aee3fc4c66551893
Open in Devin Desktop: https://app.devin.ai/desktop/session/a2d291aa33784f92aee3fc4c66551893?variant=devin
Requested by: @ucsandman