Skip to content

docs(validation): re-score prospective N=28 on the current engine (DE-54) — no OOD improvement#102

Merged
jam-sudo merged 2 commits into
mainfrom
docs/prospective-current-engine-rescore
Jul 5, 2026
Merged

docs(validation): re-score prospective N=28 on the current engine (DE-54) — no OOD improvement#102
jam-sudo merged 2 commits into
mainfrom
docs/prospective-current-engine-rescore

Conversation

@jam-sudo

@jam-sudo jam-sudo commented Jul 5, 2026

Copy link
Copy Markdown
Owner

What

Closes the README "predates FLUX-1, not re-scored on the current engine" note on the 2026-06-01 prospective set with an actual measurement, and records the (negative) result across the research docs.

Result

Re-ran the 28-drug public-clone prospective set through the current engine (HEAD f8a4663: FLUX-1 + CLF leak-free + UGT single-path + RBP honest-1.0) via scripts/score_prospective_candidates.py:

track OLD (pre-FLUX-1) NEW (current) Δ
meta 3.2084 3.2861 +0.078
engine 4.3017 4.5506 +0.249
ml 3.2704 3.2704 0.000 (bit-identical)
in-domain meta (N=16) 3.1988 3.3177 +0.119

Deep inside the N=28 CI [2.42, 4.37] → statistically indistinguishable, no improvement.

Why it matters (clean comparison + mechanism)

  • Same-stack proof: the ML track is per-drug bit-identical (max abs diff 0.0), so the 2026-06-01 artifact shares this local macOS numerics stack. The delta isolates engine/predict code, not a stack artifact.
  • Mechanism — error cancellation removed: prospective is under-predicted (geobias pred/obs 0.50). The pre-FLUX-1 E-cap held Cmax too high (offsetting the fa under-call); FLUX-1 correctly raised E → lowered Cmax → deepened the under-call (concentrated in high-extraction drugs, e.g. arimoclomol engine fold 4.96→16.10). This is the first actual full-set confirmation of the 2026-06-03 diagnosis §8 prediction / DE-43: the fixed-weight meta damps the engine's +0.25 to +0.08. Residual OOD deficit is absorption-side (fa) dispersion (DE-42 floor), not extraction.

Changes

  • New artifact data/validation/prospective_N28_current_engine_2026-07-05.json (full per-drug + provenance + comparison block).
  • DE-54 in dead-ends.md (negative result + telltale).
  • 2026-07-05 entry in experiment-log.md.
  • §8 confirmation paragraph in diagnosis.md.

Headline 2.743 untouched. The local prospective number is directional (sign + mechanism robust); a citable canonical number would need a CI-stack run.

Note

The public README's prospective-row note ("predates FLUX-1, not re-scored") is now stale; a maintainer edit to reference this re-score is provided separately.

…-54)

Closes the README 'predates FLUX-1, not re-scored on the current engine' note
on the 2026-06-01 prospective set with an actual number, and records the result.

Re-ran the 28-drug public-clone prospective set through the current engine
(HEAD f8a4663: FLUX-1 + CLF leak-free + UGT single-path + RBP honest-1.0) via
scripts/score_prospective_candidates.py. Result: Meta AAFE 3.2084 -> 3.2861
(+0.078), engine 4.3017 -> 4.5506 (+0.249), ML 3.2704 -> 3.2704 (per-drug
bit-identical). Deep inside the N=28 CI [2.42, 4.37] -> no improvement,
statistically indistinguishable. In-domain (N=16) meta 3.1988 -> 3.3177.

The comparison is clean: the ML track being per-drug bit-identical proves the
old artifact shares this local numerics stack, so the delta isolates engine
code. Mechanism: prospective is under-predicted (geobias 0.50); FLUX-1
correctly raised extraction -> lowered Cmax -> removed the E-cap's compensating
over-estimate -> deepened the under-call. Residual OOD deficit is absorption
(fa) dispersion (DE-42 floor), not extraction.

Adds the re-scored artifact + DE-54 (dead-ends) + a 2026-07-05 experiment-log
entry + a diagnosis section 8 confirmation. Headline 2.743 untouched.
@jam-sudo
jam-sudo enabled auto-merge (squash) July 5, 2026 17:31
… prospective section

Keeps the published 3.21 headline; adds that re-scoring the N=28 set on the
current engine gives 3.29 (no improvement, within CI), with the
error-cancellation mechanism (extraction fix deepens the fa-limited OOD
under-call).

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 34ff70d904

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".


**2026-06-03 — the prospective set is also co-calibration-locked; the meta damps the engine to ~18% everywhere (dead-ends.md DE-43).** The natural follow-up — "the *prospective* N=28 set (3.21) isn't in the meta co-calibration, so a first-pass lever foreclosed retrospectively should help it" — was tested and **falsified**. A `F = fa·Fg·Fh` decomposition shows the prospective catastrophes (mirdametinib, sevabertinib, kinase inhibitors) are **fa-first** (absorption starved: low Peff or low-solubility → `particle_radius=50µm` → `ka ≪` gut transit) then **gut-CYP3A `Fg`-second** — *not* the fa-saturated pure-CYP3A mode. Two levers (absorption scalar 5.25×; gut-CYP3A 0.5×) each improve the **engine** track on prospective (4.11→3.75 / 4.11→4.00; mirdametinib engine fold 58→13) — but the fixed-weight **meta-learner damps the change to ~18–19% pass-through, identically on prospective and retrospective**, so the production meta nets neutral-to-negative (absorption scalar: prospective −0.069 *vs* retrospective +0.082 = **net −0.012**; gut-CYP3A: +0.020, inside the N=28 CI). **The meta is robust to engine errors by construction, which symmetrically blocks engine improvements — on every benchmark.** This is the unifying mechanism behind the error-cancellation dead-ends (§2): the engine is structurally not a headline lever, retrospective *or* prospective. The only remaining levers are per-drug **measured-F routing** (a capability, not a global move) or a **meta-architecture change** (AD-gated engine weighting — itself likely foreclosed per the meta-learner dead-ends DE-23/24/25 and the non-generalizing AD signal DE-41, and N=28-underpowered to validate).

**2026-07-05 — confirmed by the actual full-set re-score (dead-ends.md DE-54).** The 2026-06-03 prediction above was, until now, an extrapolation from a decomposition; it has been tested directly. The N=28 prospective set, re-scored end-to-end on the post-FLUX-1 engine (+CLF/UGT/RBP; HEAD `f8a4663`), moves **Meta AAFE 3.21 → 3.29** and **engine 4.30 → 4.55** — the engine's +0.25 damped to +0.08 at the meta, exactly the ~18% pass-through predicted, and *worse* not better. The comparison is clean: the ML track is per-drug bit-identical (max abs diff 0.0), so the 2026-06-01 artifact shares this local numerics stack and the delta isolates engine/predict code. Mechanism: prospective is under-predicted (geobias pred/obs 0.50); FLUX-1 correctly raised extraction → lowered Cmax → *removed a compensating over-estimate* (the pre-fix `E`-cap held Cmax too high), deepening the net under-call. The residual is absorption-side `fa` dispersion (the DE-42 floor), not extraction. Artifact: `data/validation/prospective_N28_current_engine_2026-07-05.json`.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Fix the pass-through ratio claimed by DE-54

Using the AAFE deltas cited in this same sentence, the current re-score does not show the predicted ~18% damping: 0.0777 / 0.2489 ≈ 31% (and the log-AAFE change is ≈43%). Since this paragraph records DE-54 as a full-set confirmation of the DE-43 mechanism, future research decisions will be gated on an overstated mechanism unless the text either computes the pass-through on the intended per-drug/weight basis or stops claiming the aggregate result is “exactly” ~18%.

Useful? React with 👍 / 👎.

@jam-sudo
jam-sudo merged commit f52b1f5 into main Jul 5, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant