Skip to content
Open

merge #1044

Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
36 commits
Select commit Hold shift + click to select a range
9a5f199
fix: block orphan lab-result joins
streamentry Jul 19, 2026
5de8b1a
feat: expose scientific review readiness gate
streamentry Jul 20, 2026
2c1fbfd
fix: verify lab result certificate identities
streamentry Jul 20, 2026
b8fe61d
fix: separate raw and usable lab summaries
streamentry Jul 21, 2026
3e445b1
fix: verify lab result panel identity
streamentry Jul 21, 2026
50ecd55
expose raw assay provenance coverage
streamentry Jul 22, 2026
e465de5
feat: expose Phase Z accountability gate
streamentry Jul 22, 2026
4b38fe4
fix: validate external review calendar dates
streamentry Jul 23, 2026
7de2ea4
feat: bind review outcomes to frozen packages (#9)
streamentry Jul 23, 2026
c895a1b
Verify declared raw assay file hashes (#10)
streamentry Jul 24, 2026
6462efd
fix: validate lab result calendar dates
streamentry Jul 24, 2026
2bd3736
feat: expose baseline accountability gate (#12)
streamentry Jul 25, 2026
f06f17a
feat: expose claim integrity gate workflow (#13)
streamentry Jul 25, 2026
1279667
fix: block synthetic results from recalibration (#14)
streamentry Jul 26, 2026
259a7ed
test: keep live collection count honest (#15)
streamentry Jul 26, 2026
0c1deb3
test: align calibration fixtures with synthetic gate (#16)
streamentry Jul 27, 2026
5a3b850
test: align dry-run smoke test with public APIs
streamentry Jul 28, 2026
4a8dbf3
fix stale doc report summary (#18)
streamentry Jul 29, 2026
9bb36a0
Harden semantic documentation link checks
streamentry Aug 1, 2026
eaa2c57
docs: keep benchmark agent metrics path current (#20)
streamentry Aug 3, 2026
12d96b0
fix: route review packet generation through canonical ERP
streamentry Aug 5, 2026
d636a26
fix: fail closed on invalid review packets
streamentry Aug 7, 2026
61fd111
feat: add canonical V4 review packet schema
streamentry Aug 10, 2026
bd07b2f
fix: enforce V4 review packet consistency in schema
streamentry Aug 11, 2026
2ee6311
fix: harden V4 packet timestamp validation
streamentry Aug 13, 2026
bd3d252
fix: enforce V4 packet status consistency
streamentry Aug 18, 2026
65dfc4d
fix: surface lab result data origin (#27)
streamentry Aug 21, 2026
8e4cc5b
feat: schema-validate lab result reports
streamentry Aug 22, 2026
dbc9c14
fix: require locked pilot preregistration
streamentry Aug 22, 2026
8a6df73
fix: bind locked pilot preregistration content
streamentry Aug 23, 2026
961dabc
feat: add pilot preregistration lock helper
streamentry Aug 31, 2026
b1ad98a
feat: expose pilot preregistration CLI check
streamentry Sep 1, 2026
cb12f03
fix: harden pilot preregistration CLI input boundary
streamentry Sep 4, 2026
d897cd0
fix: reject malformed nested pilot preregistration input
streamentry Sep 5, 2026
a9b8378
docs: maintain AI engineering guidance with monthly evidence reviews
cschanhniem Oct 4, 2026
ac7830f
Merge pull request #35 from streamentry/codex/ai-engineering-practice…
streamentry Oct 4, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
71 changes: 71 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -53,8 +53,30 @@ If these docs conflict, safety and claim discipline win.
`phase-aa-reproducibility-gate-check` and
`make phase-aa-reproducibility-gate-check`; partial or not-established
verdicts exit nonzero and do not certify a pipeline run.
- Phase AB exposes the ABAG- claim-integrity aggregate through
`phase-ab-claim-integrity-gate-check` and
`make phase-ab-claim-integrity-gate-check`; partial or not-established
verdicts exit nonzero and do not authenticate reviewers, validate science,
or establish biological evidence.
- Phase Z exposes the ZAG- per-family accountability aggregate through
`phase-z-accountability-gate-check` and
`make phase-z-accountability-gate-check`; partial or not-established
verdicts exit nonzero and do not establish benchmark superiority or
adapter ranking authority.
- Phase Y exposes the YAG- baseline-vs-pipeline accountability aggregate
through `phase-y-accountability-gate-check` and
`make phase-y-accountability-gate-check`; only a complete CBR/FIA/SDA/PMC
artifact set returns success. This is a dry-lab comparison-review control,
not evidence that the pipeline beats cheap baselines or validates biology.
- Use `python3 -m pytest --collect-only -q --no-header` to verify the full test
graph before relying on targeted evidence.
- The current pytest collection count recorded in
`docs/evidence/METRICS_CURRENT.md` is checked by
`tests/test_current_state_alignment.py`; intentional test additions or
removals must update that source-of-truth note.
- The toy end-to-end smoke path in `tests/test_pipeline_dry_run_e2e.py` must
consume the current public artifact builders and dataclasses. It is an API
compatibility check only and does not create biological evidence.
- The Phase E ERP example and validator retain an explicitly legacy compatibility
bridge; new packet work must use the component-based V4 ERP API.
- Lab-result directory loading remains warning-compatible for legacy callers, but
Expand All @@ -65,12 +87,54 @@ If these docs conflict, safety and claim discipline win.
- Calibration and reporting workflows also retain duplicate result IDs and
duplicate panel candidate IDs as structured input-integrity issues; those
inputs are not clean evidence and block the recalibration gate.
- Calibration intake also retains orphan result candidate IDs when a result
references a candidate absent from the submitted panel; orphan results are
not joined to predictions and block clean intake/recalibration.
- Candidate outcome rollups retain raw failed-control observations and IDs for
audit, but interpretable outcome flags and numeric counts use only
control-passing observations; failed controls still block recalibration.
- Lab-result batch summaries retain raw qualitative counts for audit, but expose
a separate `by_usable_qualitative_result` view restricted to control-passing
observations; reports label both views explicitly.
- Calibration intake retains control-failed assay observations for audit, but
excludes them from per-assay actual predicates and cohort metrics; failed
controls remain a recalibration-gate blocker.
- Calibration intake can verify each result's required computational-certificate
hash against an optional panel column; mismatches or partial opted-in coverage
are structured input-integrity blockers. Legacy panels without that column are
reported as certificate identity not available, not silently verified.
- Calibration intake can also verify an optional frozen `panel_id` against each
matched result. Multiple panel IDs, mismatches, or partial opted-in coverage
are structured input-integrity blockers; legacy panels report panel identity
not available, not silently verified.
- Lab-result reports expose raw assay-file hash coverage as
`no_results`, `not_available`, `partial_declaration`, or
`declared_for_all`. A declared `raw_data_sha256` is provenance only, not an
independently verified file hash; this status does not change legacy intake
acceptance or recalibration policy.
- Supplying `--raw-data-dir` to `lab-result-report` or `calibration-intake`
enables independent SHA-256 verification for records that provide the
relative `raw_data_file` field. Missing files, path escape, or mismatches are
structured verification blockers; matching bytes do not validate assay
contents, reviewer identity, biology, or release readiness.
- Lab-result `assay_date` values are checked as real, canonical `YYYY-MM-DD`
calendar dates after schema validation. Impossible or non-canonical dates are
retained as structured invalid-file errors and cannot enter reports or
metrics; this is temporal input integrity, not assay validation.
- Intake reports classify explicit `SYNTHETIC` labels by result ID. Those
records remain usable for demonstrations and audit, but the recalibration
gate fails closed when any synthetic-labeled result is present. Unclassified
records are not silently asserted to be real wet-lab evidence.
- The Phase R scientific-review readiness gate is available through
`scientific-review-readiness-check` and
`make scientific-review-readiness-check`; only a
`ready_for_external_review` verdict exits successfully. This is a dry-lab
documentation gate, not biological validation or release authorization.
- Domain-review outcomes remain backward-compatible when validated by ID alone,
but `domain-review-outcome-check --package-json <frozen-pep.json>` now fails
closed unless the outcome carries a matching `pep_sha256`. This binds a
review record to the exact frozen package JSON; it does not authenticate the
reviewer or establish scientific correctness.

## The agent role

Expand Down Expand Up @@ -132,6 +196,7 @@ A rule that nothing can catch you breaking is a wish, not a contract. Wherever t
| Ranking gates are respected | `make gate-check` · `make bench-gate` |
| Outputs are reproducible | `make cert-quality-check` · `make full-reproducibility-report` |
| Docs link where they claim | `make doc-links-check` |
| Bare documentation paths remain current | `src/openamp_foundry/checks/stale_doc_detector.py` and its focused tests |
| Deprecated benchmarks stay dead | `make bench-deprecation-check` |
| Code is green and typed | `make ci` (lint + test) · `make coverage` · `make typecheck` |
| Fast pre-PR bundle | `make agent-check` then `make doctor` |
Expand Down Expand Up @@ -417,3 +482,9 @@ This is what makes the contract future-proof: every rule here is written to get
## Final sentence

Build trust, not theater.

## Maintaining AI engineering guidance

At the first repository task each month (Asia/Ho_Chi_Minh), follow [the monthly practice review](docs/operations/HUMAN_AGENT_COLLABORATION.md#monthly-ai-engineering-practice-review), starting with claude.dev. Apply evidence-backed improvements to this contract and canonical docs; preserve existing ownership, security, product, and release rules. This runs on agent entry, not a background scheduler.

Before long-task interruption/compaction, record a redacted checkpoint and revalidate actual state on resume using [the resume protocol](docs/operations/HUMAN_AGENT_COLLABORATION.md#resuming-agent-work). Claims of better prompt/skill/workflow outcomes require [independent evaluation](docs/operations/HUMAN_AGENT_COLLABORATION.md#evaluating-guidance-changes); source recommendations and green counts alone are not proof.
28 changes: 22 additions & 6 deletions Makefile
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

PYTHON := $(shell [ -f .venv/bin/python ] && echo .venv/bin/python || echo python3)

.PHONY: phase-aa-reproducibility-gate-check
.PHONY: phase-aa-reproducibility-gate-check phase-ab-claim-integrity-gate-check phase-z-accountability-gate-check scientific-review-readiness-check
PYTEST := $(shell [ -f .venv/bin/pytest ] && echo .venv/bin/pytest || echo pytest)
RUFF := $(shell [ -f .venv/bin/ruff ] && echo .venv/bin/ruff || echo ruff)

Expand Down Expand Up @@ -81,7 +81,7 @@ help:
@echo " make generate-synthetic-lab-results Generate synthetic lab results for calibration testing"
@echo " make calibration-audit-example Run calibration pipeline consistency audit on synthetic example"
@echo " make calibration-audit Run calibration pipeline consistency audit (INTAKE=[path] GATE=[path] ...)"
@echo " make test Run full test suite (2937 passing tests, >=80% coverage)"
@echo " make test Run the full test suite (coverage target: >=80%)"
@echo " make coverage Test suite with per-module coverage report"
@echo " make lint Ruff lint check on src/ tests/ scripts/"
@echo " make typecheck mypy type check on src/"
Expand Down Expand Up @@ -645,11 +645,11 @@ lab-batch-pack:

generate-review-packet:
PYTHONPATH=src $(PYTHON) scripts/generate_review_packet.py \
--format v4 \
--erp-id ERP-DEMO-$(shell date -u +%Y%m%d) \
--batch-id BATCH-DEMO-$(shell date -u +%Y%m%d) \
--pipeline-version v0.5.73 \
--git-sha $$(git rev-parse HEAD) \
--candidate-count 36 \
--proof-ladder-level 2 \
--out outputs/review_packet_skeleton.json \
--out outputs/review_packet_v4.json \
--validate

failed-candidate-report:
Expand Down Expand Up @@ -963,6 +963,22 @@ phase-aa-reproducibility-gate-check:
PYTHONPATH=src $(PYTHON) -m openamp_foundry.cli phase-aa-reproducibility-gate-check --entry-json '{"aarg_id":"AARG-001","pipeline_version":"demo","rmc_id":"RMC-001","dcr_id":"DCR-001","cfp_id":"CFP-001","sbw_id":"SBW-001","created_at":"2026-07-16"}' --format text
@echo "Phase AA reproducibility gate check complete."

phase-ab-claim-integrity-gate-check:
PYTHONPATH=src $(PYTHON) -m openamp_foundry.cli phase-ab-claim-integrity-gate-check --entry-json '{"abag_id":"ABAG-001","pipeline_version":"demo","components_present":["CSD","RDR","EGN","EHP"],"limitations":["Dry-lab claim-integrity review control; not scientific validation."],"created_at":"2026-07-26"}' --format text
@echo "Phase AB claim-integrity gate check complete."

phase-y-accountability-gate-check:
PYTHONPATH=src $(PYTHON) -m openamp_foundry.cli phase-y-accountability-gate-check --entry-json '{"yag_id":"YAG-001","pipeline_version":"demo","cbr_artifact_id":"CBR-001","fia_artifact_id":"FIA-001","sda_artifact_id":"SDA-001","pmc_artifact_id":"PMC-001","limitations":["Dry-lab baseline accountability only; not biological validation."],"created_at":"2026-07-25"}' --format text
@echo "Phase Y accountability gate check complete."

phase-z-accountability-gate-check:
PYTHONPATH=src $(PYTHON) -m openamp_foundry.cli phase-z-accountability-gate-check --entry-json '{"zag_id":"ZAG-001","pipeline_version":"demo","fbh_id":"FBH-001","bxr_id":"BXR-001","arg_id":"ARG-001","cbf_id":"CBF-001","created_at":"2026-07-23"}' --format text
@echo "Phase Z accountability gate check complete."

scientific-review-readiness-check:
@set +e; PYTHONPATH=src $(PYTHON) -m openamp_foundry.cli scientific-review-readiness-check --entry-json '{"srg_id":"SRG-DEMO-001","candidate_family_id":"FAMILY-DEMO-001","cfc_id":"CFC-DEMO-001","fnr_id":"FNR-DEMO-001","atr_id":"ATR-DEMO-001","pqg_id":"PQG-DEMO-001","readiness_verdict":"not_ready","safety_flags":["no_flags"],"failed_gates":["No qualified wet-lab result is available"],"review_scope":"internal_only","n_confirmed_hits":0,"n_total_candidates":1,"limitations":"Dry-lab readiness example; not biological proof."}' --format text; status=$$?; test $$status -eq 3
@echo "Scientific review readiness is blocked as expected until qualified evidence exists."

pre-registration-check:
openamp-foundry pre-registration-check --entry-json '{"registration_id":"PRE-001","batch_id":"BATCH-001","pipeline_version":"0.9.4","registration_date":"2026-07-10","primary_hypothesis":"Candidates selected by OpenAMP will show MIC values at least 2-fold lower than random length/charge-matched peptides in broth microdilution against E. coli ATCC 25922.","primary_outcome_metric":"mic_value","success_threshold":4.0,"baseline_comparators":["random_selection","charge_matched_random"],"candidate_ids":["AMP-001","AMP-002","AMP-003"],"assay_type":"mic_assay","statistical_test":"Mann-Whitney U test, two-sided, alpha=0.05","registered_by":"test@example.com","dry_lab_only":true}' --format text

Expand Down
87 changes: 79 additions & 8 deletions SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,11 @@ qualified lab evidence.
honesty checks and regression entrypoints.
- `src/openamp_foundry/calibration/`: lab-result intake, gate, and proposal-only
recalibration.
- `src/openamp_foundry/checks/`: deterministic repository-integrity checks,
including stale documentation reference detection.
- `tests/test_pipeline_dry_run_e2e.py`: toy-only end-to-end smoke test that
exercises the current public artifact APIs without external calls or lab
claims.

## Diagrams

Expand Down Expand Up @@ -71,9 +76,11 @@ sequenceDiagram
1. Read `AGENTS.md`, `CLAUDE.md`, `MISSION.md`, and `docs/evidence/METRICS_CURRENT.md`.
2. Treat `docs/evidence/METRICS_CURRENT.md` plus `outputs/metrics_snapshot.json` as the
current benchmark truth when docs disagree.
3. For the required disconfirming pass, use
3. Keep the live test-graph count in `METRICS_CURRENT.md` synchronized; the
current-state alignment test fails when the recorded count drifts.
4. For the required disconfirming pass, use
`docs/evidence/DISCONFIRMING_TEST_RECORD_GUIDE.md` when recording a challenge.
4. Preserve the safety boundary: dry-lab scoring and evidence only. No wet-lab
5. Preserve the safety boundary: dry-lab scoring and evidence only. No wet-lab
protocols, pathogen enablement, toxicity-maximizing objectives, or biological
proof claims.

Expand All @@ -100,6 +107,43 @@ RMC, DCR, CFP, and SBW artifact IDs are all present. This is a structural
provenance check, not proof that the underlying run is scientifically correct
or biologically valid.

The Phase AB claim-integrity gate is available through
`openamp-foundry phase-ab-claim-integrity-gate-check --entry-json ...` or
`make phase-ab-claim-integrity-gate-check`. It returns success only when CSD,
RDR, EGN, and EHP components are all present. This checks claim-review and
external-handoff assembly; it does not authenticate reviewers, validate the
science, establish biology, or authorize release.

The Phase Z per-family accountability gate is available through
`openamp-foundry phase-z-accountability-gate-check --entry-json ...` or
`make phase-z-accountability-gate-check`. It returns success only when FBH,
BXR, ARG, and CBF artifact IDs are all present. This checks that the
per-family benchmark and adapter-accountability surface was assembled; it does
not prove benchmark superiority, biological validity, or release readiness.

The Phase Y baseline-vs-pipeline accountability gate is available through
`openamp-foundry phase-y-accountability-gate-check --entry-json ...` or
`make phase-y-accountability-gate-check`. It returns success only when CBR,
FIA, SDA, and PMC artifact IDs are all present. This checks that the
cheap-baseline comparison surface was assembled; it does not prove that the
pipeline beats those baselines, validate biology, or authorize an external
pilot claim.

The Phase R scientific-review readiness gate is available as
`openamp-foundry scientific-review-readiness-check --entry-json ...` or
`make scientific-review-readiness-check`. It returns success only for
`ready_for_external_review`; conditional, incomplete, safety-blocked, and
malformed inputs fail closed. The Make example is intentionally blocked until
qualified evidence exists. This is a dry-lab documentation control, not
biological validation or release authorization.

When the frozen pilot-evidence package JSON is available, validate a domain
review outcome with `domain-review-outcome-check --entry-json ...
--package-json <pep.json>`. The package-aware path requires a matching
`pep_sha256`; ID-only validation remains available for legacy records. A
verified hash proves package identity only, not reviewer authentication,
scientific correctness, or biological validity.

External-result intake is also fail-closed at the review boundary. Use the
structured loader/report fields `invalid_lab_result_files` and
`input_validation_status` to preserve schema-invalid returns; the
Expand All @@ -108,12 +152,39 @@ to proceed while any invalid file is excluded. Missing or non-directory result
paths return an input error before a report is written. An existing empty
directory is the only valid zero-result state. Duplicate result IDs and
duplicate panel candidate IDs are also preserved as `input_integrity_issues` and
block clean intake. These controls catch incomplete or ambiguous input, not
assay-quality or biological-validity problems. Control-failed assay observations
remain visible for audit but are excluded from per-assay actual predicates,
cohort metrics, and interpretable per-candidate outcome flags. Raw outcome fields
and failed-result IDs remain available for audit; failed controls still block
recalibration.
block clean intake. Result candidate IDs absent from the submitted panel are
preserved as orphan-result integrity issues and also block clean intake, because
they cannot be joined to prior predictions. These controls catch incomplete or
ambiguous input, not assay-quality or biological-validity problems. Control-failed
assay observations remain visible for audit but are excluded from per-assay
actual predicates, cohort metrics, and interpretable per-candidate outcome
flags. Raw outcome fields, failed-result IDs, and raw batch-level qualitative
counts remain available for audit; usable batch counts are restricted to assays
with both controls passing. Failed controls still block recalibration. New
panels may also carry
`computational_candidate_certificate_hash`; when present, result hashes must
match for every tested candidate. Mismatches or partial opted-in coverage block
clean intake. Legacy panels without the optional column are reported as
certificate identity not available, not silently verified. New panels may also
carry an optional frozen `panel_id` in both panel and result records. Multiple
panel IDs, mismatches, or partial opted-in coverage block clean intake; legacy
panels report panel identity not available, not silently verified.
Lab-result reports also expose raw assay-file hash coverage as `no_results`,
`not_available`, `partial_declaration`, or `declared_for_all`. A declared
`raw_data_sha256` is provenance only, not an independently verified file hash;
this status does not change legacy intake acceptance or recalibration policy.
When a caller supplies `--raw-data-dir`, records with `raw_data_file` can be
independently checked against their declared SHA-256. Missing files, path
escape, and mismatches block clean calibration intake; matching bytes prove
file identity only, not assay validity or biology.
Lab-result `assay_date` values are additionally checked as real, canonical
`YYYY-MM-DD` calendar dates because JSON Schema's date format annotation is not
enforced by the generic validator. Impossible or non-canonical dates are
retained as structured invalid-file errors and cannot enter reports or metrics.
Intake reports also classify explicit `SYNTHETIC` labels by result ID. Synthetic
records remain available for demonstrations and audit, but the recalibration
gate rejects any report containing them; unlabeled records remain unclassified,
not asserted to be real wet-lab evidence.

- `make bench-easy-baseline`: trivial length/charge baselines.
- `make bench-charge-matched`: adversarial check that removes charge-density
Expand Down
4 changes: 4 additions & 0 deletions docs/AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,10 @@ documents remain at stable paths to preserve public links.
- `getting-started/`, `operations/`, `review/`: contributor workflows.
- `research/`: current strategy separated from historical records.
- `PROJECT_INDEX.md`: complete inventory.
- Current status pages should expose executable review gates and their expected
fail-closed behavior, not imply that a gate is biological validation.
- The current status route includes the Phase Z ZAG- per-family accountability
workflow; its presence check must remain distinct from benchmark superiority.

## Diagrams

Expand Down
Loading
Loading