Skip to content

merge - #1044

Open
streamentry wants to merge 36 commits into
Open-Problem-Lab:mainfrom
streamentry:main
Open

merge#1044
streamentry wants to merge 36 commits into
Open-Problem-Lab:mainfrom
streamentry:main

Conversation

@streamentry

Copy link
Copy Markdown
Contributor

Description

Change class

  • Class A — safe documentation or small bug fix
  • Class B — infrastructure / schema / CLI / report change
  • Class C — scientific behavior change: scoring, selection, benchmark, calibration, simulation
  • Class D — safety-sensitive or wet-lab-facing change

See docs/getting-started/MAINTAINER_GUIDE.md.

Type of change

  • Bug fix
  • New feature
  • New scorer / predictor / simulation module
  • New or changed benchmark
  • Documentation update
  • Safety policy change
  • Infrastructure / CI / tooling
  • Agent/onboarding improvement
  • External-review / collaboration artifact

Safety impact

  • No biological misuse capability added
  • No wet-lab protocol or operational biological procedure added
  • No harmful optimization objective added
  • No unscreened candidate sequences published
  • No model weights or sensitive artifacts released without policy review
  • No efficacy, safety, clinical, or therapeutic claims without evidence
  • Limitations and failure modes documented
  • Evidence is reproducible where applicable
  • Human safety review requested if this is safety-sensitive

Claim discipline

  • Claims are mapped to docs/evidence/PROOF_LADDER.md
  • Dry-lab outputs are not described as biological proof
  • No “AI discovered an antibiotic” / “drug candidate” / “safe” / “clinically useful” language unless evidence level supports it
  • Unsupported claims are explicitly listed or avoided

Baseline and benchmark discipline

Required for scorer, predictor, simulator, ranking, calibration, active-learning, or benchmark changes.

  • Cheapest meaningful baseline identified
  • Baseline comparison included or change is explicitly informational only
  • Benchmark leakage or shortcut risks considered
  • docs/evidence/METRICS_CURRENT.md updated if metrics changed
  • docs/evidence/BENCHMARKING.md / docs/evidence/BENCHMARK_GOVERNANCE.md updated if benchmark behavior changed

Verification

  • make ci passes
  • make bench-gate passes if benchmark-sensitive
  • Tests added for new functionality
  • Existing tests updated to match behavior changes
  • Documentation updated if behavior changed
  • Determinism confirmed (make regenerate-all if pipeline logic changed)
  • Schema validation passes if artifacts changed

Agent disclosure

  • This PR was not agent-generated
  • This PR was fully or partly agent-generated and human-reviewed
  • Agent-generated changes stayed within docs/getting-started/AGENT_ONBOARDING.md boundaries

Evidence

External or wet-lab-facing impact

  • Not applicable
  • Affects external review packet / handoff / pilot planning
  • Requires domain expert review
  • Requires safety review
  • Does not include operational experimental instructions

Dependencies and licenses

  • None

Reviewer focus

streamentry and others added 30 commits July 20, 2026 04:05
Merged after review: orphan lab-result joins are preserved as audit provenance and blocked from clean panel-specific calibration intake.
Expose the SRG- gate through a fail-closed CLI and Make review loop with tests and current-state documentation.
Adds optional certificate-hash identity checks to calibration intake. Mismatches and partial opted-in coverage fail closed; legacy panels remain explicitly unverified.
Keep failed-control observations auditable while excluding them from usable batch counts.
Add fail-closed optional panel identity checks to lab-result intake.
Merged after focused verification and CLEAR review comment. Full hook collection limitation was disclosed.
Expose the Phase Z ZAG- accountability gate through CLI and Make review workflows, with synchronized docs and fail-closed coverage.
Merged after focused review. Calendar validity remains metadata integrity only; it does not authenticate reviewers or establish scientific proof.
Bind domain review outcomes to the exact frozen package JSON when available.
Squash merge after focused verification and a non-approval automated review comment. Human review remains required for calibration-intake workflow impact; this change does not alter calibration policy or biological claims.
Reject impossible and non-canonical assay dates while preserving structured invalid-file provenance.
Expose the Phase Y YAG- gate through CLI and Make, with fail-closed tests and current-state documentation.
Expose the Phase AB ABAG- claim-integrity gate through CLI and Make with fail-closed integration coverage and synchronized handoff documentation.
* fix: block synthetic results from recalibration

* fix: fail closed on malformed synthetic counts

---------

Co-authored-by: OpenCode <opencode@example.com>
Synchronize the live pytest count with current metrics and harden the count regression parser.
Repairs synthetic calibration fixture/API drift while preserving fail-closed recalibration behavior.
Merged after self-review. Focused dry-run/API verification passed; repository-wide lint and the separate pre-push environment remain pre-existing limitations.
Separate clean summary metrics from stale-entry detail output.
Merged after scoped review. Local doc-link, targeted test, Ruff, and collection-alignment evidence passed; repository-wide pre-existing lint/test limitations remain documented in the PR.
Repoint active benchmark guidance to canonical metrics docs and guard the path with a regression test.
Use the V4 component-based review packet in the active Make workflow while preserving legacy compatibility. Missing components remain draft or incomplete.
Squash merge of the reviewed validation-boundary fix.
Merged after focused validation, PR-bundle verification, self-review, and documented baseline/environment caveats.
Portable V4 ERP schema now rejects contradictory component counts, missing lists, presence flags, and packet status. Focused schema tests and all 32 valid combinations pass; no biological or release claims changed.
Co-authored-by: OpenCode <opencode@example.com>
Keep Python and portable V4 packet validation fail-closed in sync.
Merged after self-review and focused verification. Full-suite limitations remain documented in the PR.
Add portable lab-result report schema and build-time validation while preserving audit duplicates and dry-lab limitations.
Fail closed on unlocked PRR records while preserving editable drafts; aligned tests and docs with the existing pre-experiment freeze contract.
Require and verify deterministic freeze_sha256 for locked PRR records while preserving editable draft construction.
streamentry and others added 6 commits September 1, 2026 06:55
Add a non-mutating helper to hash and lock PRR drafts, while rejecting silent re-freeze of already locked records.
Expose the canonical PRR lock and freeze-digest validation in the CLI with fail-closed exit codes and explicit distinction from the legacy pre-registration command.
Reject malformed PRR CLI field types with structured input errors while preserving valid and semantic-invalid exit behavior.
Reject malformed selection and amendment list entries as structured PRR CLI input errors.
…s-202610

docs: maintain AI engineering guidance with monthly evidence reviews
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants