Skip to content

About

No description, website, or topics provided.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

786 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

OpenAMP Foundry

OpenAMP Foundry is a verification-first, safety-constrained dry-lab foundry for AI-assisted antimicrobial peptide (AMP) discovery.

It is designed around a strict principle:

Computers can triage, falsify, rank, and document candidates. They do not prove biological efficacy. Wet-lab assays are still required before any scientific claim of activity.

The current repository is a rigorous dry-lab foundry.

The larger mission is more ambitious:

Build an open wet-lab compression engine for AMP discovery: a system that helps qualified scientists decide which small number of experiments are most worth running next, then learns from those outcomes.

The long-term infrastructure ambition is described in VISION.md, GOAL.md, docs/research/OPEN_BIOTECH_STACK.md, docs/research/OPEN_INFRASTRUCTURE_MOAT.md, and docs/trust/TRUST_CENTER.md.

Start here

You are Read first Then read
New human contributor docs/README.md docs/getting-started/FIRST_RUN_WALKTHROUGH.md, CONTRIBUTING.md
AI agent AGENTS.md docs/getting-started/AGENT_ONBOARDING.md, docs/operations/HUMAN_AGENT_COLLABORATION.md
Reviewer docs/getting-started/REVIEWER_ONBOARDING.md docs/evidence/CLAIM_REVIEW_CHECKLIST.md, docs/trust/RISK_REGISTER.md
Computational scientist docs/evidence/README.md docs/evidence/METRICS_CURRENT.md, docs/evidence/BENCHMARKING.md
Data/schema/model contributor docs/engineering/SCHEMA_REGISTRY.md docs/trust/DATA_GOVERNANCE.md, docs/engineering/ADAPTER_AUTHOR_GUIDE.md
Wet-lab/domain expert docs/review/WET_LAB_HANDOFF.md docs/review/EXTERNAL_REVIEW_PACKET.md, docs/review/PRE_REGISTERED_PILOT_TEMPLATE.md
Safety reviewer docs/trust/TRUST_CENTER.md SAFETY.md, SECURITY.md
Funder/institution docs/research/OPEN_INFRASTRUCTURE_MOAT.md docs/research/ADOPTION_METRICS.md, GOVERNANCE.md
Maintainer GOVERNANCE.md docs/operations/SUSTAINABILITY_AND_BUS_FACTOR.md, docs/trust/RELEASE_CHECKLIST.md

Why this repo exists

AI generation is cheap.

Trusted candidate selection is scarce.

OpenAMP Foundry exists to build the open evidence layer between AI-generated biological hypotheses and qualified experimental testing. The project is not trying to make biology look solved by software. It is trying to make experiment selection more reproducible, auditable, baseline-aware, and safe.

This repo gives you a safe starting point for:

  • building AMP candidate datasets;
  • scoring candidates with transparent baseline heuristics;
  • checking novelty against known references;
  • penalizing likely safety/synthesis risks;
  • selecting diverse candidates;
  • generating auditable JSON evidence certificates;
  • bundling expert-review packs with provenance and identity hashes;
  • maintaining run manifests, schema registry, and release-status records;
  • running a demo pipeline without downloading external biological datasets;
  • expanding later with real predictors and qualified external validation.

It also establishes the architecture and governance needed for a future virtual assay layer that can improve experiment selection without pretending to replace biology.

What this repo is

A starter implementation of a computer-first AMP candidate foundry:

candidate records
  -> validity checks
  -> physicochemical features
  -> activity-likeness score
  -> safety-risk score
  -> feasibility score
  -> novelty score against references
  -> ensemble rank
  -> evidence certificate
  -> run manifest
  -> expert-review package, if human review approves

The present repo answers:

Can we build a reproducible, leakage-aware, safety-first ranking pipeline that earns the right to guide real experiments?

The next-horizon repo should answer:

Can we compress wet-lab cost by learning which peptide experiments are worth running, better than cheap predictors alone?

What this repo is not

This repo is not:

  • a medical product;
  • a drug-discovery guarantee;
  • a wet-lab protocol collection;
  • an unsafe biological-design tool;
  • a generator for harmful biological capabilities;
  • a replacement for qualified microbiologists, toxicologists, or regulatory experts.

Safe scope

The default repo contains only toy/demo data and transparent baseline scorers. It deliberately avoids:

  • operational biological instructions;
  • unsafe optimization objectives;
  • release of unscreened high-risk candidate lists;
  • trained generator weights;
  • clinical or medical advice.

See SAFETY.md, RESPONSIBLE_USE.md, and MODEL_RELEASE_POLICY.md.

Vision ladder

OpenAMP now operates on two connected horizons:

  1. Current horizon — trustworthy dry lab Build deterministic ranking, evidence certificates, leakage-resistant benchmarks, novelty auditing, feasibility checks, and reviewable shortlist generation.
  2. Next horizon — wet-lab compression Add higher-fidelity membrane, selectivity, stability, and learned-surrogate layers that improve which small number of experiments a qualified lab should run next.

The second horizon only matters if the first one stays honest. Better simulation without calibration is not a breakthrough.

Quick start

Requires Python 3.11+.

python -m venv .venv
source .venv/bin/activate
pip install -e .[dev]
make demo

Validate one generated evidence certificate:

python -m openamp_foundry.cli validate \
  --certificate outputs/evidence/AMPF-000001.json \
  --schema schemas/candidate.schema.json

For first-run interpretation, read docs/getting-started/FIRST_RUN_WALKTHROUGH.md. For command interpretation, read docs/getting-started/COMMAND_SURFACE.md. Commands produce artifacts; artifacts support claims only when the proof ladder allows them.

Repository map

openamp-foundry/
  README.md                              # primary entrypoint
  VISION.md                              # long-term infrastructure vision
  GOAL.md                                # milestones, kill rules, metrics
  MISSION.md                             # project mission and claim boundaries
  GOVERNANCE.md                          # decision governance
  AGENTS.md                              # agent operating contract
  CLAUDE.md                              # concise collaborator guidance
  CONTRIBUTING.md                        # contributor workflow and PR checklist
  CODE_OF_CONDUCT.md                     # community and scientific integrity standard
  SAFETY.md                              # safety policy
  SECURITY.md                            # security and safety-sensitive reporting
  RESPONSIBLE_USE.md                     # allowed/disallowed use
  MODEL_RELEASE_POLICY.md                # model and artifact release policy
  DATA_LICENSE_NOTICE.md                 # data license and redistribution policy
  CITATION.cff                           # citation metadata
  .github/                               # PR template, issue templates, CODEOWNERS
  configs/                               # scoring and recalibration policy
  data/README.md                         # data directory rules
  models/README.md                       # model directory rules
  docs/README.md                         # task-based documentation front door
  docs/PROJECT_INDEX.md                  # exhaustive document catalog
  docs/trust/TRUST_CENTER.md                   # safety/evidence/governance trust front door
  docs/research/OPEN_INFRASTRUCTURE_MOAT.md       # durable infrastructure thesis
  docs/research/NUMBER_ONE_REPO_STANDARD.md       # category-leader standard
  docs/getting-started/FIRST_RUN_WALKTHROUGH.md          # first-run path
  docs/getting-started/COMMAND_SURFACE.md                # command workflows and claim boundaries
  docs/engineering/SCHEMA_REGISTRY.md                # schema and artifact registry
  docs/engineering/RUN_MANIFEST_STANDARD.md          # provenance standard
  docs/engineering/ADAPTER_AUTHOR_GUIDE.md           # safe adapter authoring
  docs/trust/RISK_REGISTER.md                  # major risks and mitigations
  docs/operations/SUSTAINABILITY_AND_BUS_FACTOR.md  # sustainability and bus-factor plan
  docs/trust/PUBLICATION_POLICY.md             # public claims policy
  docs/research/NEXT_100_PR_MAP.md                # PR-sized roadmap
  docs/engineering/CI_AND_QUALITY_GATES.md           # CI and quality gates
  docs/operations/HUMAN_AGENT_COLLABORATION.md      # human-agent collaboration model
  docs/getting-started/REVIEWER_ONBOARDING.md            # reviewer guide
  docs/research/ADOPTION_METRICS.md               # adoption metrics focused on trust
  docs/operations/DECISION_RECORD_TEMPLATE.md       # decision record template
  docs/getting-started/HUMAN_ONBOARDING.md               # human contributor onboarding
  docs/getting-started/AGENT_ONBOARDING.md               # agent task protocol
  docs/evidence/PROOF_LADDER.md                   # evidence levels and claim ladder
  docs/evidence/CLAIM_REVIEW_CHECKLIST.md         # claim review checklist
  docs/trust/DATA_GOVERNANCE.md                # data governance standard
  docs/trust/MODEL_CARD_TEMPLATE.md            # model/adapter card template
  docs/engineering/ARTIFACT_VERSIONING.md            # artifact compatibility policy
  docs/trust/RELEASE_CHECKLIST.md              # release checklist
  docs/evidence/BENCHMARKING.md                   # benchmark suite
  docs/evidence/BENCHMARK_GOVERNANCE.md           # benchmark lifecycle and governance
  docs/evidence/METRICS_CURRENT.md                # current benchmark summary
  docs/evidence/CALIBRATION_POLICY.md             # recalibration gate policy
  docs/evidence/EVIDENCE_CERTIFICATE.md           # candidate certificate spec
  docs/evidence/VIRTUAL_ASSAY_SCOPE.md            # virtual-assay scope and gates
  docs/review/WET_LAB_HANDOFF.md                # safe expert-review handoff guide
  examples/                              # toy datasets only
  outputs/.gitkeep                       # generated files ignored by git
  schemas/                               # JSON schemas
  scripts/                               # helper entrypoints and compatibility shims
  scripts/benchmarks/                    # canonical benchmark and baseline entrypoints
  scripts/calibration/                   # canonical calibration workflow entrypoints
  scripts/external/                      # canonical external predictor and handoff entrypoints
  scripts/lab/                           # canonical lab handoff entrypoints
  scripts/novelty/                       # canonical novelty DB and audit entrypoints
  scripts/release/                       # canonical demo, evidence, and reproducibility entrypoints
  scripts/research/                      # canonical exploratory generation and screening scripts
  scripts/waves/                         # canonical wave-program generation and panel scripts
  src/openamp_foundry/                   # Python package
  tests/                                 # repository tests
  tests/benchmarks/                      # benchmark and regression-gate tests
  tests/calibration/                     # calibration workflow tests
  tests/external/                        # external workflow and report tests
  tests/lab/                             # lab handoff and return-validation tests
  tests/novelty/                         # novelty scoring and novelty-pressure tests
  tests/release/                         # release artifact and reproducibility tests
  tests/waves/                           # wave-program gate tests

Philosophy

The project optimizes for honest candidate selection, not impressive claims.

A candidate is only worth lab money if it survives independent attacks:

  • basic validity;
  • novelty check;
  • feasibility review;
  • predicted activity;
  • predicted safety;
  • diversity selection;
  • reproducible evidence bundle;
  • human review.

The first serious milestone is not “AI discovered an antibiotic.”

The first serious milestone is:

A reproducible pipeline can recover known AMP positives, reject weak controls, avoid leakage, and generate a small shortlist of candidates that survives qualified external review.

The longer-range milestone is:

A calibrated virtual assay layer helps the project choose fewer, smarter experiments and improves hit-rate or safety-adjusted yield relative to cheap predictors alone.

License strategy

  • Core code: Apache-2.0.
  • Documentation: intended for CC BY 4.0 reuse where marked.
  • Third-party data: not bundled unless redistribution is allowed.
  • Generator weights and unscreened candidate lists: not released by default.
  • Project name and logo: trademark retained.

About

No description, website, or topics provided.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages