Build the world’s most rigorous open, safety-first AI antimicrobial peptide foundry: a reproducible dry-lab system that generates, filters, documents, and prepares novel antimicrobial candidates for qualified experimental validation through evidence rather than hype.
The deeper legacy mission is to turn that foundry into an open wet-lab compression engine for antimicrobial discovery: a system that helps qualified scientists ask fewer, smarter experimental questions and learn from every answer.
OpenAMP Foundry exists to discover whether an AI-directed, open, auditable candidate-selection pipeline can identify novel antimicrobial peptide families that later survive real laboratory validation.
The desired long-term result is:
An open AI pipeline discovers a novel antimicrobial peptide family, independently validated in qualified lab testing, with the full computational evidence trail released for scientific review.
This project must never claim that a computational prediction is a drug, cure, therapy, or proven antimicrobial result. The lab is the judge.
The project should be understood in two layers:
- Immediate wedge Build a dry-lab system that can generate, score, falsify, rank, and document AMP candidates honestly enough to deserve lab budget.
- Breakthrough wedge Build a calibrated virtual assay stack that improves experiment selection by modeling bacterial selectivity, mammalian membrane risk, and stability better than cheap heuristics alone.
The near-term goal is not full biological simulation.
The near-term goal is experimental compression:
use computation to reduce the number of real experiments needed to find a real, safe-enough signal.
Antimicrobial resistance is one of the clearest global scientific threats. The world needs more credible ways to search for new antimicrobial candidates, but most AI-for-science projects fail by overclaiming, cherry-picking, or hiding negative results.
OpenAMP Foundry is designed around the opposite principle:
Every candidate must carry its evidence, limitations, failure modes, and reproducible selection history.
The project should create value even before any lab hit by producing cleaner benchmarks, safer filters, auditable candidate certificates, and reusable negative-result infrastructure.
OpenAMP Foundry is a dry-lab pipeline for:
- Loading and cleaning antimicrobial peptide reference data.
- Checking benchmark leakage and near-duplicate contamination.
- Generating or importing candidate peptide sequences.
- Extracting physicochemical features.
- Scoring predicted antimicrobial activity.
- Penalizing predicted hemolysis, cytotoxicity, and unsafe properties.
- Checking novelty against known AMP references.
- Ranking diverse candidate batches.
- Producing machine-readable evidence certificates.
- Preparing qualified, safety-reviewed candidate batches for external experimental validation.
Over time, it should also become a platform for:
- Benchmarking virtual membrane and selectivity proxies.
- Learning from wet-lab outcomes through active feedback loops.
- Measuring whether new modeling layers actually save experiments.
OpenAMP Foundry is not:
- a wet-lab protocol repository;
- a dangerous pathogen research guide;
- a toxin-design system;
- a clinical drug-discovery claim machine;
- a hype engine for “AI discovered a cure” narratives;
- a project that publishes unscreened high-risk candidate lists without review.
The project does not provide instructions for culturing, enhancing, or misusing biological agents.
A result is not meaningful because a model ranks it highly.
A candidate becomes scientifically interesting only when it passes increasingly hard gates:
valid sequence
→ reproducible features
→ predicted activity
→ predicted low toxicity / low hemolysis
→ novelty check
→ synthesis feasibility
→ diversity selection
→ evidence certificate
→ expert review
→ qualified lab assay
→ independent replication
Computational evidence can justify testing. It cannot prove biological efficacy.
The higher bar beyond dry-lab credibility is not prettier scoring.
It is this:
design
→ simulate
→ test
→ learn
→ improve experiment selection
If simulation layers are added, they must be judged by whether they improve real decision quality, not by whether they sound advanced.
A result is headline-grade only if all conditions below are met:
| Requirement | Minimum standard |
|---|---|
| Novelty | Candidate family is not a trivial near-duplicate of known AMPs |
| Activity | At least one candidate shows meaningful antimicrobial activity in qualified lab testing |
| Safety | Candidate shows low hemolysis / low mammalian cytotoxicity in initial screening |
| Relevance | Testing uses an appropriate, safety-reviewed bacterial panel through qualified partners |
| Reproducibility | The dry-lab pipeline can be rerun from versioned inputs |
| Evidence | Every selected candidate has a valid evidence certificate |
| Independence | Key result is reproduced by an external lab or CRO |
| Honesty | Negative results and failed candidates are documented where safe |
| Safety | Release is reviewed for dual-use and misuse risk |
| Review | Claims are reviewed by qualified domain experts |
Only then may the project claim:
An open AI discovery pipeline produced a newly validated antimicrobial peptide family.
Every important claim must point to code, data, tests, benchmark results, literature, evidence certificates, or expert review.
Unsupported claims must be removed or explicitly marked as speculation.
A boring result that reproduces is more valuable than an impressive result that cannot be checked.
Major outputs should include:
- input data hash;
- model and scorer versions;
- config file;
- command used;
- random seed;
- code commit;
- output hash;
- timestamp.
Computational scores are pre-lab triage, not biological proof.
Allowed terms before lab validation:
- computationally nominated candidate;
- predicted antimicrobial peptide;
- dry-lab candidate;
- selected by reproducible pipeline.
Forbidden terms before sufficient evidence:
- drug;
- cure;
- safe;
- effective in humans;
- clinically useful;
- proven therapy;
- breakthrough treatment.
Preserve failures, rejected candidates, benchmark weaknesses, model disagreements, and negative results where safe.
A project that only publishes successes is not trustworthy.
The default objective must combine:
high predicted antimicrobial activity
+ low predicted mammalian toxicity
+ low predicted hemolysis
+ novelty
+ synthesis feasibility
+ candidate diversity
The project must not optimize for mammalian toxicity, virulence, immune evasion, harmful delivery, pathogen enhancement, or misuse against humans, animals, or crops.
AI agents may propose, implement, score, rank, and report.
AI agents may not make final scientific, safety, legal, release, lab-testing, or public-claim decisions.
The first serious milestone is not a lab discovery. It is a trustworthy dry-lab system.
Phase 1 is successful when a clean checkout can:
- Run the demo pipeline.
- Rank candidate peptide sequences.
- Validate evidence certificates against JSON Schema.
- Generate a batch report.
- Run tests in CI.
- Penalize obvious unsafe or low-quality candidates.
- Produce deterministic outputs from fixed inputs.
Phase 2 is successful when retrospective benchmarks show that the pipeline:
- Recovers hidden known active AMPs better than random or naive baselines.
- Remains meaningful under cluster-split evaluation.
- Does not rely on near-duplicate leakage.
- Down-ranks negative or toxic examples.
- Produces stable rankings under repeated runs.
Phase 3 is successful when the project has a lab-ready candidate pack with:
- 50–100 selected candidates.
- Evidence certificate for each candidate.
- Novelty report.
- Toxicity / hemolysis risk report.
- Diversity clustering report.
- Synthesis feasibility report.
- Pre-registered selection rule.
- Pass/fail criteria.
- External expert review.
- Dual-use safety review.
Phase 4 is successful when the project can responsibly begin calibrating a virtual assay layer using qualified wet-lab outcomes.
That phase requires:
- A written scope for what the emulator does and does not model.
- A benchmark separating bacterial-selective references from clearly hemolytic references.
- Confidence and abstention rules for simulator outputs.
- A pre-registered active-learning policy for choosing informative experiments.
- Evidence that the added layer improves prioritization relative to simpler baselines.
Stop, downgrade, or redesign the project if:
| Failure | Required response |
|---|---|
| Pipeline cannot beat simple baselines | Do not proceed to lab preparation |
| Rankings depend on leakage | Fix benchmark before claiming progress |
| Top candidates are near-duplicates of known AMPs | Strengthen novelty filters |
| Toxicity risk is ignored | Block candidate release |
| Outputs are not reproducible | Block external claims |
| Agents propose harmful objectives | Remove, review, and harden policy |
| Human reviewers cannot understand the evidence | Improve reports before continuing |
| Lab tests show no signal after repeated cycles | Publish negative result and reassess |
Failure is allowed. Unverifiable success is not.
The project should work toward this headline and no weaker hype version:
Open AI pipeline discovers a new antimicrobial peptide family, independently validated in lab tests, with full reproducible evidence trail released for scientific review.
The deeper systems headline, if the work truly earns it, would be:
OpenAMP built an open wet-lab compression engine for AMP discovery, showing that small, honest AI-guided experiment loops can find validated candidates with far less wasted assay effort.
Everything in this repository should move the project closer to that standard.