An individual-based simulation framework for studying antimicrobial use, antimicrobial resistance (AMR), infection outcomes, and potential policy interventions across a broad bacterial and antibacterial ecosystem.
The current executable configuration represents:
- 42 bacteria
- 62 individual antibacterial drugs
- 39 internal drug classes
- 46 resistance mechanisms
- 6 world regions
- daily simulation from 1930 to 2025 for calibration, or to 2035 for policy branches
The framework is under active development and calibration. It is intended for research and policy analysis, not for diagnosis, prescribing, or individual clinical decision-making. A formal open-source licence has not yet been selected, so this repository is not yet a released software version.
The model follows simulated people through demographic change, bacterial infection and carriage, symptoms and sepsis, diagnostic testing, antibacterial treatment, resistance dynamics, and mortality.
Resistance can be introduced or altered through:
- resistance profiles sampled at infection or carriage acquisition
- de novo mechanism emergence under relevant drug pressure
- horizontal gene transfer for eligible mobile mechanisms
- mechanism reversion and fitness costs
- treatment selection and differential bacterial clearance
- regional and care-setting resistance history
- a bounded local persistence archive representing resistant strains circulating outside the finite simulated sample
Infection incidence is externally parameterised. The model does not dynamically generate infection incidence from the prevalence of infected people. Its dynamic population interaction concerns resistance profiles and treatment pressure rather than explicit person-to-person transmission.
The framework deliberately does not model organism-drug MIC values or detailed
drug-specific PK/PD. Resistance severity is represented by bounded model
quantities such as any_r, while Figure 2 calibration prevalence is based on
whether any_r > 0.
See the technical model description for the complete scientific and implementation specification.
The executable inventories are defined in:
BACTERIA_LISTinsrc/simulation/population.rsDRUG_SHORT_NAMESinsrc/simulation/population.rsResistanceMechanism::all()insrc/simulation/population.rs
Current parameter values are defined in src/config.rs. Files under archive/
are historical evidence only and are not model or analysis inputs. Repository
tests protect the dimensions and relationships between the executable
inventories, parameters, targets, and output schema.
- Rust stable with Cargo; the crate uses Rust edition 2021
- Python 3.10 or later for analysis
- Substantial RAM and runtime for research-scale simulations
The Python environment is currently specified by lower bounds in
requirements.txt; a frozen release environment is still to be added.
cargo build --release
cargo test --all-targetsThe default binary is executable_amr.
cargo run --releaseRun settings are currently selected near the top of main() in
src/main.rs. The checked-in configuration uses a population of 10,000,000,
CalibrationMode::Full, random seeding, and no individual or infection-journey
logging. This is a long research run, not a quick installation test.
For a smoke test only, temporarily use a much smaller population_size.
Outputs from a small population must not be interpreted as calibrated model
results.
| Mode | Years and output |
|---|---|
FullMinimal |
Sparse 2022-2025 output containing drug share and bacteria-drug resistance fields |
Full |
Sparse 2022-2025 output containing all fields required by calibration_summary.py |
Partial |
Daily 1930-2025 output for historical time-series analysis |
Partial25Counterfactual |
Daily policy-0 rows for 1930-2025 plus no-resistance policy-2 rows for 2022-2025 |
Full25Counterfactual |
Sparse 2022-2025 output for policy 0 and the no-resistance policy 2 |
None |
Full 1930-2035 run with selected policy branches from 2027 |
time_steps is selected from the mode: 35,040 days for calibration and 2025
counterfactual modes, and 38,325 days for the full policy horizon. The two 2025
counterfactual modes checkpoint immediately before the first 2022 timestep,
retain the completed policy-0 trajectory, then restore that checkpoint and run
the resistance-suppressed policy 2 through the end of 2025.
Branch-enabled runs use disk-backed checkpoints by default. Checkpoint capture streams borrowed population and mechanism-cache state directly to a temporary file, then flushes, checksums, and atomically publishes it without cloning the population in memory. After the baseline summaries have been retained, the completed active population is released before a checkpoint is deserialized; the restored state is moved into the policy branch. Multi-policy runs reread the checkpoint for each branch to keep memory bounded. These modes therefore still require space for the serialized checkpoint and retained summary rows, but disk checkpointing does not require another simultaneous full population. The explicitly selectable in-memory checkpoint path remains suitable only when the additional population copy is known to fit. Before rerunning at 10 million, use fixed-seed 300k, 1M, and 3M trials to review the emitted checkpoint size, write/read duration, phase RSS/PSS, and runner cgroup peak. Provision checkpoint disk capacity with explicit headroom; an interrupted process can leave a uniquely named stale file in the configured checkpoint directory for deliberate cleanup.
To compare model-scope infection mortality between those two branches, set
SIMULATION_CSV near the top of
amr_simulation_output_analysis/counterfactual_2025_death_rates.py, then run:
python amr_simulation_output_analysis/counterfactual_2025_death_rates.pyThe script reports mean annual model-scope infection deaths, population-scaled
to the same world-population target used by calibration_summary.py, over
2022-2025 for policy 0 and policy 2. It also reports deaths per 100,000
person-years and the corresponding bacterium-associated values for every
organism under each policy. The terminal report is also written to
output_graphs/counterfactual_2025_death_rates_NNNNNN.txt, where NNNNNN is
the run ID in the simulation CSV filename. It accepts output from either
Full25Counterfactual or Partial25Counterfactual. A CSV path supplied on the
command line overrides SIMULATION_CSV; --output-dir overrides the report
directory.
The launcher generates and records a random u64 seed unless fixed seeding is
enabled. AMR_RNG_SEED overrides the source setting and is the preferred way to
replay a run.
PowerShell:
$env:AMR_RNG_SEED = "1234567890"
cargo run --releaseBash:
AMR_RNG_SEED=1234567890 cargo run --releaseFixed-seed runs use named ChaCha RNG streams and deterministic population
chunks. For the same source, configuration, and seed, summary output is
expected to be reproducible across repeated runs and different
RAYON_NUM_THREADS settings.
Rayon uses the available logical cores unless RAYON_NUM_THREADS is set.
PowerShell:
$env:RAYON_NUM_THREADS = "4"
cargo run --releaseBash:
RAYON_NUM_THREADS=4 cargo run --releaseSimulation outputs are written under
amr_simulation_output_analysis_outputs/. A completed run normally produces:
simulation_summary_NNNNNN.csvrun_metadata_<timestamp>_seed_<seed>.txtconfig_validation_<timestamp>.txt
The summary CSV uses output schema version 6. Its fields depend on the selected
run mode and can number in the tens of thousands. Optional diagnostic-cascade
columns are accompanied by diagnostic_cascade_collection_enabled, so an
uncollected metric is not mistaken for a genuine zero count. The metadata
records the source hash, seed and seed source, run ID, population, time steps,
mode, policies, thread count, duration, output path, CSV SHA-256 hash,
validation status, and completion or failure state.
Schema 4 introduced resistance snapshots by home region. With regional collection enabled
(including the default Full calibration mode), the calibration summary reports
resistance prevalence and conditional mean any_r for all six regions, with the
number of contributing bacterium-drug pairs. These fields count living active
infections at the end of the day. regional_resistance_collected distinguishes
observed zeros from disabled collection; FullMinimal leaves this group disabled.
Schema-1 to schema-3 CSVs cannot supply this table. Generate a new simulation CSV
to obtain missing regional resistance results.
Schema 5 aligns global and hospital/community resistance counts and sums with the same end-of-day living active infections used for their denominators. It also aligns the associated infection-mechanism, testing and legacy MIC-proxy stocks. Hospital/community counts retain the pre-rule care-setting classification for both numerator and denominator. Regional resistance retains its schema-4 timing and home-region definition. This reporting correction leaves model dynamics and resistance-profile feedback unchanged.
Calibration summaries still run on historical schemas 1-4 but warn that the global drug-specific resistance counts used a different observation time from their infection denominators. Existing CSVs cannot be repaired by clipping percentages; a new simulation run is required for corrected global resistance results. Schema-5 and later calibration validates matching counts and sums and rejects inconsistent snapshots rather than silently clipping them.
Schema 6 applies the headline model scope to the existing regional and age-specific sepsis/non-sepsis infection-death counters; the CSV layout is unchanged. Calibration-summary death counts and rates now exclude deaths whose eligible infection contributors are only H. pylori or MDR-TB. A death with an eligible contributor inside the headline scope is counted once, including mixed infections. These fields retain the existing regional collection marker. Global death-cause counters remain broad, and regional background/toxicity counters are unchanged. Schemas 1-5 recorded broader regional/age infection deaths, so the restricted calibration tables are reported unavailable for those files. A new simulation run with regional collection enabled is required; the missing split cannot be reconstructed by subtracting bacterium-associated deaths.
Schema-1 and schema-2 summaries remain readable only through the audited calibration-snapshot compatibility path:
python -m amr_simulation_output_analysis.calibration_summaryPaper outputs from those legacy snapshots require the explicit compatibility flag and omit Supplementary Figure S5:
python -m amr_simulation_output_analysis.make_paper_tables `
--legacy-without-sf5 `
output_graphs/calibration_summary_123456.txtComprehensive analysis and SF5 accept schemas 3, 4, 5 and 6. Schemas 1 and 2 retain the explicit compatibility restrictions above. Matching column names across versions do not imply matching observation times.
The source hash can be supplied by AMR_SOURCE_HASH or source_hash.txt.
Otherwise the launcher uses the current Git commit and marks a dirty worktree.
For formal analyses, retain the metadata file and exact source snapshot with
the CSV.
Parameter validation is strict by default. AMR_CONFIG_VALIDATION=warn permits
a diagnostic run to continue despite validation errors, but such a run should
not be used as a calibrated research result.
Create an environment and install the current analysis dependencies:
python -m venv .venv
python -m pip install -r requirements.txtActivate the environment using the command appropriate for the operating
system. Select the input CSV through DataConfig.simulation_file in
amr_simulation_output_analysis/config.py, then run:
python -m amr_simulation_output_analysis.amr_analysisThe analysis writes calibration summaries and configured plots under
output_graphs/. Plot selection, policies, output format and memory settings
are controlled by PlotConfig; the input CSV and Parquet-cache options are
controlled by DataConfig in amr_simulation_output_analysis/config.py.
The calibration snapshot includes overall infection acquisition rates by region
for its baseline calibration window and exports the table to
output_graphs/infection_incidence_by_region_<run_id>.csv. Rates sum acquisition
events across all bacteria and both care settings, divided by summed daily
population / 365. They count bacterial events, including repeat acquisitions,
rather than unique people. Regional rates are approximate: the existing CSV
assigns events to home region but population to effective location, including
travel. The CSV includes this interpretation and its source and window.
Regional resistance adds 31,501 CSV columns with the current inventories, so these runs produce larger files. Preprocessing converts only the columns needed to calculate derived plotting fields and passes the unchanged raw matrices through without a Pandas–Polars roundtrip. Only derived columns are converted back; other wide conversions use bounded column batches. All regional data remains available to the calibration summary. These Python optimisations work with existing CSVs and do not require another simulation run.
The calibration snapshot also reports infection-death counts by age group and region, with raw calibration-window totals and mean annual counts scaled using the headline world-population factor. These combine sepsis and non-sepsis infection deaths using the same organism scope as the headline: deaths with only H. pylori or MDR-TB contributors are excluded. The existing regional and age counters use this scope from schema 6, and both totals reconcile with the headline before display rounding. Regional counts use the person's effective location at the start of the death day. Missing, disabled or historical broader-scope mortality data is reported as unavailable for these restricted calibration tables; rerun the simulation to obtain the revised counts.
make_paper_tables adds Figure 14 with two panels: annual infection-death counts
by region and by age group, using those calibration-summary tables. Both panels
use the same complete set of schema-6 runs and show world-scaled annual counts
for 2022-2025 in millions. Multiple runs receive equal weight, with 95% confidence
intervals for the mean; one run is shown without an interval. The figure exports
HTML, PNG and SVG, and its HTML page includes a table of the counts. Historical,
missing or incomplete breakdowns are excluded with an explanation; if no run
qualifies, the figure displays an unavailable-data notice. Region here means
effective location on the death day, including travel.
The separate table "Deaths among people actively infected, by bacterium and
home region" uses the existing {bacterium}_deaths_infected_{region} fields.
It reports world-scaled mean annual counts for each bacterium and six home
regions. These include deaths from all causes, including background and
toxicity deaths. A person with several active infections can appear in several
rows, so no grand total across bacteria is shown. This descriptive table is
available from existing CSVs when the core per-bacterium fields were collected,
including FullMinimal; the regional-resistance collection flag does not
control these fields.
Run the Python regression tests from the repository root with:
python -m unittest discover -s tests -p "test_*.py"The current calibration combines sourced estimates, transformed comparisons, evidence-informed benchmarks, and transparent expert-informed placeholders. These categories must not be treated as interchangeable observations.
Key provenance documents are:
- Resistance target provenance
- Comparison overlay provenance
data/resistance_targets_v1.manifest.json
Best-guess placeholder overlays are disabled by default and are not calibration score inputs.
CalibrationMode::None can run five independent branches from 2027:
| ID | Branch |
|---|---|
| 0 | Baseline continuation |
| 1 | Antimicrobial stewardship example |
| 2 | Resistance-suppressed AMR counterfactual |
| 3 | Near-complete diagnostics bound |
| 4 | Equal global access example |
CalibrationMode::Partial25Counterfactual and
CalibrationMode::Full25Counterfactual instead compare policy 0 with policy 2
from the start of 2022 through the end of 2025. The former retains the complete
policy-0 history from 1930; the latter retains only the 2022-2025 calibration
window. In both modes, calibration summaries select policy 0 and do not mix the
counterfactual rows into baseline fit statistics.
These branches are research scenarios and should not be interpreted as validated policy forecasts without scenario-specific calibration, uncertainty analysis, and suitable comparison across well-fitting parameter sets.
src/
main.rs Run launcher and provenance
config.rs Model parameters
config_validation.rs Parameter validation
observability.rs Run/source observability
rules/mod.rs Daily individual-level rules
simulation/
population.rs State, inventories, and enums
simulation.rs Simulation loop, caches, branches, CSV export
journey_logger.rs Optional sampled infection journeys
amr_simulation_output_analysis/ Python analysis package
data/ Calibration targets and comparison data
model_description/ Technical model description
tests/ Rust integration and Python regression tests
paper_tables/ Generated manuscript tables and figures
archive/ Historical, non-executable material
Sampled infection-journey logging can be enabled in src/main.rs for
illustrative or diagnostic work. It records much denser individual traces and
can materially slow a run, so it is disabled for routine calibration.
Contribution guidance, citation metadata, a stable run configuration interface, a locked Python environment, and a formal release archive are still being prepared. Until those are available, please treat the repository as an active research workspace rather than a stable public API.
No software licence has yet been selected. A standard open-source licence and the appropriate UCL/contributor copyright notice must be added before formal public release and reuse.