Short-term Memory Active Learning for molecular property prediction.
This repository contains the code, configurations, and analysis materials needed to reproduce the experiments and figures from:
Yan Xiang, Jamie Wang, Zachary Fralish, and Daniel Reker. Short-Term Memory Active Learning for Drug Development. Journal of Chemical Information and Modeling 66(16), 9816–9828 (2026). doi:10.1021/acs.jcim.6c00687
Traditional active learning only ever adds data to a training set. SMAL additionally lets a model remove training data during learning. We implement several forgetting strategies built on random forest out-of-bag estimates and benchmark them across six ADMET datasets, showing that SMAL produces more balanced training sets and can lower labeling costs through data recycling. On datasets with artificially introduced label errors, SMAL autonomously identifies and discards the corrupted points, and across 99 virtual-screening datasets it matches or exceeds classical active learning in 97.7% of comparisons.
SMAL is built on top of MolALKit. All experiments in this repository were run with MolALKit v1.2.0 — please install that version following the instructions in the MolALKit README to reproduce the published results.
The scripts/ folder contains one bash script per (dataset, model, learning_type) combination. See scripts/README.md for the file-naming convention and usage.
The results_processed/ folder contains processed figure inputs. The figures/ folder contains one notebook per manuscript figure (figure2 through figure6 and figureS1 through figureS14); running a notebook regenerates the corresponding SVG from the processed inputs.
The datasets/ folder contains the CSV input datasets used in this study:
- 6 ADMET benchmark datasets:
ames,CYP2D6_Veith,CYP3A4_Veith,MDR1_MDCK_classification2,PAMPA_NCATS, andpgp_broccatelli. - 99 SIMPD ChEMBL Ki assay datasets, named
CHEMBL*.csv.
These files are copied from the local MolALKit packaged datasets and the SIMPD source-data directory so the study inputs are available directly from this repository.
For a single-run, browser-only walkthrough, upload colab_demo.ipynb to Google Colab. The notebook installs MolALKit v1.2.0 and runs an SMAL example on the pgp_broccatelli dataset by default.
If you use this code or the SMAL method in your work, please cite:
@article{xiang2026smal,
title = {Short-Term Memory Active Learning for Drug Development},
author = {Xiang, Yan and Wang, Jamie and Fralish, Zachary and Reker, Daniel},
journal = {Journal of Chemical Information and Modeling},
volume = {66},
number = {16},
pages = {9816--9828},
year = {2026},
doi = {10.1021/acs.jcim.6c00687}
}