Skip to content

Latest commit

 

History

11 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

SMAL

Short-term Memory Active Learning for molecular property prediction.

This repository contains the code, configurations, and analysis materials needed to reproduce the experiments and figures from:

Yan Xiang, Jamie Wang, Zachary Fralish, and Daniel Reker. Short-Term Memory Active Learning for Drug Development. Journal of Chemical Information and Modeling 66(16), 9816–9828 (2026). doi:10.1021/acs.jcim.6c00687

About

Traditional active learning only ever adds data to a training set. SMAL additionally lets a model remove training data during learning. We implement several forgetting strategies built on random forest out-of-bag estimates and benchmark them across six ADMET datasets, showing that SMAL produces more balanced training sets and can lower labeling costs through data recycling. On datasets with artificially introduced label errors, SMAL autonomously identifies and discards the corrupted points, and across 99 virtual-screening datasets it matches or exceeds classical active learning in 97.7% of comparisons.

Installation

SMAL is built on top of MolALKit. All experiments in this repository were run with MolALKit v1.2.0 — please install that version following the instructions in the MolALKit README to reproduce the published results.

Reproducing the experiments

The scripts/ folder contains one bash script per (dataset, model, learning_type) combination. See scripts/README.md for the file-naming convention and usage.

Processed results and figures

The results_processed/ folder contains processed figure inputs. The figures/ folder contains one notebook per manuscript figure (figure2 through figure6 and figureS1 through figureS14); running a notebook regenerates the corresponding SVG from the processed inputs.

Datasets

The datasets/ folder contains the CSV input datasets used in this study:

  • 6 ADMET benchmark datasets: ames, CYP2D6_Veith, CYP3A4_Veith, MDR1_MDCK_classification2, PAMPA_NCATS, and pgp_broccatelli.
  • 99 SIMPD ChEMBL Ki assay datasets, named CHEMBL*.csv.

These files are copied from the local MolALKit packaged datasets and the SIMPD source-data directory so the study inputs are available directly from this repository.

Colab demo

For a single-run, browser-only walkthrough, upload colab_demo.ipynb to Google Colab. The notebook installs MolALKit v1.2.0 and runs an SMAL example on the pgp_broccatelli dataset by default.

Citation

If you use this code or the SMAL method in your work, please cite:

@article{xiang2026smal,
  title   = {Short-Term Memory Active Learning for Drug Development},
  author  = {Xiang, Yan and Wang, Jamie and Fralish, Zachary and Reker, Daniel},
  journal = {Journal of Chemical Information and Modeling},
  volume  = {66},
  number  = {16},
  pages   = {9816--9828},
  year    = {2026},
  doi     = {10.1021/acs.jcim.6c00687}
}

About

Short-term Memory Active Learning

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages