scPertEval is a toolkit for experimenting with and sharing reference implementations of
evaluation protocols in single-cell perturbation studies, usable both as a command-line
interface and as a native Python API. The same catalog of protocols backs three actions:
score (score a model's predictions against ground truth), calibrate (calibrate a
protocol against empirical positive/negative controls per perturbation, reporting the Dynamic
Range Fraction (DRF) and Bound Discrimination Score (BDS)), and de (export per-gene
differential expression).
scPertEval is introduced in Towards Principled Evaluation of Single-Cell Perturbation Prediction Models, where we develop a taxonomy of evaluation protocols — decomposing them into representation, metric, score transformation, and reporting strategy — and use this package to assess protocol behavior across seven public perturbation datasets.
Schäfer, P. S. L., Reid, K. A., Boldyga, Z., Aksu, E. D., Hakem, H., & Saez-Rodriguez, J. (2026). Towards Principled Evaluation of Single-Cell Perturbation Prediction Models. bioRxiv. https://doi.org/10.64898/2026.07.23.740433
If you use scPertEval in your work, please cite that paper (see CITATION.cff).
→ Full documentation at https://scperteval.readthedocs.io/
pip install scpertevalOr from this repo:
pip install "scperteval @ git+https://github.com/Virtual-Cell-Research-Community/scPertEval.git"The Sinkhorn / optimal-transport metrics (the sinkhorn_w2_* protocols) need PyTorch and
GeomLoss, which are optional to keep the base
install light. Enable them with the sinkhorn extra:
pip install "scperteval[sinkhorn]"From the command line:
# calibrate protocols against built-in controls (DRF/BDS)
scperteval calibrate data/wessels23.h5ad -p all --de-method t-test
# score a model's predictions against ground truth
scperteval score data/wessels23.h5ad predictions.h5ad -p all
scperteval list protocols # also: de-methods | spaces | sources | calibratorsOr from Python — the same protocols, returning results in memory (see the Python API guide):
import scperteval as sp
prep = sp.prepare("data/wessels23.h5ad", "pearson_ctrl") # read + index once, reusable
result = sp.calibrate(prep, "pearson_ctrl", de_method="t-test")
result.aggregate # {"mean": …, "median": …} — calibrated DRF summary
result.per_perturbation # the per-perturbation detail tableSample datasets are available at
https://storage.googleapis.com/scperteval/processed/<dataset>_processed_complete.h5ad.
If you use scPertEval, please cite the paper it accompanies:
@article{Schafer_2026_scPertEval,
author = {Sch{\"a}fer, Philipp S. L. and Reid, Kendall A. and Boldyga, Zach and
Aksu, Ekin D. and Hakem, Hugo and Saez-Rodriguez, Julio},
title = {Towards Principled Evaluation of Single-Cell Perturbation Prediction Models},
journal = {bioRxiv},
year = {2026},
doi = {10.64898/2026.07.23.740433},
}scPertEval originates with the authors of that paper — Philipp S. L. Schäfer, Kendall A. Reid, Zach Boldyga, Ekin D. Aksu, Hugo Hakem, and Julio Saez-Rodriguez — and is maintained as a community project under the Virtual Cell Research Community.
Contributing: see CONTRIBUTORS.md.