A community-driven, cross-species gene expression atlas for drug discovery.
Comparing gene/transcript expression patterns across human, mouse, rat, and cynomolgus monkey (cyno) is a routine but critical step in target validation, translational safety assessment, and PK/PD modeling during drug discovery. Despite how often this comparison is needed, there is currently no public, unified dataset that lets a scientist query a single gene or transcript across all four species and all major organs/tissues at once.
Efforts like GTEx provide deep human tissue expression data, but there is no equivalent harmonized resource spanning the other three species commonly used in preclinical research, nor a resource that puts all four side by side.
Build and maintain an open, community-contributed compendium of RNA-seq TPM (transcripts per million) expression values for:
- Species: human, mouse, rat, cynomolgus monkey
- Coverage: as many organs/tissues as possible per species (e.g., brain, heart, liver, kidney, lung, muscle, skin, blood, spleen, testis, ovary, etc.)
- Unit: gene- and transcript-level TPM, harmonized to a common gene/ortholog identifier scheme across species
The end result should let anyone answer: "How is gene X expressed across organs in human, mouse, rat, and cyno?" — in one query, from one place.
This project is a community collaboration and welcomes contributions from any scientist. Ways to help:
- Contribute data — Point us to or upload public RNA-seq TPM matrices (with proper source/license attribution) for any of the four species and any organ/tissue not yet covered.
- Contribute mapping/annotation — Help harmonize gene/transcript IDs and orthology mapping across species (e.g., Ensembl, NCBI, HomoloGene/Orthology resources).
- Contribute code — Scripts for data processing, QC, normalization, and the query interface are all welcome.
- Open an issue — Propose a dataset, report a data quality concern, or suggest a tissue/organ to prioritize.
Please open a pull request or issue to get started. All are welcome regardless of experience level.
Early stage — dataset structure and contribution guidelines are still being established. Watch this repo for updates.
TBD (data contributions must be from sources compatible with open redistribution; please note the original source/license when contributing).