A fast Rust reimplementation of MS-GF+ significance scoring — the generating-function spectral E-value (SpecEValue) for high-resolution tandem mass spectrometry — validated to be bit-exact against the reference Java MS-GF+.
Clean-room scorer (2026-09-30). The scorer and generating function were reimplemented clean-room from a written specification based on the published papers (
docs/cleanroom/). The implementer never saw MS-GF+'s source or the earlier Java-derived code, which was removed unread. The new code is byte-identical to the previous release on every test set. No MS-GF+ code remains in the scoring path. See LICENSING.md §3.
Given an MS/MS spectrum and a candidate peptide, MSGF_Rust reproduces MS-GF+'s three scores:
| Score | Meaning |
|---|---|
| RawScore | Integer match score: how strongly the observed peaks support the peptide's b/y fragment ions. |
| DeNovoScore | The maximum RawScore achievable by any peptide of the precursor mass (the best path through the de novo graph). |
| SpecEValue | The significance: P(score ≥ RawScore) over the ensemble of all peptides of that mass — the spectral E-value / p-value. |
On the validation set the Java implementation is the numeric oracle: RawScore and DeNovoScore
match exactly (30/30 PSMs), and SpecEValue agrees to f64 accumulation noise (|Δlog10| ≤ 0.05,
observed worst ≈3e-5). (Measured with the previous implementation; the clean-room one is
byte-identical to it on F13 with MS-GF+'s models, so the result carries over.) For speed, see
PERFORMANCE.md. Its numbers describe the previous implementation, and the
integration report gives the clean-room implementation's timings.
Database search is implemented (
msgf search): FASTA digestion, modification-aware candidate generation, RawScore + SpecEValue, and target-decoy q-values. Candidates are generated by a mass-sorted peptide index rather than the fragment index Sage uses — see plans/PLAN.md for the roadmap and plans/PLAN2.md for the target-decoy/FDR design.
msgf rescoreremains available for scoring a PSM list you already have.
Tagged releases attach prebuilt msgf binaries for Linux (glibc + static musl), macOS (Intel +
Apple Silicon), and Windows. Download the archive for your platform from the
Releases page, verify the checksum, and unpack:
tar -xzf msgf-v0.1.0-x86_64-unknown-linux-musl.tar.gz
./msgf-v0.1.0-x86_64-unknown-linux-musl/msgf --helpBinaries are built by .github/workflows/release.yml on every
v* tag (a manual workflow run also publishes them as downloadable CI artifacts).
Requires a recent stable Rust (developed on 1.94). The workspace lives under rust/:
cd rust
cargo build --release -p msgf-cli # binary at rust/target/release/msgf
./target/release/msgf --helpOther common commands:
cargo test --workspace # unit tests + golden validation (skips if data absent)
cargo clippy --workspace --all-targets
cargo bench -p msgf-genfunc --bench genfunc # SpecEValue DP, 1-core + rayonThe msgf binary has four subcommands: search, rescore, decoy, and fdr. Run
msgf <command> --help for the full flag list.
msgf search \
--spectra run.mgf \
--fasta human.revCat.fasta \
--fixed-mod C+57.021464 --var-mod M+15.994915 --num-mods 2 \
--precursor-tol 10ppm --ti 0,1 \
--out psms.tsvFASTA → enzymatic digestion → modification-aware candidates → RawScore → SpecEValue →
target-decoy q-values, parallel over spectra. Output uses MS-GF+'s TSV column set
(SpecEValue, EValue, QValue, PepQValue, …), so existing MS-GF+ tooling reads it.
Three defaults are worth knowing:
- The scoring model is bundled. With no
--param, MSGF_Rust scores with its own HCD/high-res/tryptic model, trained from the CC0 MassIVE-KB corpus — nothing to download, and nothing UC-licensed on the path. Pass--paramfor another activation/instrument/enzyme. This is not the model the bit-exactness claims are measured with — those use MS-GF+'sHCD_HighRes_Tryp.param. A different model is a different scoring function, so the default ranks candidates differently and is not a drop-in for reproducing MS-GF+ output; pass--paramif that is your goal. Different is not worse: on held-out ground truth the two models rank the true peptide above mass-identical decoys equally often (0.9988 vs 0.9988). Details and numbers. - Missed cleavages are unlimited, matching MS-GF+'s own default (
-maxMissedCleavageshas no limit there). This makes the index much larger; pass-c 2for a conventional search. - Q-values need decoys. Either search a concatenated database (
*.revCat.fasta— decoys are detected by accession prefix) or pass--tdato generate them. Without decoys theQValuecolumn is 0 and is not an FDR estimate; the CLI warns.
Amino-acid background frequencies for the generating function are computed from the database being searched, so SpecEValue reflects that database rather than a uniform alphabet.
Other useful flags: -e/--enzyme (number 0–10, name, or a full Name,CleaveAt,Terminus,Desc
definition), --ntt (2 fully enzymatic, 1 semi, 0 non-enzymatic), --min-len/--max-len,
-n/--num-matches, --mods (an MS-GF+ Mods.txt file), --unroll, --threads.
msgf decoy -d human.fasta -o human.revCat.fasta # byte-identical to MS-GF+ -tda 1
msgf fdr -i psms.tsv -o psms.q.tsv # append QValue / PepQValue to any PSM tablemsgf fdr works on any TSV with peptide, protein and score columns — including msgf rescore
output — so you can get MS-GF+-compatible q-values without running a search.
rescore takes a spectra file and a list of peptide-spectrum matches, and recomputes RawScore,
DeNovoScore and SpecEValue for each. The generating function depends only on the spectrum (not on
any one peptide), so it is built once per (scan, charge) and cached; every PSM against that
spectrum is then a cheap RawScore + tail lookup. Groups are independent, so they run in parallel
(--threads, default all cores); rows are still written in input order.
msgf rescore \
--spectra spectra.mgf \
--psms identifications.tsv \
--out rescored.tsv| Flag | Description | |
|---|---|---|
-s, --spectra <FILE> |
required | MS/MS spectra in MGF format. |
-p, --param <FILE> |
optional | Scoring model (.param). Defaults to the bundled MassIVE-KB-trained HCD/HighRes/Tryptic model; pass a file to match a different acquisition (activation × resolution × enzyme). |
-i, --psms <FILE> |
required | PSMs to rescore (TSV: scan, peptide, optional charge). |
-o, --out <FILE> |
optional | Output TSV path (default: stdout). |
--ti <LO,HI> |
optional | Isotope-error range, like MS-GF+ -ti (default 0,1). |
--aa-probs <FILE> |
optional | Amino-acid background probabilities (TSV). Default: uniform 0.05. |
--ox-m |
optional | Add variable oxidation on M (+15.994915) to the graph alphabet. |
--db-size <N> |
optional | Also emit evalue = SpecEValue × N (candidate count). |
--threads <N> |
optional | Worker threads (default: all cores). The (scan, charge) groups are scored in parallel; output and stderr are identical for every thread count. |
PSMs whose scan is missing from the spectra file, that have no charge, or whose peptide can't be
parsed are skipped with a note on stderr; a summary line (rescored N; skipped M) is printed at
the end.
Standard Mascot Generic Format. Each spectrum needs a SCANS= id (used to join with the PSM list),
CHARGE=, and PEPMASS= (precursor m/z), followed by m/z intensity peak lines:
BEGIN IONS
TITLE=scan=2018
PEPMASS=384.5719
CHARGE=3+
SCANS=2018
110.07130 2905.6
120.08084 1809.4
...
END IONS
The precursor's neutral mass is derived as PEPMASS × charge − charge × proton. (mzML support via
the mzdata crate is planned; today the reader is MGF only.)
A trained fragment-scoring model, one per (activation × resolution × enzyme × protocol).
You do not need one to run MSGF_Rust. The default is bundled: MSGFRust_HCD_HighRes_Tryp_v1,
counted by this project's own trainer (msgf-train) from 258k PSMs of the CC0 MassIVE-KB
peptide libraries. It is MIT-clean, ~1 MB, embedded in the binary, and announced on stderr at the
start of every run so results stay traceable. Provenance and how it compares to MS-GF+'s model:
rust/crates/msgf-scorer/models/README.md.
Which model you pass decides whether output is comparable to MS-GF+. The fidelity results quoted throughout this README — RawScore/DeNovoScore exact, SpecEValue within
|log10| ≤ 0.05— are measured with MS-GF+'sHCD_HighRes_Tryp.param. The bundled default is a different trained model, so it ranks candidates differently by construction and will not reproduce MS-GF+'s output.That is a difference, not an error. On 4,000 held-out MassIVE-KB spectra with ground-truth peptides, both models rank the true peptide above 5 mass-identical shuffled decoys 0.9988 of the time — identically — at 3919 vs 3939 IDs on a 1 %-decoy threshold (ρ = 0.973 on log₁₀ SpecEValue). Where the two disagree on a real search there is usually no right answer to be had: on F13, MS-GF+'s own top hits are 50.0 % decoy, i.e. exactly chance. Use
--paramto reproduce MS-GF+; use the default to run MSGF_Rust as its own engine.
Pass --param when your data is not high-resolution HCD tryptic. MS-GF+'s own models
(HCD_QExactive_Tryp.param, CID_HighRes_Tryp.param, ETD_HighRes_Tryp.param, …) are read
unchanged, but they are UC-licensed and not distributed here; fetch them with:
cd validation && ./fetch_reference_data.sh # models + test spectra + tiny FASTAsTraining a model for another identity is a counting pass over annotated spectra — see
docs/training.md.
Tab-separated, one PSM per line. Columns are located by a header row if present (a line containing
peptide, case-insensitive); otherwise the columns are assumed to be scan, peptide, charge
in that order.
scan peptide charge
2018 RSRRRRKR 3
6044 RTLMARPM+15.995IKEAR 2
6061 K.SIKNIQKITK.A
scan— matches an MGFSCANS=value.peptide— MS-GF+ peptide string:- bare sequence
PEPTIDEK; - optional enzyme context
K.PEPTIDEK.A(residue before./ after.;-marks a protein terminus). Context is used to score the N-terminal (neighboring) cleavage of semi-tryptic peptides; a bare peptide is assumed fully tryptic. - inline modification deltas
+d/-don the preceding residue, e.g.M+15.995,+42.011SAM…. - Only the 20 standard residues are accepted; an unknown residue skips the row.
- bare sequence
charge— optional; if omitted, the spectrum'sCHARGE=is used.
One residue per line, residue<TAB>probability; # comment lines allowed. These are the
background (edge) probabilities the generating function weights amino-acid transitions by. Default
is uniform 0.05 (MS-GF+ de novo). To reproduce a real search's SpecEValue you must supply
the searched database's composition (see below):
# residue probability (e.g. iPRG-2013 human FASTA composition)
A 0.069428
R 0.056718
...
scan peptide charge raw_score denovo_score spec_evalue
2018 RSRRRRKR 3 40 55 1.057136e-8
| Column | Description |
|---|---|
scan, peptide, charge |
Echoed from the input (charge resolved from spectrum if it was omitted). |
raw_score |
Integer RawScore = node+edge match score + terminal cleavage. |
denovo_score |
DeNovoScore: max achievable score for this precursor mass. |
spec_evalue |
SpecEValue: P(score ≥ raw_score). |
evalue |
Only with --db-size N: spec_evalue × N. |
Everything is MIT-licensed. The simplest route is the umbrella msgf crate — one dependency that
re-exports the whole pipeline:
[dependencies]
msgf = { git = "https://github.com/mwang87/MSGF_Rust" }
# scoring only, without the search engine and its rayon dependency:
msgf = { git = "https://github.com/mwang87/MSGF_Rust", default-features = false }use msgf::prelude::*;
let model = msgf::scorer::ScoreModel::from_file("HCD_HighRes_Tryp.param")?;
let db = msgf::db::fasta::ProteinDb::read("human.revCat.fasta", "XXX_")?;
let index = msgf::search::PeptideIndex::build(&db, &DigestParams::default(), &Default::default());The individual crates can also be depended on directly — the facade adds no wrappers, so
msgf::chem::residue_mass and msgf_chem::residue_mass are the same function.
msgf-io ──► msgf-scorer ──► msgf-genfunc
(MGF read) (.param model + (null model + score-distribution DP → SpecEValue) ← hot core
prepared spectrum
→ RawScore)
msgf-chem (masses, residues, fragment ions, tolerance, mass-grid scaling) — used by all
| Crate | Key public API |
|---|---|
msgf-chem |
mass::{PROTON, WATER, …}, scaling::{nominal_bin, high_res_bin}, residue_mass, peptide::{parse, nominal_prefix_masses, accurate_prefix_masses, num_mods}, round_half_up, Tolerance. |
msgf-io |
Spectrum, Peak, MgfReader, read_mgf_file. |
msgf-scorer |
ScoreModel::{from_file, from_bytes}, bundled::{model, score_model}, prepare → PreparedSpectrum, Ms2, Candidate, Residue, Cleavage, match_and_terminal_score; the .param record read_param_file → ScoringModel / write_param; preprocess, peak_by_mass (trainer). |
msgf-genfunc |
NullModel::{uniform, from_probs, from_probs_fixed}, AlphabetEntry, score_distribution → NullTail::{best_possible, tail_mass}. |
msgf-db |
fasta::ProteinDb, enzyme::{Enzyme, DigestParams, digest}, decoy::{write_decoy_database, DecoyOptions, validate_concatenated}. |
msgf-fdr |
TargetDecoyAnalysis::{psm_q_value, pep_q_value}, QValueMap, PsmRecord, peptide_key, is_decoy_match. |
msgf-search |
PeptideIndex::build, SearchEngine::{new, run}, SearchParams, Psm, mods::ModSet, assign_q_values, report::write_tsv. |
This is the whole pipeline — the same flow msgf rescore runs. The null distribution (tail)
depends only on the spectrum, so scoring more peptides against the same spectrum reuses it.
use msgf::{genfunc, io, scorer};
let model = scorer::bundled::score_model()?; // or scorer::ScoreModel::from_file("HCD_HighRes_Tryp.param")?
let spectrum = &io::read_mgf_file("run.mgf")?[0];
let peaks: Vec<(f64, f64)> = spectrum.peaks.iter().map(|p| (p.mz, p.intensity)).collect();
let prep = scorer::prepare(&model, &scorer::Ms2 {
peaks: &peaks,
precursor_mz: spectrum.precursor_mz.unwrap(),
charge: spectrum.charge.unwrap_or(2),
}).expect("precursor nominal mass in 50..=10000");
// Residues with their modification deltas; the N-terminal context is the caller's decision.
let parsed = msgf::chem::peptide::parse("K.SAM+15.995PLER.A").unwrap();
let residues: Vec<scorer::Residue> = parsed.iter()
.map(|r| scorer::Residue { letter: r.aa, delta: r.mod_delta })
.collect();
let cand = scorer::Candidate { residues: &residues, n_term_credit: true }; // flank K → enzymatic
// Null model: residue alphabet + background probabilities, trypsin terminus rule, isotope range.
let null = genfunc::NullModel::from_probs(|_| 0.05, &[(b'M', 15.994915)],
scorer::Cleavage::trypsin(), (0, 1));
let raw_score = scorer::match_and_terminal_score(&prep, &cand, &null.cleavage).unwrap();
let tail = genfunc::score_distribution(&prep, &null, Some(raw_score)).unwrap();
let (denovo, spec_evalue) = (tail.best_possible(), tail.tail_mass(raw_score));Gotchas that matter for fidelity:
- RawScore includes the terminal cleavage terms.
match_and_terminal_scoreadds +2 credit / −11 penalty at each terminus (C-terminal: last residue is a cleavage residue; N-terminal:Candidate::n_term_credit, which the caller sets from the flanking residue, protein start or initiator-Met excision).match_and_terminal_partsreturns the two parts separately. - DB-composition probabilities are required for a search-exact SpecEValue. The null alphabet's probabilities (and so the cleavage mixture weight) must match the searched database's composition, not the uniform 0.05 default, or the tail will differ from MS-GF+.
- Pass a pruning threshold.
score_distribution(.., Some(θ))is exact for every score ≥ θ and much cheaper; use the lowest RawScore you will query (Nonebuilds the full distribution). - Build the distribution once per spectrum. It is independent of the candidate peptide;
amortize it across every PSM for that
(scan, charge).
To match a specific MS-GF+ search bit-for-bit (RawScore & DeNovoScore exact, SpecEValue to f64 noise), the scoring configuration must match the search:
- Model — MS-GF+'s own
.paramfor the acquisition (e.g.HCD_HighRes_Tryp.paramfor-inst 1), passed with--param. The bundled default will not do: it is a different trained model, so it is a different scoring function. This is the requirement people miss. - Amino-acid probabilities — the searched database's composition via
--aa-probs(DBScanner.setAminoAcidProbabilities), not the uniform default. - Variable mods — the same mod set in the graph alphabet (e.g.
--ox-mfor oxidation on M). - Isotope error — the same
--tirange as the search (default0,1).
Under this configuration the CLI matches MS-GF+ 30/30 on the F13 iPRG-2013 golden set (RawScore
and DeNovoScore exact; SpecEValue worst |Δlog10| ≈ 3e-5). This was measured with the previous
implementation, and the clean-room one is byte-identical to it on F13 with MS-GF+'s models. This is exercised by
rust/crates/msgf-cli/tests/golden_rescore.rs and skips gracefully when the (UC-licensed, un-vendored)
reference data is absent — run validation/fetch_reference_data.sh to enable it.
Fidelity is the contract. The value of this project is that the output is bit-exact to MS-GF+,
not an approximation — given MS-GF+'s model and matching configuration, as above. Integer scores
match exactly; SpecEValue within |Δlog10| ≤ 0.05. See
CLAUDE.md and plans/PLAN.md for the validation strategy and the data-absence
contract.
MSGF_Rust/
├── README.md # this file
├── plans/ # design + execution plans (PLAN, PLAN1 model, PLAN2 FDR, PLAN3 p-value speed)
├── ALGORITHMIDEAS.md # index of algorithm/performance research
├── research-trials/ # detailed, reproducible reports behind that index
├── PERFORMANCE.md # Rust-vs-Java timings (of the pre-clean-room implementation)
├── docs/cleanroom/ # SPEC.md, SPEC_ISSUES.md, PROVENANCE.md of the clean-room scorer
├── rust/ # the Cargo workspace
│ ├── README.md # workspace/crate status
│ └── crates/
│ ├── msgf-chem/ # masses, residues, fragment ions, tolerance, mass scaling
│ ├── msgf-io/ # Spectrum types + MGF reader
│ ├── msgf-scorer/ # .param loader/writer + spectrum preparation + RawScore (clean-room)
│ │ └── models/ # the bundled MIT/CC0-clean default scoring model
│ ├── msgf-genfunc/ # generating-function DP → DeNovoScore, SpecEValue (clean-room) ← hot core
│ ├── msgf-train/ # trainer: annotated spectra → a .param we own
│ ├── msgf-db/ # FASTA, digestion, decoys
│ ├── msgf-fdr/ # target-decoy q-values
│ ├── msgf-search/ # the database search engine
│ ├── msgf/ # umbrella library crate (facade over the msgf-* crates)
│ └── msgf-cli/ # the `msgf` binary (search / rescore / decoy / fdr)
├── LICENSING.md # what ships vs. what is fetched; the clean-room boundary
├── validation/ # cross-language oracle: goldens generated locally (see LICENSING.md)
└── .github/workflows/ # release.yml — builds + publishes msgf binaries on tag
License: MIT (LICENSE) — including the bundled scoring model, which was trained
here from the CC0 MassIVE-KB corpus. Nothing MSGF_Rust distributes is UC-licensed: MS-GF+'s
.param models, its test spectra, and every golden derived by running it are fetched or
regenerated on demand by validation/fetch_reference_data.sh and
validation/reference/build_all_golden.sh. The scorer and generating function are a clean-room
reimplementation (2026-09-30) from a written specification. The implementer never saw MS-GF+
source, and the earlier Java-derived files were deleted unread
(docs/cleanroom/PROVENANCE.md). The full accounting is in
LICENSING.md. It covers what ships and what does not, the clean-room boundary, the
remaining behaviour-compatible components (FDR, decoys), and the caveat that git history still
contains the previously committed goldens and the replaced files.