Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 4 additions & 2 deletions .github/workflows/crisprworks-fit.yml
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,8 @@ jobs:
python-version: ${{ matrix.python }}
- name: Install experimental package
run: python -m pip install ./packages/crisprworks-fit
- name: Require compiled engine
run: python -c 'from crisprworks_fit.kernels import resolved_engine; assert resolved_engine("native") == "native"'
- name: Test scientific and workflow parity
run: python -m unittest discover -s packages/crisprworks-fit/tests -v
- name: Check installed command
Expand Down Expand Up @@ -90,7 +92,7 @@ jobs:
echo 'run=true' >> "$GITHUB_OUTPUT"
else
git diff --name-only "$BEFORE" HEAD > /tmp/fit-changed-paths
if grep -Eq '^packages/crisprworks-fit/(src/|benchmarks/.*\.py|pyproject\.toml|container-constraints\.txt)' /tmp/fit-changed-paths; then
if grep -Eq '^packages/crisprworks-fit/(src/|benchmarks/.*\.py|pyproject\.toml|setup\.py|container-constraints\.txt)' /tmp/fit-changed-paths; then
echo 'run=true' >> "$GITHUB_OUTPUT"
else
echo 'run=false' >> "$GITHUB_OUTPUT"
Expand All @@ -104,7 +106,7 @@ jobs:
if: steps.changed.outputs.run == 'true'
- name: Paired public full-library benchmark with numerical parity
if: steps.changed.outputs.run == 'true'
run: python packages/crisprworks-fit/benchmarks/public_hap1.py --out-dir fit-hap1 --repeats 3 --cohort "${{ inputs.cohort || 'four-guide' }}"
run: python packages/crisprworks-fit/benchmarks/public_hap1.py --out-dir fit-hap1 --repeats 3 --kernel native --compare-numpy --cohort "${{ inputs.cohort || 'four-guide' }}"
- uses: actions/upload-artifact@v7
if: always()
with:
Expand Down
4 changes: 4 additions & 0 deletions packages/crisprworks-fit/.gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -4,3 +4,7 @@ __pycache__/
build/
dist/
.venv/

*.so
*.pyd
*.o
48 changes: 32 additions & 16 deletions packages/crisprworks-fit/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,8 @@
**Experimental acceleration of MAGeCK2 MLE for pooled CRISPR screens.**

Fit runs MAGeCK2's count-table-to-gene-results workflow with faster numerical
kernels and worker reuse. It belongs to the CRISPRWorks family alongside
kernels and worker reuse. Alpha 2 adds a compiled C EM loop; the earlier
NumPy engine remains selectable for reproducible comparisons. It belongs to the CRISPRWorks family alongside
CRISPRWorks Count, powered by DotMatch. This is a separate installable package;
the existing DotMatch package remains independent of its dependencies.

Expand All @@ -23,7 +24,11 @@ crisprworks-fit demo --out-dir fit-demo/

MAGeCK2 is an installation dependency and currently builds its bundled C++
helpers, so its installation requires a C++ compiler and `make`. Python 3.10+
is required by Fit. The package has not been published to PyPI.
is required by Fit. Building the optional native loop requires a C compiler.
`--kernel auto` uses the native loop when available and otherwise NumPy;
`--kernel native` fails explicitly if the compiled engine is unavailable;
`--kernel numpy` selects the previous engine. The provenance manifest records
the engine actually used. The package has not been published to PyPI.

The demo creates a small synthetic count table and design, then writes gene
and guide summaries. It demonstrates software behavior, not biological accuracy.
Expand Down Expand Up @@ -80,6 +85,9 @@ processes; `--blas-threads` controls BLAS threads within each worker.

## What is faster

- Execute the EM loop in compiled C, with working storage reused across
iterations. For finite inputs, process exact nonzero design entries while
preserving row summation order; nonfinite inputs use the dense calculation.
- Apply IRLS weights by vector multiplication and solve the smaller ridge
system directly. No dense diagonal weight matrix is constructed.
- Calculate the sandwich covariance after the last iteration, and compute the
Expand Down Expand Up @@ -108,18 +116,25 @@ checks dense-versus-compact covariance algebra, restarts, strict permutation
tails including ties and NaNs, FDR, source compatibility, failure restoration,
invalid designs, the public upstream count-table fixture and spawned workers.

The [genome-wide HAP1 cohort measurement](benchmarks/HAP1.md) covered 17,445
complete four-guide gene labels: median wall time was **199.74 s reference vs
75.63 s accelerated (2.64×)** across three runs per backend. Printed gene
summaries were byte-identical, permutation p-values/FDR matched exactly, and
full-precision beta differences were below `1.25e-14`. This selects complete
four-guide labels before normalization; it is not a timing for the unfiltered
table or evidence of biological hit accuracy.

The unfiltered 71,090-guide table also passed a separate cross-environment
full-precision comparison with identical printed gene results; see the HAP1
record for how the CI reference and local accelerated outputs were compared.
No unfiltered-table timing claim is made.
The [native-engine HAP1 measurement](benchmarks/NATIVE.md) covered 17,445
complete four-guide gene labels and 69,780 guides. Three paired runs per engine
measured **332.45 s MAGeCK2 reference, 131.06 s NumPy accelerator, and 38.68 s
native accelerator**: **8.59× versus reference and 3.39× versus NumPy**.
All nine printed gene summaries were byte-identical; permutation p-values and
FDR matched exactly. Maximum full-precision beta differences were `2.49e-14`.
These timings include startup and diagnostic output on one shared CI runner.

This cohort selects complete four-guide labels before normalization and
excludes larger control bins. The unfiltered 71,090-guide, 18,056-label table
also matched the reference in a separate cross-environment full-precision
comparison. That comparison establishes numerical agreement, with no
unfiltered-table timing ratio claimed. These checks preserve the reference
results; they do not establish superior biological hit accuracy or a universal
speedup. Chronos and JACKS have not been benchmarked in this comparison.

The [earlier NumPy measurement](benchmarks/HAP1.md) remains archived with its
original implementation commit and runner timings. Compare engines within each
paired experiment rather than combining times from different environments.

Initial measurements are documented in [benchmarks/RESULTS.md](benchmarks/RESULTS.md).
The initial records cover two bounded workloads at commit `e900ecdb`; they do
Expand Down Expand Up @@ -162,8 +177,9 @@ commands and input hashes. The underlying covariance algebra is tested at
CNV correction and `--debug-gene` are rejected until separately evaluated.
Experimental Bayes options retain upstream's explicit rejection. The upstream
active fitting path does not apply `--remove-outliers`; Fit preserves that
behavior. Fit is currently an optional Python/NumPy accelerator, with native
linear algebra provided by BLAS. It does not implement counting or change
behavior. Fit offers a compiled C EM engine and a NumPy fallback, with BLAS
used for covariance calculations. The native loop omits exact zero design
entries for finite inputs and retains a dense path for nonfinite inputs. It does not implement counting or change
DotMatch's read-assignment rules.

Keep counting comparisons separate from inference comparisons. A faster fit
Expand Down
64 changes: 64 additions & 0 deletions packages/crisprworks-fit/benchmarks/NATIVE.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
# Native-engine HAP1 benchmark, 2026-09-30

CRISPRWorks Fit alpha 2 at implementation commit
`2938b8e6e2219e80404280facf14f323e55ee0c3` analyzed **17,445 gene labels and
69,780 guides** from the pinned public Hart Lab BAGEL HAP1 table. All three
engines used identical counts, design, options and random seed. Execution order
rotated across three repetitions per engine on the same shared CI runner.

| Engine | Median wall time | Individual wall times |
| --- | ---: | --- |
| Unmodified MAGeCK2 0.3.0 fitting functions | 332.452 s | 336.551, 330.837, 332.452 s |
| Previous Fit NumPy engine | 131.060 s | 131.060, 129.386, 132.549 s |
| Fit native engine | 38.681 s | 38.838, 38.590, 38.681 s |

**8.59× faster than reference; 3.39× faster than the NumPy accelerator.**
Wall time includes interpreter startup, the count-table workflow and
full-precision diagnostic output. All nine printed gene summaries were
byte-identical. All permutation p-values and FDR values, including both tails,
matched exactly. Maximum absolute differences were `2.49e-14` for beta,
`3.27e-14` for Wald z-scores and `2.01e-14` for Wald FDR.

The design has one T0 baseline and three T18 samples with one treatment effect.
Settings: updated guide efficiencies, mean-variance modeling using 1,000 genes,
two permutation rounds, seed 42, one worker and one BLAS thread. This cohort
selects labels with exactly four guides before normalization, excluding
incomplete labels and larger control bins. It is not the unfiltered table.

The [machine-readable record](hap1-native-four-guide.json) includes dataset
provenance, hashes, all commands and run manifests, versions, BLAS configuration
and numerical differences. The [successful CI run](https://github.com/dnncha/dotmatch/actions/runs/36724219194)
retains all nine complete result sets. Artifact ID: `11102737825`; archive
SHA-256: `cf30d6e98a92efb8e2252e66592e9c59fd6afcc3bf8e4a2c8645e9c954b097ae`.

To reproduce with a compiled install:

```bash
python packages/crisprworks-fit/benchmarks/public_hap1.py \
--out-dir hap1-native --cohort four-guide --repeats 3 \
--kernel native --compare-numpy
```

The native implementation uses a checked SciPy special-function C API,
partial-pivot LU without extra regularization, and compiler settings that
exclude fast-math and floating-point contraction. It omits exact zero design
entries for finite inputs while retaining a dense path for nonfinite inputs.
Tests cover multiple designs, efficiency modes, low counts, matrix adapters,
nonfinite initialization, singular solves, invalid buffers, engine fallback and
state restoration. Linux, macOS, installed wheels and the container passed CI.
Local fitting tests also passed AddressSanitizer and undefined-behavior checks.

The [unfiltered-table parity record](hap1-native-full-parity.json) separately
covers all 71,090 guides and 18,056 labels. It compares a completed CI reference
with local native outputs using identical input hashes, settings and seed.
Printed gene summaries matched byte for byte, all permutation statistics
matched exactly, and maximum absolute differences were below `5.53e-14`.
It is a cross-environment numerical comparison; no timing ratio is claimed.
Upstream's default skip and permutation fallback for the 96-guide label remain.

These results establish a measured performance improvement with numerical
compatibility on this dataset. They do not establish universal performance,
biological superiority or overall state of the art. Chronos, JACKS, legacy
MAGeCK releases and independent multi-screen hit-quality benchmarks remain
outside this comparison. Earlier NumPy timing records use different runners
and must not be combined with these times to calculate a speedup.
Loading
Loading