Skip to content

perf(io): parallel order-preserving BGZF compression for BAM output - #276

Open
BenjaminDEMAILLE wants to merge 1 commit into
mainfrom
perf/parallel-bgzf-writer
Open

BenjaminDEMAILLE wants to merge 1 commit into
mainfrom
perf/parallel-bgzf-writer

Conversation

@BenjaminDEMAILLE

Copy link
Copy Markdown
Contributor

Refs #223 (option 1: parallel compression; record encoding still single-threaded).

Wall time, 2M SE 100 bp reads, synthetic 20 Mb genome, BAM Unsorted, 16-core Mac (median of 3):

threads no output BAM before BAM after
1 12.35s 12.88s 12.94s
4 5.17s 5.35s 5.42s
8 2.87s 3.08s 2.97s
12 2.30s 3.40s 2.54s
16 2.32s 4.18s 2.83s

At 16 threads with --outBAMcompression 6: 5.00s → 2.92s.

Follow-up (not here): move BAM record encoding into the parallel align stage (conflicts with #222/#261), and bam_dedup.rs's own writer.

Tests: fmt, clippy 0 warnings, cargo test 636 passed.

🤖 Generated with Claude Code

New BgzfWriter: the caller fills 64 KiB blocks, --runThreadN dedicated
threads compress them with libdeflate, and one writer thread emits them in
submission order. Output is byte-identical to noodles_bgzf at every thread
count. All four BAM writers use it. finish() now writes the EOF marker and
propagates errors instead of relying on drop.

Refs #223

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant