Skip to content

Tracking: PRs for #31, #95, #172, #218, #223, #270, #279, #280 #277

Description

@BenjaminDEMAILLE

Tracking the batch of PRs opened for the six open issues that had no PR or assignee, plus two bugs found along the way (#279, #280).

PRs

Issue PR Status Closes issue?
#31 NH cap vs STAR #271 Ready for review. Stacked on #254 (base fix/pair-subset-loci), adds a same chromosome/strand guard Via #254
#218 bz2/zstd/xz input #272 Ready for review Yes
#95 paraseq input #273 Ready for review Yes
#172 solo CellRanger flags #274 Ready for review Partially (Refs)
#270 hdf5-pure .h5 output #275 Proposal, design open No (Refs)
#223 thread scaling #276 Ready for review (parallel BGZF compression) No (Refs)
#223 thread scaling #278 Ready for review. Stacked on #276: SAM/BAM record encoding on the align workers, gated on core saturation With #276, yes
#279 SE WithinBAM crash #282 Ready for review. STAR hard-clip rules for SE supplementary records Yes
Bulk total RNA (unspliced targets) #283 Ready for review. Opt-in <gene_id>-I targets in TranscriptomeSAM, sp tag, GeneSplicing table; fixes a TranscriptomeSAM soft-clip crash n/a (Refs COMBINE-lab/salmon#1229)
#280 per-read 10k reservation #281 Ready for review. Peak RSS ~8x lower, 6-17% faster Yes

Checklist

Suggested merge order and known conflicts

  1. fix(io): decode multi-member gzip input instead of truncating it #219 / feat(io): optional pure-Rust parallel gzip input decoding (rapidgzip-core) #225 vs feat(io): detect input compression by magic bytes, optional bz2/zstd/xz #272: all three touch the gzip branch of src/io/fastq.rs. feat(io): detect input compression by magic bytes, optional bz2/zstd/xz #272 already uses MultiGzDecoder (keep its version over fix(io): decode multi-member gzip input instead of truncating it #219's; fix(io): decode multi-member gzip input instead of truncating it #219's solo fixes are still needed). feat(io): optional pure-Rust parallel gzip input decoding (rapidgzip-core) #225's rapidgzip call moves into the Gzip arm of src/io/compression.rs. feat(io): detect input compression by magic bytes, optional bz2/zstd/xz #272 and feat(io): optional pure-Rust parallel gzip input decoding (rapidgzip-core) #225 both add a [features] table to Cargo.toml.
  2. feat(io): optional paraseq FASTQ backend #273 after feat(io): detect input compression by magic bytes, optional bz2/zstd/xz #272: paraseq should read from feat(io): detect input compression by magic bytes, optional bz2/zstd/xz #272's open_decoded so it inherits format detection and multi-member gzip.
  3. fix(solo): port STAR's cbMinP, QSmax and oneExact for 1MM_multi barcode correction #274 vs solo: cbMinP posterior threshold, oneExact guard, adapter-anchored geometry (replaces #150) #165: fix(solo): port STAR's cbMinP, QSmax and oneExact for 1MM_multi barcode correction #274 supersedes the cbMinP / oneExact part of solo: cbMinP posterior threshold, oneExact guard, adapter-anchored geometry (replaces #150) #165. Conflicts in resolve_multi_cb and CHANGELOG.md.
  4. feat(solo): optional CellRanger v3 .h5 count matrices via hdf5-pure #275 vs feat: output anndata #237: trivial conflict in src/lib.rs next to write_gene_matrix. Different flag and feature (--soloOutH5 / hdf5-out vs --soloOutputFormat / anndata-out).
  5. perf(io): parallel order-preserving BGZF compression for BAM output #276: independent of perf(align): cut the align batch size and stop cloning every FASTQ record #222, perf(pipeline): size alignment batches so the result vector stays off the arena path #261, feat(io): optional pure-Rust parallel gzip input decoding (rapidgzip-core) #225, chore(deps): bump noodles 0.113 -> 0.115, noodles-bgzf 0.49 -> 0.51 #211, chore(deps): bump noodles from 0.113.0 to 0.116.0 #259. perf(io): encode SAM/BAM records on the align workers (stacked on #276) #278 (stacked on perf(io): parallel order-preserving BGZF compression for BAM output #276): simulated 3-way merge of src/lib.rs against perf(align): cut the align batch size and stop cloning every FASTQ record #222 and perf(pipeline): size alignment batches so the result vector stays off the arena path #261 gives 0 conflicts.

Open points

🤖 Generated with Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions