Hello! Thank you for making your data accessible on GitHub! I have a few questions regarding your table 3:
- Just to confirm: are the files labeled scSplit_barcodes_0-8 the scSplit classifications of the cell barcodes?
- When discussing this data, the paper mentions comparisons of scSplit and demuxlet on known truth from cell hashing. Are the known truth calls represented in the files labeled A-H, or are these files something else like the demuxlet calls? Also, can you please confirm how the known truth calls were generated—are these the classifications that were made by Seurat’s demultiplexing algorithm in the Stoeckius et al. paper?
- The 7,392 cell barcodes in the table 3 files are a subset of the >12,000 cells in figure 2 of Stoeckius et al. How was this subset obtained—are these the cell barcodes that remained after the scSplit pipeline (perhaps because many barcodes were filtered out due to an inability to classify with such low-read depth), or was there other processing involved?
- Finally, I noticed that if I merged the data from A-H and the data from scSplit_barcodes_0-8, there are cases where some cell barcodes only have one of the two classifications. Because scSplit_barcodes_8.csv has barcodes for doublets, I inferred that the cases without a scSplit_barcode classification are negatives—is this correct? What should the A-H classification be for cells with an scSplit_barcode classification but currently without an A-H classification (my analysis seems to suggest these are doublets)?
Thank you!
Hello! Thank you for making your data accessible on GitHub! I have a few questions regarding your table 3:
Thank you!