diff --git a/docs/tutorials/compounds/dose_response.ipynb b/docs/tutorials/compounds/dose_response.ipynb index dd37729..b8fa76e 100644 --- a/docs/tutorials/compounds/dose_response.ipynb +++ b/docs/tutorials/compounds/dose_response.ipynb @@ -12,7 +12,7 @@ "\n", "The dataset is [the OASIS pilot](../../datasets/oasis_pilot.ipynb). OASIS is a liver toxicity screen: its main cell\n", "line is HepaRG, a hepatic progenitor line that differentiates into hepatocyte-like cells and keeps much of the\n", - "drug-metabolising machinery a real liver has. That is the point of using it — a compound that only becomes toxic\n", + "drug-metabolising machinery a real liver has. That is the point of using it: a compound that only becomes toxic\n", "after the liver metabolises it will not show up in a cell line that cannot metabolise it. Most of its thirty-six\n", "compounds were dosed over ten concentrations, a three-fold ladder from 15 nM to 300 uM, with eight replicate wells\n", "at each concentration spread over eight plates; a few, such as berberine, over a different range.\n", @@ -81,11 +81,7 @@ } }, "outputs": [], - "source": [ - "import plotly.io as pio\n", - "\n", - "pio.renderers.default = \"notebook_connected\"" - ] + "source": [] }, { "cell_type": "markdown", @@ -803,7 +799,7 @@ "`Cytoplasm_AreaShape_Zernike_*` are Zernike moments: an orthogonal decomposition of the cytoplasm outline, so a\n", "change means the cells are changing shape rather than getting brighter. `Nuclei_RadialDistribution_FracAtD_DNA_3of4`\n", "is the fraction of the DNA signal falling in the third of four concentric rings measured out from the nucleus\n", - "centre — chromatin moving toward the nuclear periphery, as chromatin does when it condenses against the nuclear envelope early in apoptosis.\n", + "centre, chromatin moving toward the nuclear periphery, as chromatin does when it condenses against the nuclear envelope early in apoptosis.\n", "\n", "`spearman` and `bmd` disagree in the first row. A monotone trend across the whole range and a threshold crossing\n", "are different questions: a feature that steps up early and then flattens has a low Spearman and a low benchmark\n", @@ -946,7 +942,7 @@ "\n", "Staurosporine is the case a single number hides. Its amplitude reaches 1.5 at 0.41 uM and 1.9 at 1.2 uM,\n", "then stops: 2.1 at 33, 2.0 at 300. As a distance it saturates more than two decades below the top of the range. But its plate\n", - "halves agree from 0.41 uM on, 0.76 and above, so each of those concentrations has a direction that reproduces —\n", + "halves agree from 0.41 uM on, 0.76 and above, so each of those concentrations has a direction that reproduces,\n", "and `cosine_to_top` stays between 0.16 and 0.29. The cell is doing something different at 0.41 uM than at 300 uM.\n", "The next section shows what.\n", "\n", @@ -1046,13 +1042,13 @@ "\n", "At 300 uM a different set takes over, and all of them are channel-to-channel correlations inside the nucleus\n", "region: AGP against Mito, Mito against RNA, AGP against ER. AGP is the actin, Golgi and plasma-membrane stain.\n", - "Every stain agreeing with every other stain in the same place is not a phenotype in the usual sense — it is what\n", + "Every stain agreeing with every other stain in the same place is not a phenotype in the usual sense: it is what\n", "the image looks like when the cell has lost its internal organisation and the compartments no longer separate.\n", "`Nuclei_Correlation_Correlation_AGP_Mito` reaches 10.8 MADs.\n", "\n", "So the distance saturating from about 1 uM while the direction keeps turning is a real biological sequence: the\n", "kinase-inhibition phenotype arrives first and stops, and cell disintegration takes over later. Two of the\n", - "plotted features change sign at the top: the RNA radial fraction falls below zero and the AGP–Mito correlation\n", + "plotted features change sign at the top: the RNA radial fraction falls below zero and the AGP-Mito correlation\n", "rises from below it. A curve fitted to distance\n", "alone reports one EC50 for both and describes neither." ] @@ -1068,11 +1064,11 @@ "\n", "`dose_direction` labels each concentration with a `phase`:\n", "\n", - "- `silent` — no reproducible response.\n", - "- `responding` — still moving. The step from the concentration below, `step_amplitude`, beats the noise two\n", + "- `silent`: no reproducible response.\n", + "- `responding`: still moving. The step from the concentration below, `step_amplitude`, beats the noise two\n", " independent groups of wells carry.\n", - "- `saturated` — reproducible, but it has stopped changing.\n", - "- `cytotoxic` — more than half the cells are gone, so the profile is the morphology of dying cells whatever else\n", + "- `saturated`: reproducible, but it has stopped changing.\n", + "- `cytotoxic`: more than half the cells are gone, so the profile is the morphology of dying cells whatever else\n", " is true of it. The US EPA's phenotypic pipeline drops these before fitting anything.\n", "\n", "The window is a run, not a scatter: `split_half_cosine` over eight wells is itself noisy, and a response that\n", @@ -1504,7 +1500,7 @@ "metadata": {}, "source": [ "Mupirocin is grey the whole way. Amperozide and cycloheximide are grey until about 20 uM and then respond to the\n", - "top, so seven of their ten concentrations sit below their effective range — which is the expected result of\n", + "top, so seven of their ten concentrations sit below their effective range, which is the expected result of\n", "covering three decades to find a window you cannot predict, not a wasted experiment.\n", "\n", "Staurosporine responds from 0.41 uM to the top without settling, which is the turning direction above seen from\n", @@ -1513,7 +1509,7 @@ "concentration below does not beat the noise.\n", "\n", "Actinomycin D is responding at the lowest concentration tested. It intercalates DNA and blocks RNA polymerase at\n", - "nanomolar concentrations, so 15 nM is already well inside its range and its onset lies below this ladder — the\n", + "nanomolar concentrations, so 15 nM is already well inside its range and its onset lies below this ladder, and the\n", "screen cannot say where. Only its 300 uM\n", "concentration falls below half the cells, where the band turns red. Both edges of a dose series can fall outside it." ] @@ -1612,7 +1608,7 @@ "should have happened yet, the two sit at 0.99 and 0.87.\n", "\n", "A cytotoxicity call is a claim about cell counts, and cell counts are the most batch-sensitive number in\n", - "a screen — check them against the plate layout before believing one." + "a screen: check them against the plate layout before believing one." ] }, { diff --git a/docs/tutorials/compounds/enrichment.ipynb b/docs/tutorials/compounds/enrichment.ipynb index d48bae1..47ae9bb 100644 --- a/docs/tutorials/compounds/enrichment.ipynb +++ b/docs/tutorials/compounds/enrichment.ipynb @@ -53,11 +53,7 @@ } }, "outputs": [], - "source": [ - "import plotly.io as pio\n", - "\n", - "pio.renderers.default = \"notebook_connected\"" - ] + "source": [] }, { "cell_type": "markdown", diff --git a/docs/tutorials/compounds/hits.ipynb b/docs/tutorials/compounds/hits.ipynb index 1f4dad5..19e9dd1 100644 --- a/docs/tutorials/compounds/hits.ipynb +++ b/docs/tutorials/compounds/hits.ipynb @@ -55,11 +55,7 @@ } }, "outputs": [], - "source": [ - "import plotly.io as pio\n", - "\n", - "pio.renderers.default = \"notebook_connected\"" - ] + "source": [] }, { "cell_type": "markdown", diff --git a/docs/tutorials/compounds/mechanism_of_action.ipynb b/docs/tutorials/compounds/mechanism_of_action.ipynb index 7b56efa..b95b942 100644 --- a/docs/tutorials/compounds/mechanism_of_action.ipynb +++ b/docs/tutorials/compounds/mechanism_of_action.ipynb @@ -54,11 +54,7 @@ } }, "outputs": [], - "source": [ - "import plotly.io as pio\n", - "\n", - "pio.renderers.default = \"notebook_connected\"" - ] + "source": [] }, { "cell_type": "markdown", @@ -1346,7 +1342,7 @@ "wrong answer in the blind test above; here their profiles are not even mutually extreme,\n", "which is the same finding with no classifier in the way. Kinase inhibitors recovers one pair of three; the\n", "label covers different kinases with different substrates, so there is no reason for its members\n", - "to converge on one morphology — the annotation is broad, not the assay blind.\n", + "to converge on one morphology: the annotation is broad, not the assay blind.\n", "\n", "That distinction is the point of reading the per-mechanism table rather than the single\n", "number. A low overall recall can mean the map is poor, or it can mean the annotation groups\n", diff --git a/docs/tutorials/data/embeddings.ipynb b/docs/tutorials/data/embeddings.ipynb index 93317a7..ffdb1fa 100644 --- a/docs/tutorials/data/embeddings.ipynb +++ b/docs/tutorials/data/embeddings.ipynb @@ -62,7 +62,7 @@ "metadata": {}, "source": [ "384 numbers per well from OpenPhenom, a Cell Painting foundation model. The metadata is ordinary JUMP\n", - "metadata — source, batch, plate, well, and the compound each well received." + "metadata: source, batch, plate, well, and the compound each well received." ] }, { @@ -74,7 +74,7 @@ "\n", "A CellProfiler feature name parses into a compartment, a feature group and a channel. `openphenom_nahualX_17`\n", "has neither. The loader supplies the annotation columns the schema requires and leaves them empty, rather\n", - "than letting the name parser loose on names with no structure in them — that parser would read\n", + "than letting the name parser loose on names with no structure in them, since that parser would read\n", "`openphenom_nahualX_17` as the `nahualX` group of an `openphenom` object, and the screen would quietly\n", "acquire feature families named after the model's own tensors." ] @@ -324,7 +324,7 @@ "This is not a defect of the embedding, it is the honest answer: it has no notion of a channel or a\n", "compartment. The consequence is that anything keyed on the feature annotation has nothing to work with.\n", "`pl.effect_sizes` colours its bars by feature family, `tl.feature_sets` groups features into sets, and the\n", - "CellProfiler blocklist names features to drop — none of them have anything to key on.\n", + "CellProfiler blocklist names features to drop: none of them have anything to key on.\n", "\n", "What survives is everything that treats a feature as an anonymous number: normalization, sphering, distances,\n", "hit calling, consensus, and every metric. A block of named features measured on the same wells can name the\n", @@ -386,7 +386,7 @@ "resolution, so the object validates and can be written to h5ad. Supplying them empty is the point: the parser\n", "would read `mymodel_17` as structure it does not have.\n", "\n", - "Two things worth copying from that cell. `values` is cast to **float32** — a model's output often arrives\n", + "Two things worth copying from that cell. `values` is cast to **float32**: a model's output often arrives\n", "as float64 once it has been through pandas or a parquet file, and at 100,000 wells by 1,536 dimensions that is\n", "1.2 GB against 600 MB. And the warning is\n", "*printed*, not counted: it names `Metadata_CellCount` as missing. Keep the count if your pipeline has one:\n", diff --git a/docs/tutorials/genetics/crispr.ipynb b/docs/tutorials/genetics/crispr.ipynb index 059f1b2..02e3f96 100644 --- a/docs/tutorials/genetics/crispr.ipynb +++ b/docs/tutorials/genetics/crispr.ipynb @@ -81,11 +81,7 @@ } }, "outputs": [], - "source": [ - "import plotly.io as pio\n", - "\n", - "pio.renderers.default = \"notebook_connected\"" - ] + "source": [] }, { "cell_type": "markdown", @@ -146,7 +142,7 @@ "`non-targeting` is not quite empty, and which wells define the reference matters.\n", "\n", "We test it with a retrieval question on the sphered negcon wells, where plate effects are already gone: does a\n", - "control well retrieve others of its own kind — `no-guide` or `non-targeting` — beyond chance? `pos_sameby` makes a\n", + "control well retrieve others of its own kind (`no-guide` or `non-targeting`) beyond chance? `pos_sameby` makes a\n", "positive pair two wells of one kind and `neg_diffby` makes a negative pair two wells of different kinds, with\n", "`reference=None` so no wells are held out. An mAP near chance means the two kinds are one baseline; an mAP above\n", "chance means they separate. `null_size` can stay modest, as there are only two groups to score." @@ -319,7 +315,7 @@ "source": [ "The table gives an mAP per control kind and a p-value for whether each kind retrieves its own beyond chance; the figure shows the same as two point clouds.\n", "\n", - "The non-targeting wells are told apart from the empty ones beyond chance: their mAP of 0.383 carries a corrected p-value of 0.011, while the no-guide row's higher mAP of 0.639 does not reach significance (p 0.105) — a reminder that an mAP is read against its permutation null, not on its face. So the two negative controls are not one inert baseline. What separates non-targeting from empty is the burden of delivery and expression — the lentiviral integration, the sgRNA and the Cas9 the empty wells never received — not a cut, since a non-targeting guide cuts nothing. Because the sphering above pools both kinds as `negcon`, every knockout is scored against a reference that carries this delivery signature, and the choice of reference — empty wells, non-targeting, or both — moves the map. Worth checking, rather than assuming the two agree, before leaning on one baseline." + "The non-targeting wells are told apart from the empty ones beyond chance: their mAP of 0.383 carries a corrected p-value of 0.011, while the no-guide row's higher mAP of 0.639 does not reach significance (p 0.105), a reminder that an mAP is read against its permutation null, not on its face. So the two negative controls are not one inert baseline. What separates non-targeting from empty is the burden of delivery and expression (the lentiviral integration, the sgRNA and the Cas9 the empty wells never received), not a cut, since a non-targeting guide cuts nothing. Because the sphering above pools both kinds as `negcon`, every knockout is scored against a reference that carries this delivery signature, and the choice of reference (empty wells, non-targeting, or both) moves the map. Worth checking, rather than assuming the two agree, before leaning on one baseline." ] }, { @@ -460,7 +456,7 @@ "id": "c6dda962", "metadata": {}, "source": [ - "The bar names the measurements that carry the difference: `Nuclei_Granularity_6_Mito`, `Cytoplasm_Granularity_1_Mito`, the radial distribution of mitochondrial tubeness across the cell, a `Cytoplasm_Correlation_AGP_DNA` term and a handful of Mito texture and DNA–Mito overlap features. By kind they are granularity, radial-distribution and texture features, sitting mostly in the mitochondrial (Mito) channel, with a little RNA intensity. Whatever the exact list, the signature is one of delivery and expression, not cutting: the non-targeting wells carry a lentiviral integration, an sgRNA and Cas9 that the empty wells never received, and that added burden tilts their staining and morphology — a metabolic, mitochondrial lean is what one would expect from it — while a non-targeting guide makes no cut to leave a sharp local mark. The effects are modest in size: a median absolute Cohen's d of 0.03, reaching 0.25 at the top. The two controls are distinguishable, as the retrieval test above found, but only modestly, so this is a lean of the baseline rather than a phenotype of its own — though modest is not nothing when these wells define the reference every knockout is scored against." + "The bar names the measurements that carry the difference: `Nuclei_Granularity_6_Mito`, `Cytoplasm_Granularity_1_Mito`, the radial distribution of mitochondrial tubeness across the cell, a `Cytoplasm_Correlation_AGP_DNA` term and a handful of Mito texture and DNA-Mito overlap features. By kind they are granularity, radial-distribution and texture features, sitting mostly in the mitochondrial (Mito) channel, with a little RNA intensity. Whatever the exact list, the signature is one of delivery and expression, not cutting: the non-targeting wells carry a lentiviral integration, an sgRNA and Cas9 that the empty wells never received, and that added burden tilts their staining and morphology (a metabolic, mitochondrial lean is what one would expect from it), while a non-targeting guide makes no cut to leave a sharp local mark. The effects are modest in size: a median absolute Cohen's d of 0.03, reaching 0.25 at the top. The two controls are distinguishable, as the retrieval test above found, but only modestly, so this is a lean of the baseline rather than a phenotype of its own, though modest is not nothing when these wells define the reference every knockout is scored against." ] }, { @@ -475,8 +471,8 @@ "Start from a fresh load, so the already-sphered `wells` is left alone, and normalize it four ways: to the empty\n", "wells alone, to the non-targeting wells alone, to both together (`negcon`), and to every well on the plate. Each\n", "is a per-plate robust MAD normalization; they differ only in which rows set the centre and spread. A gene's\n", - "activity is then read cheaply as the L2 norm of its modz consensus — its distance from the reference-centred\n", - "origin — with no permutations. If the reference barely matters, a gene keeps its activity whichever rows define\n", + "activity is then read cheaply as the L2 norm of its modz consensus (its distance from the reference-centred\n", + "origin) with no permutations. If the reference barely matters, a gene keeps its activity whichever rows define\n", "zero." ] }, @@ -690,7 +686,7 @@ "source": [ "The table reads the reference off the ranking directly. Between the two negative controls the gene activities correlate at Spearman r = 0.78, and their top 100 genes by activity overlap by 72% (Jaccard). The scatter shows the same pair gene by gene: points on the dashed diagonal keep their activity whichever baseline defines zero, and the spread away from it is what the reference moves; the labelled genes shift rank the most.\n", "\n", - "Read against the diagonal, the choice of reference **moves** the hits. Even the two negative-control baselines — empty and non-targeting — agree only at r = 0.78 and share 72% of their top 100, so more than a quarter of the strongest hits turn on which controls define zero. Whole-plate normalization, which depends on no control at all, diverges further still: r = 0.43 against the empty baseline and r = 0.66 against the pooled `negcon`. The reference is a first-order analytic decision, not a formality." + "Read against the diagonal, the choice of reference **moves** the hits. Even the two negative-control baselines (empty and non-targeting) agree only at r = 0.78 and share 72% of their top 100, so more than a quarter of the strongest hits turn on which controls define zero. Whole-plate normalization, which depends on no control at all, diverges further still: r = 0.43 against the empty baseline and r = 0.66 against the pooled `negcon`. The reference is a first-order analytic decision, not a formality." ] }, { @@ -703,9 +699,9 @@ "- **Prefer whole-plate normalization.** When perturbations are spread evenly across plates, as they are in this\n", " screen, per-plate robust normalization over *all* wells is the field's primary recommendation, and it sidesteps\n", " the negcon-choice problem entirely: no subset of wells defines zero, so there is no baseline to pick wrong. This\n", - " is the JUMP recipe {cite:p}`Chandrasekaran_2023`, in line with the Carpenter–Singh lab's normalization guidance.\n", + " is the JUMP recipe {cite:p}`Chandrasekaran_2023`, in line with the Carpenter-Singh lab's normalization guidance.\n", "- **Normalize to controls only when the layout forces it.** Reach for control-based normalization when the plate\n", - " layout is uneven, or when many wells are expected to be active and would pull a whole-plate centre off zero — and\n", + " layout is uneven, or when many wells are expected to be active and would pull a whole-plate centre off zero, and\n", " only where there are enough control wells (at least 16, more is better) that are not confined to a single row or\n", " column.\n", "- **When you do, prefer non-targeting over empty wells.** Non-targeting guides are delivery-matched: a knockout is\n", @@ -715,7 +711,7 @@ "- **Cutting toxicity needs safe-targeting guides.** The gold standard for separating Cas9 *cutting* toxicity from\n", " the biology is safe-targeting guides, which cut at non-genic loci and so carry the double-strand-break burden\n", " without knocking out a gene ({cite:p}`Morgens_2017`; {cite:p}`Alvarez_2022`). This screen's control pool has\n", - " none, so cutting toxicity cannot be isolated here — the empty-versus-non-targeting contrast measures delivery and\n", + " none, so cutting toxicity cannot be isolated here, and the empty-versus-non-targeting contrast measures delivery and\n", " expression, not the cut.\n", "\n", "For this notebook the sphering above used both negative controls pooled; the whole-plate route is the cleaner\n", diff --git a/docs/tutorials/genetics/interpreting_a_screen.ipynb b/docs/tutorials/genetics/interpreting_a_screen.ipynb index 1c34ebe..8a95fa1 100644 --- a/docs/tutorials/genetics/interpreting_a_screen.ipynb +++ b/docs/tutorials/genetics/interpreting_a_screen.ipynb @@ -61,13 +61,11 @@ "import numpy as np\n", "import pandas as pd\n", "import plotly.express as px\n", - "import plotly.io as pio\n", "import scanpy as sc\n", "\n", "import mantispy as mt\n", "\n", "# Serialize plotly figures into the notebook so the interactive twins embed under nbconvert.\n", - "pio.renderers.default = \"notebook_connected\"\n", "warnings.filterwarnings(\"ignore\")\n", "SEED = 0\n", "\n", diff --git a/docs/tutorials/multisite/cross_laboratory.ipynb b/docs/tutorials/multisite/cross_laboratory.ipynb index 5bb5777..19cb9df 100644 --- a/docs/tutorials/multisite/cross_laboratory.ipynb +++ b/docs/tutorials/multisite/cross_laboratory.ipynb @@ -53,11 +53,7 @@ } }, "outputs": [], - "source": [ - "import plotly.io as pio\n", - "\n", - "pio.renderers.default = \"notebook_connected\"" - ] + "source": [] }, { "cell_type": "markdown", diff --git a/docs/tutorials/multisite/learned_embeddings.ipynb b/docs/tutorials/multisite/learned_embeddings.ipynb index 217a40f..87fc1fb 100644 --- a/docs/tutorials/multisite/learned_embeddings.ipynb +++ b/docs/tutorials/multisite/learned_embeddings.ipynb @@ -64,11 +64,7 @@ } }, "outputs": [], - "source": [ - "import plotly.io as pio\n", - "\n", - "pio.renderers.default = \"notebook_connected\"" - ] + "source": [] }, { "cell_type": "markdown", @@ -215,13 +211,13 @@ "The question is not whether an embedding works, but how it compares with the alternatives on the same wells.\n", "Two readouts:\n", "\n", - "- **`tl.map`** — mean average precision. Rank every other profile by similarity to a query and ask whether\n", + "- **`tl.map`**: mean average precision. Rank every other profile by similarity to a query and ask whether\n", " its replicates come first. The pairs are defined so that a replicate must come from a *different\n", " laboratory*, which is the cross-site reproducibility question, and copairs gives every compound a\n", " permutation p-value that is then BH-corrected. This is the metric the JUMP benchmarks report\n", " {cite:p}`Kalinin_2025,Arevalo_2024`, and it is the only one of the two that scales: the\n", " similarity-matrix readouts build a dense n-by-n matrix.\n", - "- **`metrics.known_relationships`** — of the compounds annotated to act on the same gene, how many end up in\n", + "- **`metrics.known_relationships`**: of the compounds annotated to act on the same gene, how many end up in\n", " either tail of the similarity distribution?\n", "\n", "Each is measured under three alignments: the principal components alone, `pp.tvn`, and `pp.harmony`, which\n", @@ -510,7 +506,7 @@ "small decision.\n", "\n", "**The CellProfiler-equivalent set leads without a correction.** `cp_measure` tops the uncorrected column\n", - "and calls the most compounds under Harmony — though MorphEm's mean mAP under Harmony is higher, SubCell's is\n", + "and calls the most compounds under Harmony, though MorphEm's mean mAP under Harmony is higher, SubCell's is\n", "level with it and DINOv2's slightly below, and under TVN `cp_measure` falls behind MorphEm and DINOv2.\n", "\n", "Each of those compounds had to be told apart from *other compounds*, across laboratories. The JUMP-Lite\n", @@ -1328,7 +1324,7 @@ "Once you do, the readout says something clean, and something different from what the raw numbers suggest.\n", "**`pp.tvn` lifts recall clearly above its own null on three of the four trained embeddings and on\n", "`cp_measure`, and on DINOv2 only marginally (p = 0.06, five of the 100 shuffles scoring as high).\n", - "Under either of the other two alignments no embedding clears its own null — but `cp_measure` clears it under\n", + "Under either of the other two alignments no embedding clears its own null, but `cp_measure` clears it under\n", "all three.** And the\n", "untrained model never clears it: its apparently high raw recall, the highest in the unaligned column, is\n", "entirely a property of its null, which is high because a near-random projection spreads pairs into the tails.\n", @@ -1517,8 +1513,8 @@ "source": [ "Harmony takes the laboratory down to a fraction of a percent on every feature set. TVN does not\n", "consistently reduce it: it is unchanged on three blocks, up by half on DINOv2, and down on SubCell and the\n", - "untrained model. That is consistent with what the two do — Harmony mixes the batches in every local\n", - "neighbourhood, while TVN aligns each batch's control covariance to the pooled one — and a plausible reason the\n", + "untrained model. That is consistent with what the two do (Harmony mixes the batches in every local\n", + "neighbourhood, while TVN aligns each batch's control covariance to the pooled one) and a plausible reason the\n", "two readouts disagree.\n", "\n", "Read this together with the tables above, not instead of them. TVN leaves the laboratory in place and is still\n", @@ -1902,7 +1898,7 @@ "cross-fits a ridge from the embedding to each `cp_measure` feature and reports the out-of-fold R^2, averaged\n", "over each feature family: near one means the embedding linearly reconstructs that family, near zero means it\n", "does not carry it. The two blocks are the same 1,536 wells, so the metric aligns them on `obs_names` and needs\n", - "nothing else. `classical` is the block to recover — normalized but not selected, so it still names every\n", + "nothing else. `classical` is the block to recover, normalized but not selected, so it still names every\n", "feature and its `var[\"feature_group\"]`." ] }, @@ -2071,8 +2067,8 @@ "id": "254f5aed", "metadata": {}, "source": [ - "The trained embeddings carry the ferret and texture families of `cp_measure` well — an image model\n", - "reconstructs how large each object is and how patterned each channel is almost for free — and the zernike and shape families\n", + "The trained embeddings carry the ferret and texture families of `cp_measure` well (an image model\n", + "reconstructs how large each object is and how patterned each channel is almost for free) and the zernike and shape families\n", "less, since those turn on the exact segmentation the embedding never performed. `dinov2_random` carries little\n", "of any family: the untrained control sits well below the trained models everywhere, which is the point. A\n", "learned embedding stops being a black box once you can say what of the classical block it linearly rebuilds;\n", diff --git a/docs/tutorials/overview.ipynb b/docs/tutorials/overview.ipynb index 44ea23a..f0013a2 100644 --- a/docs/tutorials/overview.ipynb +++ b/docs/tutorials/overview.ipynb @@ -50,11 +50,7 @@ } }, "outputs": [], - "source": [ - "import plotly.io as pio\n", - "\n", - "pio.renderers.default = \"notebook_connected\"" - ] + "source": [] }, { "cell_type": "markdown", @@ -107,8 +103,8 @@ "metadata": {}, "source": [ "Rows are wells, columns are CellProfiler features, and every `Metadata_` column in `obs` says what a well is: its\n", - "plate, its compound and concentration, its mechanism label, its cell count. Nothing here is mantispy-specific yet\n", - "— a scanpy user already knows this object." + "plate, its compound and concentration, its mechanism label, its cell count. Nothing here is mantispy-specific yet.\n", + "A scanpy user already knows this object." ] }, { @@ -173,7 +169,7 @@ "source": [ "The controls sit together and the treated wells fan out around them, so there is a signal to call. Cell count\n", "varies smoothly across the same map, a reminder that a well with few cells can move for reasons unrelated to the\n", - "compound — a confounder we return to under hits." + "compound, a confounder we return to under hits." ] }, { @@ -259,7 +255,7 @@ "metadata": {}, "source": [ "Agreement rises from one replicate to two and is still climbing at the deepest the data allows. BBBC021's two to\n", - "three replicates per treatment are on the low side for separating mechanisms — worth knowing before trusting a\n", + "three replicates per treatment are on the low side for separating mechanisms, worth knowing before trusting a\n", "marginal call." ] }, @@ -471,7 +467,7 @@ "id": "q19", "metadata": {}, "source": [ - "Most treatments clear the threshold and the DMSO control does not — the assay separates active from inactive. The\n", + "Most treatments clear the threshold and the DMSO control does not: the assay separates active from inactive. The\n", "volcano for the strongest hit names the individual measurements carrying its phenotype, which is the handle for\n", "asking what the compound does." ] @@ -544,7 +540,7 @@ "id": "q22", "metadata": {}, "source": [ - "Groups in the upper left are far from the controls and have lost most of their cells — suspect, not necessarily\n", + "Groups in the upper left are far from the controls and have lost most of their cells: suspect, not necessarily\n", "biology. A large distance with intact viability is the hit worth following." ] }, @@ -556,8 +552,8 @@ "## What does it do?\n", "\n", "Mechanism is a retrieval question: does a treatment's nearest neighbour share its mechanism, and where the assay\n", - "fails, which mechanisms does morphology confuse? `tl.nn_moa_classify` with the not-same-compound rule — a compound\n", - "must be recognised from a *different* molecule with the same mechanism — answers the first." + "fails, which mechanisms does morphology confuse? `tl.nn_moa_classify` with the not-same-compound rule (a compound\n", + "must be recognised from a *different* molecule with the same mechanism) answers the first." ] }, { @@ -619,8 +615,8 @@ "metadata": {}, "source": [ "The diagonal dominates: most treatments retrieve a different compound with the same mechanism. The largest\n", - "off-diagonal block, Eg5 inhibitors against microtubule destabilizers, is biology — both arrest mitosis and look\n", - "alike here — not a failure of the classifier." + "off-diagonal block, Eg5 inhibitors against microtubule destabilizers, is biology (both arrest mitosis and look\n", + "alike here) and not a failure of the classifier." ] }, { @@ -690,7 +686,7 @@ "id": "q28", "metadata": {}, "source": [ - "Mechanism blocks stand out along the diagonal — compounds that share a mechanism sit close — while the off-diagonal\n", + "Mechanism blocks stand out along the diagonal (compounds that share a mechanism sit close) while the off-diagonal\n", "shows which mechanisms the assay cannot separate. This is the same structure the confusion matrix summarised, now\n", "laid out treatment by treatment." ] @@ -712,7 +708,7 @@ "id": "q31", "metadata": {}, "source": [ - "`mt.io.write` saves the object and every one of these tables in one `.h5ad`, and `mt.io.read` brings them back —\n", + "`mt.io.write` saves the object and every one of these tables in one `.h5ad`, and `mt.io.read` brings them back,\n", "the same file a scanpy user can open." ] }, @@ -725,15 +721,15 @@ "\n", "Each modality has its own tutorial that goes deeper than this page:\n", "\n", - "- **Compound screens** — [hits, effects and cell loss](compounds/hits.ipynb): activity, what moved, and whether a\n", + "- **Compound screens**, [hits, effects and cell loss](compounds/hits.ipynb): activity, what moved, and whether a\n", " hit is a phenotype or dead cells.\n", - "- **Genetic screens** — [CRISPR knockouts](genetics/crispr.ipynb): which knockouts have a phenotype, and whether\n", + "- **Genetic screens**, [CRISPR knockouts](genetics/crispr.ipynb): which knockouts have a phenotype, and whether\n", " the map recovers known biology.\n", - "- **Single cells** — [what the well median hides](single_cells/heterogeneity.ipynb): the same questions asked below\n", + "- **Single cells**, [what the well median hides](single_cells/heterogeneity.ipynb): the same questions asked below\n", " the well.\n", - "- **Across laboratories** — [cross-laboratory reproducibility](multisite/cross_laboratory.ipynb): does the screen\n", + "- **Across laboratories**, [cross-laboratory reproducibility](multisite/cross_laboratory.ipynb): does the screen\n", " reproduce at another site.\n", - "- **How-to recipes** — [from a CellProfiler run](data/cellprofiler.ipynb) and the other short data-loading\n", + "- **How-to recipes**, [from a CellProfiler run](data/cellprofiler.ipynb) and the other short data-loading\n", " examples." ] } diff --git a/docs/tutorials/phenotypes/which_features_moved.ipynb b/docs/tutorials/phenotypes/which_features_moved.ipynb index 1523baa..0286f7e 100644 --- a/docs/tutorials/phenotypes/which_features_moved.ipynb +++ b/docs/tutorials/phenotypes/which_features_moved.ipynb @@ -49,11 +49,7 @@ } }, "outputs": [], - "source": [ - "import plotly.io as pio\n", - "\n", - "pio.renderers.default = \"notebook_connected\"" - ] + "source": [] }, { "cell_type": "markdown", @@ -1660,7 +1656,7 @@ "metadata": {}, "source": [ "That is a vocabulary for the embedding, and it is specific enough to act on. Both trained models reconstruct\n", - "the *texture* measurements best, and the untrained model reconstructs little of anything — the gap between the\n", + "the *texture* measurements best, and the untrained model reconstructs little of anything, and the gap between the\n", "grey bars and the coloured ones is what training bought, measured in units a CellProfiler user already\n", "understands.\n", "\n", @@ -1874,7 +1870,7 @@ "The first two components have a named measurement that tracks them closely; by the fifth the best match\n", "explains little. So the names are worth having for the leading components and not beyond. The same idea\n", "applied to the embedding's own dimensions, rather than to its components, would give `pl.effect_sizes` and\n", - "`tl.feature_sets` something to key on again — borrowed rather than parsed, and honest as long as you say so." + "`tl.feature_sets` something to key on again, borrowed rather than parsed, and honest as long as you say so." ] }, { diff --git a/docs/tutorials/profiles/artifacts.ipynb b/docs/tutorials/profiles/artifacts.ipynb index c938876..fb57a1f 100644 --- a/docs/tutorials/profiles/artifacts.ipynb +++ b/docs/tutorials/profiles/artifacts.ipynb @@ -81,11 +81,7 @@ } }, "outputs": [], - "source": [ - "import plotly.io as pio\n", - "\n", - "pio.renderers.default = \"notebook_connected\"" - ] + "source": [] }, { "cell_type": "markdown", diff --git a/docs/tutorials/profiles/normalize_and_select.ipynb b/docs/tutorials/profiles/normalize_and_select.ipynb index 06b5784..7431392 100644 --- a/docs/tutorials/profiles/normalize_and_select.ipynb +++ b/docs/tutorials/profiles/normalize_and_select.ipynb @@ -47,11 +47,7 @@ } }, "outputs": [], - "source": [ - "import plotly.io as pio\n", - "\n", - "pio.renderers.default = \"notebook_connected\"" - ] + "source": [] }, { "cell_type": "markdown", @@ -720,14 +716,14 @@ "### Removing multivariate redundancy\n", "\n", "`correlation_threshold` is pairwise: it drops one feature from each pair correlated above the cutoff. A\n", - "feature that is a linear combination of *several* others — a texture that tracks the sum of two neighbouring\n", - "scales, say — need not correlate strongly with any single one, so the pairwise cut keeps it even though it\n", + "feature that is a linear combination of *several* others, a texture that tracks the sum of two neighbouring\n", + "scales, say, need not correlate strongly with any single one, so the pairwise cut keeps it even though it\n", "carries no new information.\n", "\n", "`decorrelate` removes exactly that. A rank-revealing, column-pivoted QR {cite:p}`Businger_1965,Golub_2013`\n", "orders the features so that each is the one least explained by those already kept, and drops a feature once\n", - "the features kept before it predict it at multiple correlation `decorr_threshold`. One decomposition does it —\n", - "cheap enough to run at several strengths — and it keeps original features rather than projecting to components.\n", + "the features kept before it predict it at multiple correlation `decorr_threshold`. One decomposition does it,\n", + "cheap enough to run at several strengths, and it keeps original features rather than projecting to components.\n", "It is experimental and not part of pycytominer, so it is opt-in via `decorrelate=True`.\n", "\n", "How hard to decorrelate is a judgement call, so rather than trust one number, run it at a few thresholds and\n", @@ -879,7 +875,7 @@ "source": [ "`decorrelate` is a dial, and there is no free lunch: a lower threshold removes more features, but because the\n", "redundant ones still carry a sliver of unique signal, retrieval falls with them. The technical replicate\n", - "retrieval is nearly flat here; the biological activity retrieval is the one that pays — gently at 0.99, where\n", + "retrieval is nearly flat here; the biological activity retrieval is the one that pays, gently at 0.99, where\n", "the set nearly halves, and more by 0.9. On a screen with real mechanism-of-action structure the trade is often\n", "kinder still, because dropping redundant, noisy features can tighten same-MOA neighbourhoods, but pki has too\n", "few compounds per MOA to show that here.\n", diff --git a/docs/tutorials/profiles/quality_control.ipynb b/docs/tutorials/profiles/quality_control.ipynb index cabf893..298aabd 100644 --- a/docs/tutorials/profiles/quality_control.ipynb +++ b/docs/tutorials/profiles/quality_control.ipynb @@ -66,11 +66,7 @@ } }, "outputs": [], - "source": [ - "import plotly.io as pio\n", - "\n", - "pio.renderers.default = \"notebook_connected\"" - ] + "source": [] }, { "cell_type": "markdown", diff --git a/docs/tutorials/profiles/screen_quality.ipynb b/docs/tutorials/profiles/screen_quality.ipynb index 9893c23..a3bf219 100644 --- a/docs/tutorials/profiles/screen_quality.ipynb +++ b/docs/tutorials/profiles/screen_quality.ipynb @@ -42,11 +42,7 @@ } }, "outputs": [], - "source": [ - "import plotly.io as pio\n", - "\n", - "pio.renderers.default = \"notebook_connected\"" - ] + "source": [] }, { "cell_type": "markdown", diff --git a/docs/tutorials/single_cells/heterogeneity.ipynb b/docs/tutorials/single_cells/heterogeneity.ipynb index 178f779..71aaf1b 100644 --- a/docs/tutorials/single_cells/heterogeneity.ipynb +++ b/docs/tutorials/single_cells/heterogeneity.ipynb @@ -63,11 +63,7 @@ } }, "outputs": [], - "source": [ - "import plotly.io as pio\n", - "\n", - "pio.renderers.default = \"notebook_connected\"" - ] + "source": [] }, { "cell_type": "markdown",