diff --git a/docs/introduction/performance.md b/docs/introduction/performance.md index 413bbe4..9090195 100644 --- a/docs/introduction/performance.md +++ b/docs/introduction/performance.md @@ -37,13 +37,13 @@ modern-di variant against the rivals whose API matches it, because a single colu flatter modern-di against half the set. By-type resolution used to add a fixed lookup cost on top of `resolve_provider`: 54-65 ns through 3.2.0, then 21/17/23 ns on C1/C2/C3 once 3.3.0 inlined `resolve_provider`'s body into `resolve`. With the template resolver, `Container.resolve` memoizes -type → resolver directly, and the two columns are within run-to-run noise of each other (253 vs -247 ns on C1, 151 vs 147 on C2, 789 vs 798 on C3 in the cells below). +type → resolver directly, and the two columns differ by a few nanoseconds in either direction (242 +vs 239 ns on C1, 146 vs 142 on C2, 760 vs 764 on C3 in the cells below). C4 does not split this way: modern-di's C4 body resolves by reference throughout, while dishka and wireup can only be measured by type. That asymmetry cuts against modern-di's C4 ratios, though with the by-type surcharge now inside noise, levelling it would not -move the dishka ratio below (1.08). The C1-C3 leveling does not apply to C4. +move the dishka ratio (1.12) below 1.0. The C1-C3 leveling does not apply to C4. C6 does not split either, for the same reason: modern-di's C6 body resolves by reference and there is no by-type C6 variant to pair against dishka and wireup, so a split would leave that half of the @@ -59,14 +59,14 @@ for why that matters and which cells it moved. ## Results -Measured 2026-09-27 with modern-di `main` at `630de77` (3.5.0 plus #540, #541 and #542, -unreleased when measured) on an Apple M2 (macOS 26.6.2), CPython 3.14.7, median over 5 runs -(ratios paired within each run); the footnote under each table bounds the across-run dispersion -of each side's own median. Rival versions: dishka 1.10.1, dependency-injector 4.49.1, -that-depends 4.1.0, wireup 2.12.0. Generated by `just bench-report`. +Measured 2026-10-05 with modern-di `main` at `f300c2e` (the 4.0 code, unreleased when measured) +on an Apple M2 (macOS 26.6.2), CPython 3.14.7 with the GIL, median over 5 runs (ratios paired +within each run); the footnote under each table bounds the across-run dispersion of each side's +own median. Rival versions: dishka 1.10.1, dependency-injector 4.49.1, that-depends 4.1.0, +wireup 2.12.0. Generated by `just bench-report`. -This publication is on the same machine and CPython build as the 3.5.0 one, so this time the -absolute cells are comparable with the previous tables as well as the ratios; the +The previous publication (3.x, at `630de77`) used the same machine, CPython build and rival +versions, so the absolute cells are comparable with it as well as the ratios. The [what moved](#what-moved-in-this-publication) section below reads both. > **These numbers are a snapshot of the version named above; a newer release does not update @@ -86,107 +86,106 @@ as a verdict. | Scenario | modern-di | vs dependency-injector | vs that-depends | |---|---|---|---| -| C1 transient | 253 ns ±1.9% | **0.42** ±2.8% | **0.58** ±2.4% | -| C2 warm singleton | 151 ns ±0.6% | 2.23 ±7.8% | 1.78 ±1.0% | -| C3 deep chain (6) | 789 ns ±1.4% | **0.36** ±5.3% | **0.53** ±0.2% | +| C1 transient | 242 ns ±0.4% | **0.41** ±0.5% | **0.56** ±0.4% | +| C2 warm singleton | 146 ns ±0.6% | 2.20 ±0.6% | 1.70 ±1.6% | +| C3 deep chain (6) | 760 ns ±1.4% | **0.35** ±1.3% | **0.51** ±1.3% | -_Across-run IQR of each side's own median (5 runs): modern-di ≤1.9%, rivals ≤7.3%. The ± on each ratio cell is a different quantity: the spread of the paired per-run ratios._ +_Across-run IQR of each side's own median (5 runs): modern-di ≤1.4%, rivals ≤0.9%. The ± on each ratio cell is a different quantity: the spread of the paired per-run ratios._ ### By-type resolution | Scenario | modern-di | vs dishka | vs wireup | |---|---|---|---| -| C1 transient | 247 ns ±3.4% | **0.72** ±3.1% | **0.82** ±2.4% | -| C2 warm singleton | 147 ns ±0.7% | **0.62** ±0.5% | 1.47 ±1.0% | -| C3 deep chain (6) | 798 ns ±3.2% | 1.25 ±0.7% | **0.87** ±2.9% | +| C1 transient | 239 ns ±0.1% | **0.70** ±0.5% | **0.80** ±0.1% | +| C2 warm singleton | 142 ns ±0.3% | **0.60** ±1.2% | 1.41 ±1.1% | +| C3 deep chain (6) | 764 ns ±1.4% | 1.21 ±1.9% | **0.83** ±4.3% | -_Across-run IQR of each side's own median (5 runs): modern-di ≤3.4%, rivals ≤1.3%. The ± on each ratio cell is a different quantity: the spread of the paired per-run ratios._ +_Across-run IQR of each side's own median (5 runs): modern-di ≤1.4%, rivals ≤2.8%. The ± on each ratio cell is a different quantity: the spread of the paired per-run ratios._ ### Request lifecycle (batched, published per request) | Scenario | modern-di | vs dependency-injector | vs that-depends | vs dishka | vs wireup | |---|---|---|---|---|---| -| C4 request lifecycle | 2.29 µs ±4.8% | **0.02** ±6.9% | **0.18** ±7.9% | 1.08 ±0.3% | **0.14** ±0.9% | +| C4 request lifecycle | 2.28 µs ±0.4% | **0.02** ±0.9% | **0.18** ±0.4% | 1.12 ±0.8% | **0.14** ±0.6% | -_Across-run IQR of each side's own median (5 runs): modern-di ≤4.8%, rivals ≤7.2%. The ± on each ratio cell is a different quantity: the spread of the paired per-run ratios._ +_Across-run IQR of each side's own median (5 runs): modern-di ≤0.4%, rivals ≤1.5%. The ± on each ratio cell is a different quantity: the spread of the paired per-run ratios._ ### Per-request context | Scenario | modern-di | vs dependency-injector | vs that-depends | vs dishka | vs wireup | |---|---|---|---|---|---| -| C6 context | 1.45 µs ±1.1% | **0.38** ±1.4% | **0.48** ±1.6% | 1.15 ±2.2% | 1.02 ±2.4% | +| C6 context | 1.09 µs ±0.1% | **0.28** ±0.4% | **0.36** ±1.2% | **0.88** ±1.3% | **0.78** ±2.3% | -_Across-run IQR of each side's own median (5 runs): modern-di ≤1.1%, rivals ≤1.9%. The ± on each ratio cell is a different quantity: the spread of the paired per-run ratios._ +_Across-run IQR of each side's own median (5 runs): modern-di ≤0.1%, rivals ≤1.6%. The ± on each ratio cell is a different quantity: the spread of the paired per-run ratios._ ## What the numbers show -- Against `dependency-injector`, modern-di is faster by reference on C1 (**0.42**) and - C3 (**0.36**), and far faster on the batched C4 request lifecycle (**0.02**). +- Against `dependency-injector`, modern-di is faster by reference on C1 (**0.41**) and + C3 (**0.35**), and far faster on the batched C4 request lifecycle (**0.02**). dependency-injector's C4 body calls `init_resources()`/`shutdown_resources()` every cycle in addition to resolving; the suite doesn't decompose how much of its per-request cost is that lifecycle work versus the resolve itself, so C4 should be read as a whole-lifecycle comparison, not an isolated resolve (see the caveat below). dependency-injector is still - faster on C2 warm-singleton (2.23, an implied ~68 ns cache hit against modern-di's 151 ns): + faster on C2 warm-singleton (2.20, an implied ~66 ns cache hit against modern-di's 146 ns): its hit is a C-level slot read on a Cython-compiled core, where modern-di's is a Python - dict lookup plus a slot read on the cache item. Pure Python does not reach ~67 ns, so this + dict lookup plus a slot read on the cache item. Pure Python does not reach ~66 ns, so this cell is expected to stay above 1.0 however much of modern-di's own overhead is removed. -- Against `that-depends` (now 4.1.0), modern-di leads by reference on C1 (**0.58**) and C3 - (**0.53**). The C1 series across publications is 1.08, 1.12, 0.98, 0.98, 0.97, 0.89, 0.91, - 0.65, 0.58, now 0.58. that-depends remains faster on C2 warm-singleton (1.78); the suite does not - decompose its `resolve_sync` cache-hit path, so no mechanism is asserted for the remaining gap. +- Against `that-depends` (4.1.0), modern-di leads by reference on C1 (**0.56**) and C3 + (**0.51**). The C1 series across publications is 1.08, 1.12, 0.98, 0.98, 0.97, 0.89, 0.91, + 0.65, 0.58, 0.58, now 0.56. that-depends remains faster on C2 warm-singleton (1.70); the suite + does not decompose its `resolve_sync` cache-hit path, so no mechanism is asserted for the + remaining gap. - Against the two `exec`-codegen frameworks, modern-di leads most of the by-type table. - modern-di is faster than `dishka` on C1 (**0.72**) and C2 (**0.62**), and faster than `wireup` - on C1 (**0.82**) and C3 (**0.87**). dishka keeps its lead on C3 (1.25), the deepest graph, and - wireup keeps C2 (1.47). Since #470 modern-di also generates its resolvers from a source - template, so both sides run one generated frame per node, and the suite does not decompose what dishka does - differently on a six-node chain. No mechanism is asserted for that cell. -- By-type resolution carries no surcharge. `Container.resolve` memoizes type → resolver directly, so - the by-type and by-reference cells differ by 4-6 ns on C1 and C2 and swap sign on C3, all - inside the run-to-run spread. The two tables measure the same resolve; only the rival set + modern-di is faster than `dishka` on C1 (**0.70**) and C2 (**0.60**), and faster than `wireup` + on C1 (**0.80**) and C3 (**0.83**). dishka keeps its lead on C3 (1.21), the deepest graph, and + wireup keeps C2 (1.41). Since #470 modern-di also generates its resolvers from a source + template, so both sides run one generated frame per node, and the suite does not decompose what + dishka does differently on a six-node chain. No mechanism is asserted for that cell. +- By-type resolution carries no surcharge. `Container.resolve` memoizes type → resolver directly, + so the by-type and by-reference cells differ by 3-4 ns, with by-type the faster of the two on C1 + and C2 and the slower on C3. The two tables measure the same resolve; only the rival set differs. -- On C6 (per-request context) modern-di is faster than `dependency-injector` (**0.38**) and - `that-depends` (**0.48**), slower than `dishka` (1.15), and level with `wireup` (1.02 ±2.4%, a - cell whose spread straddles 1.00). The direction - matches the by-type table, but the cells are not on one basis: each framework supplies the - request value through its own idiom, and two of those are structural analogs rather than - equivalents (see the caveat below). No mechanism is asserted for the gaps; the suite does not - decompose any framework's context lookup. +- On C6 (per-request context) modern-di is faster than all four rivals: `dependency-injector` + (**0.28**), `that-depends` (**0.36**), `dishka` (**0.88**) and `wireup` (**0.78**). In the 3.x + publication it was slower than dishka (1.15) and level with wireup (1.02). The cells are not on + one basis: each framework supplies the request value through its own idiom, and two of those + are structural analogs rather than equivalents (see the caveat below). No mechanism is asserted + for the gaps; the suite does not decompose any framework's context lookup. - On C4 (request lifecycle), the corrected batching does not *remove* the ~35 µs asyncio floor, it amortizes it. The guard tier's `test_g7c_event_loop_floor_control` times the same batch shape with an empty body and puts the residual at ~0.3 µs per request still inside every - C4 cell (~13% of modern-di's C4 figure), shared identically by all five frameworks. With the - floor amortized, dishka is measurably faster than modern-di here (1.08); modern-di remains + C4 cell (~14% of modern-di's C4 figure), shared identically by all five frameworks. With the + floor amortized, dishka is measurably faster than modern-di here (1.12); modern-di remains far faster than that-depends, dependency-injector, and wireup on this scenario. ### What moved in this publication -The changes in this publication moved the per-request cells and left the resolve cells alone. -The three changes since 3.5.0 sit on paths C1-C3 never take: #540 inlined the cross-scope hop into the generated resolver, #541 shares one -lock per container tree instead of allocating an `RLock` per child, and #542 stops allocating a -coroutine per finalizer-less cached item in `close_async`. A C1-C3 body resolves inside one -warm container, so every one of those cells is within its own run-to-run spread of the 3.5.0 -value (253 vs 252 ns on C1 by reference, 151 vs 151 on C2, 789 vs 792 on C3, and the same on -the by-type side). C4 and C6 build a child and, on C6, hop from the REQUEST child to an APP -dependency, and both moved. Same machine, same CPython, so the absolutes carry across as well -as the ratios. - -| | 3.5.0 (`18bbe64`) | this publication (`630de77`) | +4.0 moved C6 the most. C6 builds a REQUEST child seeded with a context value, resolves a handler +that needs that value and an APP dependency, and closes the child. It takes #557's context +resolver, #585's cheaper child build and close, and #596's cross-scope check. C4 builds a child, +first-resolves one request-scoped cached factory with an async finalizer and closes the child, so +it gets #585's cheaper child build and pays for #597's lock allocation. + +| | 3.x (`630de77`) | 4.0 (`f300c2e`) | |---|---|---| -| C4 request lifecycle, modern-di | 2.42 µs | 2.29 µs | -| C4 vs that-depends / dishka / wireup | 0.19 / 1.14 / 0.15 | **0.18** / 1.08 / **0.14** | -| C6 context, modern-di | 1.62 µs | 1.45 µs | -| C6 vs dependency-injector / that-depends | 0.42 / 0.53 | **0.38** / **0.48** | -| C6 vs dishka / wireup | 1.28 / 1.16 | 1.15 / 1.02 | - -C6 fell 10% and C4 5%, in line with the guard tier, which measured each change in isolation -against the commit before it: the hop inline is worth −15 to −17% on a cross-scope resolve and -−7.5 to −11% on a context resolve (G9, C6's guard twin); the shared lock is worth −14% on a bare -child build and −3.5% on a request cycle (G7, C4's guard twin); and the coroutine skip is worth -−8% on a batch of request cycles closing ten finalizer-less cached providers (G13b), a shape -neither C4 nor C6 has, since C4's one cached item has an async finalizer and C6's child caches -nothing and closes synchronously. -C4's own cell carries a ±4.8% spread this time, so read its 5% as the sign and the G7 figure -as the size. The wireup C6 cell has crossed from 1.16 to 1.02 ±2.4%, which is level, not a lead. +| C1 / C2 / C3 by reference, modern-di | 253 / 151 / 789 ns | 242 / 146 / 760 ns | +| C4 request lifecycle, modern-di | 2.29 µs | 2.28 µs | +| C4 vs that-depends / dishka / wireup | **0.18** / 1.08 / **0.14** | **0.18** / 1.12 / **0.14** | +| C6 context, modern-di | 1.45 µs | 1.09 µs | +| C6 vs dependency-injector / that-depends | **0.38** / **0.48** | **0.28** / **0.36** | +| C6 vs dishka / wireup | 1.15 / 1.02 | **0.88** / **0.78** | + +C6 fell 25%, close to what the guard tier predicts: G9, C6's guard twin, fell 28% with #557 and another 5.9% +with #585. That moved modern-di ahead of dishka and wireup on this scenario. C4 did not move. On +G7, C4's guard twin, #585 measured −2.9% and #597 +4.6%, so the two roughly cancel. The dishka C4 +ratio went from 1.08 to 1.12 because dishka's own cell is about 3% lower than the 3.x ratio +implies, while modern-di's stayed put. + +The C1-C3 cells fell 3-4% on both the by-reference and by-type sides, while the rival cells +implied by the 3.x ratios stayed within about 2% of this run. Every generated `Factory` resolver +checks its scope first, and #596 changed that check from `==` to `is`, which made a same-scope resolve about 5 ns +faster in that PR's own measurement. That fits the size of the drop (11 ns on C1, 5 ns on C2, +29 ns on C3). No guard run isolated it. The C4 gain recorded at 3.1.1 was a library fix, and it stands: every `Container` used to store itself in its own ancestor map, making it a reference cycle that reference counting could @@ -210,11 +209,10 @@ own way: modern-di seeds a child container's context and resolves by reference; placeholder factory; that-depends supplies it through `container_context(global_context=)`; and dependency-injector injects by reference via `providers.Dependency` + `.override()`, a structural analog rather than an equivalent. modern-di's timed body builds the child, resolves, and closes -it. It calls no `open()`: a freshly built child is already open as of 3.1, so timing one would -charge modern-di a redundant lock acquire (81 ns, ~6% of the cell) with no counterpart in any -rival's body. It does close, because all four rivals exit their scope inside the timed body; that -teardown is ~110 ns, and omitting it would have flattered modern-di by more than the `open()` -would have cost it. +it. It calls no `open()`: a freshly built child is already open as of 3.1, and no rival's body has a +counterpart. It does close, because all four rivals exit their scope inside the timed body. When +C6 was first published, the `open()` was a lock acquire worth 81 ns and the close about 110 ns, so +leaving the close out would have flattered modern-di by more than the `open()` would have cost it. Each framework runs at its default thread-safety configuration, and the defaults differ. dishka's `make_container` defaults to `lock_factory=`, so every `get()` behind its C1-C3 cells @@ -310,23 +308,40 @@ extra Python frame on the hop (−15 to −17% on a cross-scope resolve, −7.5 resolve). #541 gave a container tree one `RLock`, created by the root and shared by every child, where each child used to allocate its own (−14% on a child build, −3.5% on a request cycle). #542 made `close_async` clear a finalizer-less cached item directly rather than await a -coroutine that only did that (−8% on a request cycle closing ten such items). The C4 and C6 -movement above is the sum of the first two; the third has no cell on this page. - -4.0 sped up two resolve paths. The tables above still show the 3.x publication, so these are -guard-tier figures. #557 made a context value an ordinary dependency: a `ContextProvider` compiles -to its own resolver, and a factory that depends on one takes the same generated path as any other -factory. The per-resolve loop that looked up each context value and decided what to pass when one -was missing is gone (−28% on a context resolve, G9, C6's guard twin). #559 compiles an `Alias` to its source's -resolver, so a hop through an alias runs no frame of its own (−33% on an alias hop, G18, which -now matches a plain cached resolve). An error that crosses an alias still shows the alias in its -chain: the parent puts the hop back when it adds its own step. - -4.0 also replaced the tree lock from #541 with one lock per cache item (#569), so creating one -cached factory no longer waits for another. A child build still allocates no lock, because the -lock comes with the cache item. The cost moved to the first resolve of a cached factory in each -container. A request that resolves one request-scoped cached factory pays about 120 ns to -allocate its lock: +7.7% on G7b, one request cycle with a sync close, and +4.6% on G7. +coroutine that only did that (−8% on a request cycle closing ten such items). In the 3.x +publication at `630de77`, C6 fell 10% and C4 5%, the sum of the first two; the third has no cell +on this page. + +Several 4.0 changes moved these numbers. The figures in the next three paragraphs are guard-tier +numbers from the PR that made each change, measured against the commit before it. [What moved in this publication](#what-moved-in-this-publication) reads the comparative +cells that moved. + +#557 made a context value an ordinary dependency: a `ContextProvider` compiles to its own +resolver, and a factory that depends on one takes the same generated path as any other factory. +The per-resolve loop that looked up each context value and decided what to pass when one was +missing is gone (−28% on a context resolve, G9, C6's guard twin). #559 compiles an `Alias` to its +source's resolver, so a hop through an alias runs no frame of its own (−33% on an alias hop, G18, +which now matches a plain cached resolve). An error that crosses an alias still shows the alias in +its chain: the parent puts the hop back when it adds its own step. + +#585 made the request cycle and the cold path cheaper. A child container builds its registries +without a keyword-only dataclass constructor, compiled resolver code is memoized by shape, the +`ContextProvider` resolver reads the context values directly, and closing a container that created +nothing skips the cache teardown. That is −15.6% on a child build (G6, 716 to 605 ns), −13.4% on a +cold first resolve (G8, 24.58 to 21.29 µs), −5.9% on a context resolve (G9) and −2.9% on a batch +of request cycles (G7). #593 rebuilt the resolver templates from shared fragments and left the +generated source byte-identical, so nothing moved. #596 matches scopes by enum member instead of +integer value. A same-scope cached resolve got about 5 ns faster and a cross-scope one about 5 ns +slower, from the extra identity check on the hop. + +#561 removed `Container(use_lock=)`, so every container locks a cache miss. A cache hit never +reaches the lock, and an uncontended acquire costs about 54 ns, about 0.2% of a cold first resolve. +#597 then replaced the tree lock from #541 with one lock per cache item, so creating one cached +factory no longer waits for another, and a miss builds the factory's dependencies under that lock. +A child build still allocates no lock, because the lock comes with the cache item. The cost moved +to the first resolve of a cached factory in each container. A request that resolves one +request-scoped cached factory pays about 120 ns to allocate its lock: +7.7% on G7b, one request +cycle with a sync close, and +4.6% on G7. ## Reproduce it yourself