Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
117 changes: 65 additions & 52 deletions docs/introduction/performance.md
Original file line number Diff line number Diff line change
Expand Up @@ -37,13 +37,13 @@ modern-di variant against the rivals whose API matches it, because a single colu
flatter modern-di against half the set. By-type resolution used to add a fixed lookup cost on
top of `resolve_provider`: 54-65 ns through 3.2.0, then 21/17/23 ns on C1/C2/C3 once 3.3.0 inlined
`resolve_provider`'s body into `resolve`. As of the template resolver `Container.resolve` memoizes
type → resolver directly, and the two columns are within run-to-run noise of each other (252 vs
248 ns on C1, 151 vs 146 on C2, 792 vs 796 on C3 in the cells below).
type → resolver directly, and the two columns are within run-to-run noise of each other (253 vs
247 ns on C1, 151 vs 147 on C2, 789 vs 798 on C3 in the cells below).

C4 does not split this way: modern-di's C4 body resolves **by reference** throughout, while
dishka and wireup can only be measured by type. That asymmetry cuts **against** modern-di's
C4 ratios, not for them, though with the by-type surcharge now inside noise, levelling it would not
move the dishka ratio below (1.14). The C1-C3 leveling does not apply to C4.
move the dishka ratio below (1.08). The C1-C3 leveling does not apply to C4.

C6 does not split either, for the same reason: modern-di's C6 body resolves by reference and there
is no by-type C6 variant to pair against dishka and wireup, so a split would leave that half of the
Expand All @@ -59,15 +59,15 @@ for why that matters and which cells it moved.

## Results

Measured 2026-09-12 with modern-di at `18bbe64` (the template resolver of #470, first released
in 3.5.0) on an Apple M2 (macOS 26.6.2), CPython 3.14.7, median over 5 runs
Measured 2026-09-27 with modern-di `main` at `630de77` (3.5.0 plus #540, #541 and #542,
unreleased when measured) on an Apple M2 (macOS 26.6.2), CPython 3.14.7, median over 5 runs
(ratios paired within each run); the footnote under each table bounds the across-run dispersion
of each side's own median. Rival versions: dishka 1.10.1, dependency-injector 4.49.1,
that-depends 4.1.0, wireup 2.12.0. Generated by `just bench-report`.

This publication is on a **different machine** from the previous one (an M4 with CPython 3.14.6).
Absolute cells are therefore not comparable with the earlier tables; the ratios are, and the
[what moved](#what-moved-in-this-publication) section below reads them that way.
This publication is on the **same machine and CPython build** as the 3.5.0 one, so this time the
absolute cells are comparable with the previous tables as well as the ratios; the
[what moved](#what-moved-in-this-publication) section below reads both.

> **These numbers are a snapshot of the version named above; a newer release does not update
> them.** The tables are regenerated by hand, so they lag a release rather than ship with one. If
Expand All @@ -86,37 +86,37 @@ as a verdict.

| Scenario | modern-di | vs dependency-injector | vs that-depends |
|---|---|---|---|
| C1 transient | 252 ns ±1.0% | **0.42** ±3.4% | **0.58** ±0.2% |
| C2 warm singleton | 151 ns ±3.6% | 2.27 ±5.7% | 1.79 ±2.9% |
| C3 deep chain (6) | 792 ns ±3.8% | **0.37** ±2.9% | **0.53** ±3.1% |
| C1 transient | 253 ns ±1.9% | **0.42** ±2.8% | **0.58** ±2.4% |
| C2 warm singleton | 151 ns ±0.6% | 2.23 ±7.8% | 1.78 ±1.0% |
| C3 deep chain (6) | 789 ns ±1.4% | **0.36** ±5.3% | **0.53** ±0.2% |

_Across-run IQR of each side's own median (5 runs): modern-di ≤3.8%, rivals ≤1.3%. The ± on each ratio cell is a different quantity: the spread of the paired per-run ratios._
_Across-run IQR of each side's own median (5 runs): modern-di ≤1.9%, rivals ≤7.3%. The ± on each ratio cell is a different quantity: the spread of the paired per-run ratios._

### By-type resolution

| Scenario | modern-di | vs dishka | vs wireup |
|---|---|---|---|
| C1 transient | 248 ns ±0.8% | **0.71** ±5.5% | **0.82** ±1.1% |
| C2 warm singleton | 146 ns ±5.3% | **0.61** ±1.2% | 1.45 ±2.3% |
| C3 deep chain (6) | 796 ns ±3.5% | 1.25 ±1.7% | **0.87** ±3.4% |
| C1 transient | 247 ns ±3.4% | **0.72** ±3.1% | **0.82** ±2.4% |
| C2 warm singleton | 147 ns ±0.7% | **0.62** ±0.5% | 1.47 ±1.0% |
| C3 deep chain (6) | 798 ns ±3.2% | 1.25 ±0.7% | **0.87** ±2.9% |

_Across-run IQR of each side's own median (5 runs): modern-di ≤5.3%, rivals ≤5.7%. The ± on each ratio cell is a different quantity: the spread of the paired per-run ratios._
_Across-run IQR of each side's own median (5 runs): modern-di ≤3.4%, rivals ≤1.3%. The ± on each ratio cell is a different quantity: the spread of the paired per-run ratios._

### Request lifecycle (batched, published per request)

| Scenario | modern-di | vs dependency-injector | vs that-depends | vs dishka | vs wireup |
|---|---|---|---|---|---|
| C4 request lifecycle | 2.42 µs ±0.7% | **0.02** ±1.4% | **0.19** ±2.3% | 1.14 ±0.7% | **0.15** ±1.4% |
| C4 request lifecycle | 2.29 µs ±4.8% | **0.02** ±6.9% | **0.18** ±7.9% | 1.08 ±0.3% | **0.14** ±0.9% |

_Across-run IQR of each side's own median (5 runs): modern-di ≤0.7%, rivals ≤2.0%. The ± on each ratio cell is a different quantity: the spread of the paired per-run ratios._
_Across-run IQR of each side's own median (5 runs): modern-di ≤4.8%, rivals ≤7.2%. The ± on each ratio cell is a different quantity: the spread of the paired per-run ratios._

### Per-request context

| Scenario | modern-di | vs dependency-injector | vs that-depends | vs dishka | vs wireup |
|---|---|---|---|---|---|
| C6 context | 1.62 µs ±1.5% | **0.42** ±1.0% | **0.53** ±1.6% | 1.28 ±0.8% | 1.16 ±0.8% |
| C6 context | 1.45 µs ±1.1% | **0.38** ±1.4% | **0.48** ±1.6% | 1.15 ±2.2% | 1.02 ±2.4% |

_Across-run IQR of each side's own median (5 runs): modern-di ≤1.5%, rivals ≤5.3%. The ± on each ratio cell is a different quantity: the spread of the paired per-run ratios._
_Across-run IQR of each side's own median (5 runs): modern-di ≤1.1%, rivals ≤1.9%. The ± on each ratio cell is a different quantity: the spread of the paired per-run ratios._

## What the numbers show

Expand All @@ -126,27 +126,28 @@ _Across-run IQR of each side's own median (5 runs): modern-di ≤1.5%, rivals
in addition to resolving; the suite doesn't decompose how much of its per-request cost is
that lifecycle work versus the resolve itself, so C4 should be read as a whole-lifecycle
comparison, not an isolated resolve (see the caveat below). dependency-injector is still
faster on C2 warm-singleton (2.27, an implied ~67 ns cache hit against modern-di's 151 ns):
faster on C2 warm-singleton (2.23, an implied ~68 ns cache hit against modern-di's 151 ns):
its hit is a C-level slot read on a Cython-compiled core, where modern-di's is a Python
dict lookup plus a slot read on the cache item. Pure Python does not reach ~67 ns, so this
cell is expected to stay above 1.0 however much of modern-di's own overhead is removed.
- Against `that-depends` (now 4.1.0), modern-di leads by reference on C1 (**0.58**) and C3
(**0.53**). The C1 series across publications is 1.08, 1.12, 0.98, 0.98, 0.97, 0.89, 0.91,
0.65, now 0.58. that-depends remains faster on C2 warm-singleton (1.79); the suite does not
0.65, 0.58, now 0.58. that-depends remains faster on C2 warm-singleton (1.78); the suite does not
decompose its `resolve_sync` cache-hit path, so no mechanism is asserted for the remaining gap.
- Against the two `exec`-codegen frameworks, the by-type table is now mostly modern-di's.
modern-di is faster than `dishka` on C1 (**0.71**) and C2 (**0.61**), and faster than `wireup`
modern-di is faster than `dishka` on C1 (**0.72**) and C2 (**0.62**), and faster than `wireup`
on C1 (**0.82**) and C3 (**0.87**). dishka keeps its lead on C3 (1.25), the deepest graph, and
wireup keeps C2 (1.45). Since #470 modern-di also generates its resolvers from a source
wireup keeps C2 (1.47). Since #470 modern-di also generates its resolvers from a source
template, so the frame-count story this page used to tell about dishka's C3 no longer applies:
both sides run one generated frame per node, and the suite does not decompose what dishka does
differently on a six-node chain. No mechanism is asserted for that cell.
- The by-type surcharge is gone. `Container.resolve` memoizes type → resolver directly, so
the by-type and by-reference cells differ by 4-5 ns on C1 and C2 and swap sign on C3, all
the by-type and by-reference cells differ by 4-6 ns on C1 and C2 and swap sign on C3, all
inside the run-to-run spread. The two tables now measure the same resolve; only the rival set
differs.
- On C6 (per-request context) modern-di is faster than `dependency-injector` (**0.42**) and
`that-depends` (**0.53**), and slower than `dishka` (1.28) and `wireup` (1.16). The direction
- On C6 (per-request context) modern-di is faster than `dependency-injector` (**0.38**) and
`that-depends` (**0.48**), slower than `dishka` (1.15), and level with `wireup` (1.02 ±2.4%, a
cell whose spread straddles 1.00). The direction
matches the by-type table, but the cells are **not** on one basis: each framework supplies the
request value through its own idiom, and two of those are structural analogs rather than
equivalents (see the caveat below). No mechanism is asserted for the gaps; the suite does not
Expand All @@ -155,36 +156,38 @@ _Across-run IQR of each side's own median (5 runs): modern-di ≤1.5%, rivals
it amortizes it. The guard tier's `test_g7c_event_loop_floor_control` times the same batch
shape with an empty body and puts the residual at **~0.3 µs per request** still inside every
C4 cell (~13% of modern-di's C4 figure), shared identically by all five frameworks. With the
floor amortized, dishka is measurably **faster** than modern-di here (1.14); modern-di remains
floor amortized, dishka is measurably **faster** than modern-di here (1.08); modern-di remains
far faster than that-depends, dependency-injector, and wireup on this scenario.

### What moved in this publication

**A different machine, so only ratios carry across.** The 3.3.0 tables were measured on an Apple
M4 with CPython 3.14.6; these on an Apple M2 with 3.14.7. The rivals' implied absolutes say how
much slower the M2 is: dependency-injector's C2 hit 59.5 → 66.7 ns, that-depends' 82.6 → 84.6
(and it moved from 4.0.2 to 4.1.0 in between), dishka's 214.8 → 247.7, wireup's 94.6 → 101.0,
so 2-15% depending on the framework. Every ratio below moved by more than that, and all in
modern-di's favour.

| | 3.3.0 (M4) | this publication (M2) |
**The per-request cells, not the resolve cells.** The three changes since 3.5.0 sit on paths
C1-C3 never take: #540 inlined the cross-scope hop into the generated resolver, #541 shares one
lock per container tree instead of allocating an `RLock` per child, and #542 stops allocating a
coroutine per finalizer-less cached item in `close_async`. A C1-C3 body resolves inside one
warm container, so every one of those cells is within its own run-to-run spread of the 3.5.0
value (253 vs 252 ns on C1 by reference, 151 vs 151 on C2, 789 vs 792 on C3, and the same on
the by-type side). C4 and C6 build a child and, on C6, hop from the REQUEST child to an APP
dependency, and both moved. Same machine, same CPython, so the absolutes carry across as well
as the ratios.

| | 3.5.0 (`18bbe64`) | this publication (`630de77`) |
|---|---|---|
| C1 transient vs dependency-injector / that-depends | 0.53 / 0.65 | **0.42** / **0.58** |
| C2 warm singleton vs dependency-injector / that-depends | 2.64 / 1.90 | 2.27 / 1.79 |
| C3 deep chain vs dependency-injector / that-depends | 0.38 / 0.53 | **0.37** / 0.53 |
| C1 transient vs dishka / wireup | 0.91 / 1.01 | **0.71** / **0.82** |
| C2 warm singleton vs dishka / wireup | 0.81 / 1.84 | **0.61** / 1.45 |
| C3 deep chain vs dishka / wireup | 1.30 / 0.90 | 1.25 / **0.87** |
| C4 request lifecycle vs dishka | 1.22 | 1.14 |
| C6 context vs dishka / wireup | 1.47 / 1.23 | 1.28 / 1.16 |

The by-type column moved most, and by construction: the by-type surcharge that 3.3.0 had cut
to ~20 ns is now zero, so every dishka and wireup cell gains that on top of the resolver change
itself. On the guard tier, measured on this machine against the commit before #470, the
resolver change alone is worth −10% on a transient resolve, −11% on a warm cached hit, −19% on a
wide ten-dependency node, −16% on a context resolve and −23% on a by-type resolve; the one cost
is the first resolve of a cold container, +12%, because the generated source is compiled once per
resolver shape.
| C4 request lifecycle, modern-di | 2.42 µs | 2.29 µs |
| C4 vs that-depends / dishka / wireup | 0.19 / 1.14 / 0.15 | **0.18** / 1.08 / **0.14** |
| C6 context, modern-di | 1.62 µs | 1.45 µs |
| C6 vs dependency-injector / that-depends | 0.42 / 0.53 | **0.38** / **0.48** |
| C6 vs dishka / wireup | 1.28 / 1.16 | 1.15 / 1.02 |

C6 fell 10% and C4 5%, in line with the guard tier, which measured each change in isolation
against the commit before it: the hop inline is worth −15 to −17% on a cross-scope resolve and
−7.5 to −11% on a context resolve (G9, C6's guard twin); the shared lock is worth −14% on a bare
child build and −3.5% on a request cycle (G7, C4's guard twin); and the coroutine skip is worth
−8% on a batch of request cycles closing ten finalizer-less cached providers (G13b), a shape
neither C4 nor C6 has, since C4's one cached item has an async finalizer and C6's child caches
nothing and closes synchronously.
C4's own cell carries a ±4.8% spread this time, so read its 5% as the sign and the G7 figure
as the size. The wireup C6 cell has crossed from 1.16 to 1.02 ±2.4%, which is level, not a lead.

**The C4 gain recorded at 3.1.1 was a library fix**, and it stands: every `Container` used to
store itself in its own `_scope_map`, making it a reference cycle that reference counting could
Expand Down Expand Up @@ -301,6 +304,16 @@ object shared across providers turns its call sites polymorphic for the speciali
per shape shipped anyway, because `compile()` costs ~70 µs per provider and a test suite building
a container per test would pay it per test.

After 3.5.0, three smaller changes moved the per-request path rather than the resolve. #540
inlined the cross-scope hop: the generated resolver reads the container's ancestor map itself
instead of calling `find_container`, so a REQUEST resolve that needs an APP dependency spends no
extra Python frame on the hop (−15 to −17% on a cross-scope resolve, −7.5 to −11% on a context
resolve). #541 gave a container tree one `RLock`, created by the root and shared by every child,
where each child used to allocate its own (−14% on a child build, −3.5% on a request cycle).
#542 made `close_async` clear a finalizer-less cached item directly rather than await a
coroutine that only did that (−8% on a request cycle closing ten such items). The C4 and C6
movement above is the sum of the first two; the third has no cell on this page.

## Reproduce it yourself

```bash
Expand Down
Loading