diff --git a/README.md b/README.md
index e8ad3c3..a0aefd2 100644
--- a/README.md
+++ b/README.md
@@ -18,11 +18,11 @@
[](https://github.com/astral-sh/ruff)
[](https://github.com/astral-sh/ty)
-`faststream-outbox` is a [FastStream](https://faststream.ag2.ai) broker integration for the **transactional outbox pattern** — a Postgres table is the message queue.
+`faststream-outbox` is a [FastStream](https://faststream.ag2.ai) broker integration for the transactional outbox pattern, with a Postgres table as the message queue.
-A producer writes a domain entity and an outbox row in the *same* SQLAlchemy transaction by calling `broker.publish(body, queue=..., session=session)`. A separate subscriber polls the table and relays each row to a real message bus (Kafka, RabbitMQ, NATS, Redis…) with a single decorator — or processes the rows in-place if you don't have a downstream broker.
+A producer writes a domain entity and an outbox row in the *same* SQLAlchemy transaction by calling `broker.publish(body, queue=..., session=session)`. A separate subscriber polls the table and relays each row to a real message bus (Kafka, RabbitMQ, NATS, Redis…) with a single decorator, or processes the rows in place if you don't have a downstream broker.
-## Quickstart — outbox relay to Kafka
+## Quickstart: outbox relay to Kafka
Write the outbox row in your domain transaction; relay rows to Kafka with a stacked decorator.
@@ -65,9 +65,9 @@ async with session_factory() as session, session.begin():
The same one-decorator pattern works for RabbitMQ, NATS, Redis, and Confluent. See the [relay tutorial](https://faststream-outbox.modern-python.org/usage/relay/) for the FastAPI lifecycle, header propagation, router shapes, and the at-least-once contract.
-## Quickstart — standalone outbox queue
+## Quickstart: standalone outbox queue
-If you don't have a downstream broker, the same broker can process outbox rows in-place — the table *is* the queue.
+If you don't have a downstream broker, the same broker can process outbox rows in place, and the table *is* the queue.
```python
from sqlalchemy import MetaData
@@ -97,9 +97,9 @@ async with session_factory() as session, session.begin():
## How it works
-A subscriber owns two async loops: a **fetch** loop claims available rows via a single CTE (`SELECT … FOR UPDATE SKIP LOCKED → UPDATE acquired_token=:uuid, acquired_at=now() RETURNING *`), and `max_workers` **worker** loops dispatch to the handler. On success, `DELETE WHERE id=:id AND acquired_token=:token`; on failure, the retry strategy schedules another attempt or terminally drops the row. Terminal failures `DELETE` by default; pass `dlq_table=make_dlq_table(metadata)` to atomically archive them into a sibling audit table instead — see [Dead-letter queue](https://faststream-outbox.modern-python.org/usage/dlq/).
+A subscriber owns two async loops: a fetch loop claims available rows via a single CTE (`SELECT … FOR UPDATE SKIP LOCKED → UPDATE acquired_token=:uuid, acquired_at=now() RETURNING *`), and `max_workers` worker loops dispatch to the handler. On success, `DELETE WHERE id=:id AND acquired_token=:token`; on failure, the retry strategy schedules another attempt or terminally drops the row. Terminal failures `DELETE` by default; pass `dlq_table=make_dlq_table(metadata)` to atomically archive them into a sibling audit table instead (see [Dead-letter queue](https://faststream-outbox.modern-python.org/usage/dlq/)).
-The `acquired_token` is the load-bearing invariant: a slow handler whose lease expired and was re-claimed by another worker finds its terminal `DELETE` to be a no-op (the token no longer matches), preventing it from clobbering the new lease holder.
+The `acquired_token` is the load-bearing invariant. If a slow handler's lease expired and another worker re-claimed the row, the slow handler's terminal `DELETE` is a no-op because the token no longer matches, so it cannot clobber the new lease holder.
With the `asyncpg` driver, the fetch loop also `LISTEN`s on `outbox_
` and `publish` emits `pg_notify(...)`, so idle dispatch latency is ~10ms instead of up to `max_fetch_interval`.
@@ -107,15 +107,15 @@ See [How it works](https://faststream-outbox.modern-python.org/introduction/how-
## Optional extras
-- `faststream-outbox[asyncpg]` — asyncpg driver (enables `LISTEN/NOTIFY` for ~10ms idle dispatch)
-- `faststream-outbox[fastapi]` — FastAPI integration via `faststream_outbox.fastapi.OutboxRouter`
-- `faststream-outbox[validate]` — Alembic for `broker.validate_schema()`
-- `faststream-outbox[prometheus]` — Prometheus metrics adapter
-- `faststream-outbox[opentelemetry]` — OpenTelemetry metrics adapter
+- `faststream-outbox[asyncpg]`: asyncpg driver (enables `LISTEN/NOTIFY` for ~10ms idle dispatch)
+- `faststream-outbox[fastapi]`: FastAPI integration via `faststream_outbox.fastapi.OutboxRouter`
+- `faststream-outbox[validate]`: Alembic for `broker.validate_schema()`
+- `faststream-outbox[prometheus]`: Prometheus metrics adapter
+- `faststream-outbox[opentelemetry]`: OpenTelemetry metrics adapter
## Acknowledgements
-The architecture of this package is heavily informed by Arseniy Popov's [PR #2704](https://github.com/ag2ai/faststream/pull/2704) (`feat: add sqla broker`) on upstream FastStream — the FastStream broker/registrator/subscriber wiring, the `SELECT … FOR UPDATE SKIP LOCKED` fetch-and-claim CTE, the retry strategy hierarchy, and the in-transaction publish contract all originate from there. This package is a Postgres-only reimplementation that diverges in storage model (lease tokens instead of an explicit state column, archive table is opt-in), loop structure (two loops instead of four), wake-up mechanism (`LISTEN/NOTIFY`), and adds timer mechanics. Credit for the original design belongs to Arseniy.
+The architecture of this package is heavily informed by Arseniy Popov's [PR #2704](https://github.com/ag2ai/faststream/pull/2704) (`feat: add sqla broker`) on upstream FastStream. The FastStream broker/registrator/subscriber wiring, the `SELECT … FOR UPDATE SKIP LOCKED` fetch-and-claim CTE, the retry strategy hierarchy, and the in-transaction publish contract all originate from there. This package is a Postgres-only reimplementation. It differs in storage model (lease tokens in place of an explicit state column, and an opt-in archive table), loop structure (two loops where the PR has four), and wake-up mechanism (`LISTEN/NOTIFY`), and it adds timer mechanics. Credit for the original design belongs to Arseniy.
## 📚 [Documentation](https://faststream-outbox.modern-python.org)
@@ -126,4 +126,4 @@ The architecture of this package is heavily informed by Arseniy Popov's [PR #270
## Part of `modern-python`
Browse the full list of templates and libraries in
-[`modern-python`](https://github.com/modern-python) — see the org profile for the categorized index.
+[`modern-python`](https://github.com/modern-python); the org profile has the categorized index.
diff --git a/docs/assets/social-card.png b/docs/assets/social-card.png
index ed87a5d..700edf3 100644
Binary files a/docs/assets/social-card.png and b/docs/assets/social-card.png differ
diff --git a/docs/concepts/comparison.md b/docs/concepts/comparison.md
index 8af77dc..116f3f2 100644
--- a/docs/concepts/comparison.md
+++ b/docs/concepts/comparison.md
@@ -16,45 +16,44 @@ reader can lift the answer without reading the discussion.
## vs. writing your own outbox table and worker
-A bespoke outbox is the most common starting point — the pattern itself is
+A bespoke outbox is the most common starting point. The pattern itself is
straightforward, and an MVP can ship in an afternoon. What `faststream-outbox`
-buys you, in concrete terms, is the pile of pieces that turn that MVP into a
-production system:
+gives you is the pile of pieces that turn that MVP into a production system:
-- per-row **lease tokens** with a load-bearing invariant that any new
+- per-row lease tokens with a load-bearing invariant that any new
fetch / terminal path must preserve;
-- the **partial-index design** the fetch CTE depends on (without those, the
+- the partial-index design the fetch CTE depends on (without those, the
disjunctive WHERE clause falls back to seq-scan as the table grows);
-- the **fetch-and-claim CTE shape** with `FOR UPDATE SKIP LOCKED` that
+- the fetch-and-claim CTE shape with `FOR UPDATE SKIP LOCKED` that
reclaims both unleased rows and expired leases without a separate reaper;
-- the **retry-strategy hierarchy** (`ExponentialRetry`, `NoRetry`, …)
+- the retry-strategy hierarchy (`ExponentialRetry`, `NoRetry`, …)
enforcing `max_attempts` and `max_total_delay_seconds` uniformly;
-- **`validate_schema()`** via Alembic's `autogenerate.compare_metadata`;
-- **drain semantics** on stop: new fetches stop while in-flight handlers
+- `validate_schema()` via Alembic's `autogenerate.compare_metadata`;
+- drain semantics on stop: new fetches stop while in-flight handlers
finish, and subscribers drain concurrently;
-- **`LISTEN/NOTIFY`** short-circuit on top of polling, with NOTIFY
+- a `LISTEN/NOTIFY` short-circuit on top of polling, with NOTIFY
suppression on future-dated rows and `timer_id` conflict no-ops;
-- **`timer_id` dedup** via a partial unique index plus
+- `timer_id` dedup via a partial unique index plus
`on_conflict_do_nothing`;
-- the **DLQ atomicity CTE** that rolls back the DELETE when the DLQ insert
+- the DLQ atomicity CTE that rolls back the DELETE when the DLQ insert
fails.
None of these is hard individually. Together they make up most of the work
past an MVP.
-You also pick up the [Subscriber](../usage/subscriber.md),
+You also get the [Subscriber](../usage/subscriber.md),
[Publisher](../usage/publisher.md), [Dead-letter queue](../usage/dlq.md), and
-[Observability](../usage/observability.md) reference pages — written to
+[Observability](../usage/observability.md) reference pages, written to
the level you would otherwise have to write yourself.
-**TL;DR.** Build it yourself if you only ever need the MVP shape. Use
+Build it yourself if you only ever need the MVP shape. Use
`faststream-outbox` if you expect the system to live for a couple of years.
## vs. CDC (Debezium, logical replication)
Change-data capture sits one layer below outbox: instead of writing rows
to an outbox table, you read your write-ahead log directly. The producer
-code is unchanged — any write to the underlying tables becomes an event.
+code is unchanged, and any write to the underlying tables becomes an event.
Debezium and similar tools have spent the last decade hardening this path
for Postgres, MySQL, and others; the operator playbook is well-known.
@@ -62,29 +61,29 @@ CDC wins when you already need WAL-level capture for analytics or
reverse-ETL anyway, when you want to capture writes from services you do
not control (i.e. not all your producers are FastStream apps), or when
the polling overhead of an outbox is unacceptable. CDC also wins when
-the events you care about are derivable from row state — "an order
-exists with status='paid'" rather than "an
+the events you care about are derivable from row state, such as "an order
+exists with status='paid'" as opposed to "an
`OrderPaid` event was published."
`faststream-outbox` wins when you control the producer code (so the
outbox row is cheap to write inline with the domain write), when you
-need **handler-level retry, DLQ, and scheduled-delivery semantics
-inline** (CDC pushes those concerns to a separate consumer layer), and
-when the **async-Python logical-replication tooling gap** is too thin
-to lean on: there is no async-native logical-decoding client comparable
+need handler-level retry, DLQ, and scheduled-delivery semantics
+inline (CDC pushes those concerns to a separate consumer layer), and
+when the async-Python logical-replication tooling is too thin
+to lean on. There is no async-native logical-decoding client comparable
to Debezium's JVM connectors, so a Python CDC path means either running
the JVM stack alongside your app or driving `pg_recvlogical` / a thin
`psycopg` replication-protocol wrapper yourself. That gap is why this
project does not recommend CDC as the default path.
-**TL;DR.** Pick CDC when you already need WAL capture or have producers
+Pick CDC when you already need WAL capture or have producers
outside your control. Pick this when you own the producer and want
retry/DLQ/timers in-process.
## vs. Kafka transactions (or RabbitMQ publisher confirms)
-Atomic `DB-write + bus-publish` is also achievable on a real bus, just
-not for free. Kafka transactions plus two-phase commit, or an
+Atomic `DB-write + bus-publish` is also achievable on a real bus, at a
+cost. Kafka transactions plus two-phase commit, or an
idempotent-producer pattern combined with an inbox table on the consumer
side, can give you the same end-to-end at-least-once guarantee without a
DB-backed outbox.
@@ -100,12 +99,12 @@ to preserve, because the transactional boundary belongs to two different
systems.
`faststream-outbox` covers a focused subset of a real bus's delivery
-surface — one-process producer, one Postgres table — while adding
+surface (one-process producer, one Postgres table) while adding
outbox-native features a bare bus lacks (`cancel_timer`, `timer_id`
producer-side dedup, scheduled delivery), at the price of being
Postgres-only and polling-based.
-**TL;DR.** Kafka transactions / Rabbit confirms win at scale where the
+Kafka transactions / Rabbit confirms win at scale where the
bus is already running. `faststream-outbox` wins when Postgres is your
only durable store.
@@ -117,21 +116,21 @@ across listener disconnect: a NOTIFY emitted while the listener's
connection is dead, or during a reconnect, is silently dropped. There is
no replay, no persistence, no retry.
-`faststream-outbox` keeps the **outbox row** as the durability boundary
+`faststream-outbox` keeps the outbox row as the durability boundary
and uses NOTIFY only as a wake-up short-circuit on top of polling. If
-the NOTIFY is lost — listener reconnecting, or `LISTEN` setup failed at
-startup — the subscriber still finds the row on its next poll cycle. The worst case
+the NOTIFY is lost (listener reconnecting, or `LISTEN` setup failed at
+startup), the subscriber still finds the row on its next poll cycle. The worst case
is one `max_fetch_interval` of idle latency (default 10 seconds), not
data loss.
-**TL;DR.** Raw `LISTEN/NOTIFY` is a wake-up, not a delivery guarantee.
+Raw `LISTEN/NOTIFY` is a wake-up, not a delivery guarantee.
Use the outbox row for durability and let NOTIFY shave idle latency.
## vs. Celery (or RQ, Dramatiq) with a DB backend
-Celery and friends are *task queues* — you submit "go do this thing" and
+Celery and friends are *task queues*: you submit "go do this thing" and
a worker picks it up later. `faststream-outbox` is *message routing* with
-FastStream's subscriber/publisher semantics — the row is an event tied
+FastStream's subscriber/publisher semantics: the row is an event tied
to a domain write, and the handler is the consumer for events on that
queue.
@@ -139,34 +138,32 @@ The two abstractions overlap, but the right one depends on what you are
modelling. Celery wins for ad-hoc background jobs initiated from
arbitrary points in your app (request handlers, admin commands, cron),
where the relationship to a database transaction is incidental. Use
-`faststream-outbox` when you want **at-least-once dispatch of events
-that must commit atomically with a domain write**, and prefer
+`faststream-outbox` when you want at-least-once dispatch of events
+that must commit atomically with a domain write, and prefer
FastStream's `@broker.subscriber` model over Celery's task decorator.
-The two can also coexist — Celery for fire-and-forget background jobs,
-`faststream-outbox` for the transactional event tier. They are not in
-direct competition for the same problem.
+The two can also coexist: Celery for fire-and-forget background jobs,
+`faststream-outbox` for the transactional event tier.
-**TL;DR.** Celery for ad-hoc background jobs. `faststream-outbox` for
+Celery for ad-hoc background jobs. `faststream-outbox` for
events tied to DB transactions.
## vs. FastStream + `KafkaBroker` / `RabbitBroker` directly
If you have no domain write to atomically commit alongside the bus
-publish, drop the outbox entirely — use the foreign broker directly via
+publish, drop the outbox entirely and use the foreign broker directly via
FastStream's native `KafkaBroker`, `RabbitBroker`, `NatsBroker`, etc.
You skip the polling overhead and the Postgres dependency; you keep the
same `@broker.subscriber` ergonomics.
-The interesting case is **both at once**: domain code writes to Postgres
+The interesting case is both at once: domain code writes to Postgres
*and* needs the event to reach Kafka. That is the canonical
transactional-outbox shape, and it composes the two: the outbox row
captures the event in the domain transaction; a
[Relay](../usage/relay.md) subscriber forwards it to Kafka with the
-at-least-once contract preserved end to end. Don't pick between
-`faststream-outbox` and a real bus — use both, with the outbox as the
-durability boundary in front of the bus.
+at-least-once contract preserved end to end. Use both, with the outbox
+as the durability boundary in front of the bus.
-**TL;DR.** No DB write to commit with? Use the foreign broker directly.
+No DB write to commit with? Use the foreign broker directly.
Need atomicity with a DB write? Use this *plus* the foreign broker via
[Relay](../usage/relay.md).
diff --git a/docs/concepts/instrumentation-seams.md b/docs/concepts/instrumentation-seams.md
index 01f0f25..ad1c56e 100644
--- a/docs/concepts/instrumentation-seams.md
+++ b/docs/concepts/instrumentation-seams.md
@@ -1,7 +1,7 @@
# Instrumentation seams
-`faststream-outbox` exposes **two complementary instrumentation seams** —
-a *recorder* (callable) and a *native middleware* — and recommends
+`faststream-outbox` exposes two complementary instrumentation seams,
+a *recorder* (callable) and a *native middleware*, and recommends
running both. This page explains why two; the practical setup recipes
live in [Setup Prometheus and OpenTelemetry](../usage/setup-prometheus-opentelemetry.md),
and the event catalog and PromQL playbook in
@@ -11,15 +11,15 @@ and the event catalog and PromQL playbook in
A FastStream broker emits two natural observation moments:
-- `consume_scope` — wraps a single handler invocation. The middleware
+- `consume_scope` wraps a single handler invocation. The middleware
bus surfaces handler duration, message size, exception status, span
context.
-- `publish_scope` — wraps a single producer call. Same idea on the
+- `publish_scope` wraps a single producer call. Same idea on the
outbound side.
Upstream FastStream middlewares (`TelemetryMiddleware`,
`PrometheusMiddleware`) hook into these two scopes. For Kafka, Rabbit,
-NATS, that's the entire surface area — those buses don't have
+NATS, that's the entire surface area: those buses don't have
outbox-internal events because they don't *have* an outbox.
`faststream-outbox` does have outbox-internal events, and the middleware
@@ -39,25 +39,25 @@ This is the "spans + bus parity" mode the native middleware
## What the middleware seam *can't* observe
-Four events fire **outside** the handler invocation, with no
+Four events fire outside the handler invocation, with no
`StreamMessage` in scope:
-- **`fetched` ticks (including empty fetches).** Emitted by the fetch
- loop every time it claims rows from the table, *before* any handler
- runs. The middleware bus has no `consume_scope` to wrap yet — there
- is no message. Empty-fetch ticks are also load-bearing for
- detecting "polling but the queue is empty" patterns; the middleware
- bus never sees them.
-- **`lease_lost` events.** Fired after `consume_scope` has already
- closed (the handler returned successfully but its terminal `DELETE`
- matched zero rows because the lease expired). By the time we know
- the row was lost, the middleware has long since recorded a normal
- `acked`. The recorder catches the truth.
-- **`nacked_terminal(reason="max_deliveries")`.** This row exceeded
+- The fetch loop emits `fetched` ticks (including empty fetches) every
+ time it claims rows from the table, *before* any handler runs. The
+ middleware bus has no `consume_scope` to wrap yet because there is
+ no message. Empty-fetch ticks are also load-bearing for detecting
+ "polling but the queue is empty" patterns; the middleware bus never
+ sees them.
+- `lease_lost` fires after `consume_scope` has already closed (the
+ handler returned successfully but its terminal `DELETE` matched zero
+ rows because the lease expired). By the time we know the row was
+ lost, the middleware has long since recorded a normal `acked`. The
+ recorder catches the truth.
+- `nacked_terminal(reason="max_deliveries")` means the row exceeded
the `max_deliveries` ceiling and was dropped *without invoking the
- handler*. No handler call = no `consume_scope`. The middleware has
- nothing to wrap.
-- **`drain_timeout`.** Fired during `stop()` when the shutdown drain
+ handler*. No handler call means no `consume_scope`, so the middleware
+ has nothing to wrap.
+- `drain_timeout` fires during `stop()` when the shutdown drain
overruns `graceful_timeout`. There is no handler scope at all.
## What the recorder seam observes naturally
@@ -77,14 +77,14 @@ conditional `dlq_written` when the DLQ is configured, and one producer event
The recorder cannot bracket span lifecycles (it's a callable, not a
context manager), so tracing belongs to the middleware seam. It also
-runs **on the dispatch event loop and must not block** — a synchronous
+runs on the dispatch event loop and must not block. A synchronous
`Counter.inc()` is fine; an HTTP / StatsD push is not. See
[Observability § Recorder must not block](../usage/observability.md#recorder-must-not-block)
for the full contract.
## Layering: middleware seam vs. recorder seam
-Both can be registered together — each fires for events the other
+Both can be registered together, and each fires for events the other
physically cannot observe.
| Concern | Middleware seam | Recorder seam |
@@ -99,10 +99,10 @@ physically cannot observe.
## Operator implication
-**Run both.** Middleware for bus-scope metrics, distributed tracing,
-and label parity with the rest of your FastStream services. Recorder
-for the outbox-internal events that don't have a `StreamMessage` to
-attach to.
+Run both. Use the middleware for bus-scope metrics, distributed tracing,
+and label parity with the rest of your FastStream services, and the
+recorder for the outbox-internal events that don't have a `StreamMessage`
+to attach to.
The "Both seams together" recipe in [Setup Prometheus and OpenTelemetry
](../usage/setup-prometheus-opentelemetry.md#both-seams-together)
diff --git a/docs/concepts/performance.md b/docs/concepts/performance.md
index f374b03..fb17faf 100644
--- a/docs/concepts/performance.md
+++ b/docs/concepts/performance.md
@@ -1,8 +1,8 @@
# Performance
The outbox transport is Postgres rows, so its performance is governed by two
-things: **round-trips** (every publish and every terminal completion is a
-statement against the database) and **table churn** (every message is one
+things: round-trips (every publish and every terminal completion is a
+statement against the database) and table churn (every message is one
`INSERT`, one lease `UPDATE`, and one terminal `DELETE`, so dead tuples pile up at
roughly twice the message rate and autovacuum has to reclaim them). Three levers
address those. Two are automatic or opt-in code knobs; one is table hygiene you
@@ -15,20 +15,20 @@ detail page from the [three levers](#the-three-levers) below.
## What each lever buys
The left column is what the [benchmark harness](#measuring-it-yourself) locks in
-CI — deterministic statement counts, identical on any machine. The right column
+CI: deterministic statement counts, identical on any machine. The right column
is illustrative wall-clock from a one-off run on loopback Postgres: real, but
-**machine-dependent**, so treat it as a shape, not a promise.
+machine-dependent, so treat it as a shape, not a promise.
| Lever | CI-gated (deterministic) | Illustrative (loopback, one-off) |
|---|---|---|
-| Producer NOTIFY dedup | `pg_notify` statements per bulk publish: **5000 → 1** | ~1.9× bulk-publish throughput; grows with DB round-trip latency |
-| Batched terminal flush | terminal `DELETE`s: **5000 → 50** (100×) at `terminal_flush_batch_size=100` | one batched worker out-throughputs a four-worker per-row subscriber |
-| Autovacuum tuning | *(no benchmark — time-and-throughput dependent)* | churn demo: throughput knob bounds bloat under sustained load |
+| Producer NOTIFY dedup | `pg_notify` statements per bulk publish: 5000 → 1 | ~1.9× bulk-publish throughput; grows with DB round-trip latency |
+| Batched terminal flush | terminal `DELETE`s: 5000 → 50 (100×) at `terminal_flush_batch_size=100` | one batched worker out-throughputs a four-worker per-row subscriber |
+| Autovacuum tuning | *(no benchmark; time-and-throughput dependent)* | churn demo: throughput knob bounds bloat under sustained load |
-**Gated vs illustrative.** The gated counts come from `benchmarks/baseline.json`
+The gated counts come from `benchmarks/baseline.json`
and fail CI if they regress, so they are guarantees about *statement count*, not
speed. The wall-clock figures were measured once and depend on your CPU, disk, and
-especially network latency to Postgres — reproduce them on your own hardware
+especially network latency to Postgres; reproduce them on your own hardware
before quoting them. Autovacuum has no gated number at all: its benefit is
time-and-throughput dependent and cannot be captured as a deterministic count (see
[Autovacuum tuning](../operations/alembic.md#autovacuum-tuning-recommended) for
@@ -38,11 +38,11 @@ why).
| Your workload | Reach for | Why |
|---|---|---|
-| Many rows to one queue in one transaction (bulk / `publish_batch`) | Nothing — **NOTIFY dedup is automatic** | The redundant `pg_notify` round-trips are already gone; delivery is unchanged. |
-| High terminal throughput, **idempotent** handlers, few workers | `terminal_flush_batch_size` > 1 | Collapses per-message `DELETE`s into one per batch; reaches high throughput without spending more workers' connections. |
-| Exactly-once-sensitive or low-volume queue | Leave `terminal_flush_batch_size=1` | Batching widens the crash-redelivery window — not worth it here. |
+| Many rows to one queue in one transaction (bulk / `publish_batch`) | Nothing (NOTIFY dedup is automatic) | The redundant `pg_notify` round-trips are already gone; delivery is unchanged. |
+| High terminal throughput, idempotent handlers, few workers | `terminal_flush_batch_size` > 1 | Collapses per-message `DELETE`s into one per batch; reaches high throughput without spending more workers' connections. |
+| Exactly-once-sensitive or low-volume queue | Leave `terminal_flush_batch_size=1` | Batching widens the crash-redelivery window, which is not worth it here. |
| Table bloating / vacuum can't keep up under sustained churn | `outbox_autovacuum_ddl(...)`, and raise the throughput knob | Eligibility settings fire vacuum sooner; `vacuum_cost_delay` lets it keep pace. |
-| Need more parallel handler capacity | `max_workers` / `fetch_batch_size` | More concurrency — but mind the [connection budget](../usage/subscriber.md#connection-budget). |
+| Need more parallel handler capacity | `max_workers` / `fetch_batch_size` | More concurrency, but mind the [connection budget](../usage/subscriber.md#connection-budget). |
Single-publish-per-transaction workloads see no change from NOTIFY dedup (one row
still emits one NOTIFY), and the batched-flush win is largest at low `max_workers`
@@ -52,21 +52,21 @@ still emits one NOTIFY), and the batched-flush win is largest at low `max_worker
### Producer: NOTIFY dedup (automatic)
-`broker.publish` / `publish_batch` emit **one `pg_notify` per (transaction,
-queue)**, so N publishes to the same queue in one transaction cost one NOTIFY
+`broker.publish` / `publish_batch` emit one `pg_notify` per (transaction,
+queue), so N publishes to the same queue in one transaction cost one NOTIFY
round-trip. Postgres already coalesces identical notifications per transaction
at delivery, so extra NOTIFYs would be pure waste. It is default-on, has no knob, and is
behavior-preserving: the subscriber still gets its wake, fired inline at the first
-publish. You do not configure this — it just makes bulk publishing cheaper. See
+publish. See
[How it works](../introduction/how-it-works.md) for the write path.
### Subscriber: batched terminal flush (opt-in)
-By default each processed row is deleted with its own `DELETE` — one round-trip
+By default each processed row is deleted with its own `DELETE`, one round-trip
per message, which is the throughput ceiling at low `max_workers`. Set
`terminal_flush_batch_size` above `1` to coalesce completed rows into one
-`DELETE … RETURNING` per batch. The tradeoff is a **wider crash-redelivery
-window**: on an ungraceful crash, up to a full batch of already-handled rows are
+`DELETE … RETURNING` per batch. The tradeoff is a wider crash-redelivery
+window: on an ungraceful crash, up to a full batch of already-handled rows are
redelivered, so handlers must be idempotent. Enable it per subscriber for
high-throughput idempotent queues; leave it off otherwise. Full mechanics,
lease-ceiling sizing, and backlog-depth effects are in
@@ -76,12 +76,12 @@ lease-ceiling sizing, and backlog-depth effects are in
A high-churn queue table defeats Postgres' default autovacuum: the
`scale_factor = 0.2` bar is size-dependent and, on a table whose `reltuples`
-estimate has gone stale, fires rarely — the classic queue-table death-spiral.
+estimate has gone stale, fires rarely. This is the classic queue-table death-spiral.
`outbox_autovacuum_ddl("outbox")` renders the `ALTER TABLE … SET (autovacuum_*)`
statement to drop into a migration; it sets `scale_factor = 0` with a constant
-threshold (**eligibility** — fire on a fixed dead-tuple count, size-independent).
-Under heavy sustained churn the binding constraint is instead vacuum **throughput**
-— pass `vacuum_cost_delay` / `vacuum_cost_limit` to let vacuum keep pace. A
+threshold (eligibility: fire on a fixed dead-tuple count, size-independent).
+Under heavy sustained churn the binding constraint is instead vacuum throughput;
+pass `vacuum_cost_delay` / `vacuum_cost_limit` to let vacuum keep pace. A
`validate_schema(check_autovacuum=True)` probe can gate that the eligibility
settings are applied. Full guidance, the eligibility-vs-throughput split, and the
I/O caution on `vacuum_cost_delay=0` are in
@@ -92,13 +92,13 @@ I/O caution on `vacuum_cost_delay=0` are in
The gated numbers above come from the repository's benchmark harness. Contributors
can run it against a local Postgres:
-- `just bench` — runs the producer/consumer workload sweep and prints per-message
+- `just bench` runs the producer/consumer workload sweep and prints per-message
DB counters (split by leading SQL keyword).
-- `just bench-check` — gates the deterministic counts against
+- `just bench-check` gates the deterministic counts against
`benchmarks/baseline.json`; this is what CI runs.
The statement-count reductions (`select_calls`, `delete_calls`, `insert_calls`,
-…) are machine-independent and exact. Wall-clock throughput is not — it scales
+…) are machine-independent and exact. Wall-clock throughput is not: it scales
with your hardware and, most of all, round-trip latency to the database, so a
networked Postgres will show a larger absolute win from the round-trip reductions
than the loopback figures here.
diff --git a/docs/index.md b/docs/index.md
index b7275e5..94cf0d2 100644
--- a/docs/index.md
+++ b/docs/index.md
@@ -8,7 +8,7 @@
`faststream-outbox` is a [FastStream](https://faststream.ag2.ai) broker
-integration for the **transactional outbox pattern** — a Postgres table is
+integration for the transactional outbox pattern, with a Postgres table as
the message queue. A producer writes a domain entity and an outbox row in
the *same* SQLAlchemy transaction; a subscriber polls the table with
`FOR UPDATE SKIP LOCKED`, runs the handler, and deletes the row on
@@ -37,20 +37,20 @@ There are two ways to use it:
## Reach for something else when
- You're already running Kafka / Rabbit / NATS *and* don't need
- transactional atomicity with a DB write → use that broker directly.
-- You need sub-second scheduled-delivery precision → see
+ transactional atomicity with a DB write: use that broker directly.
+- You need sub-second scheduled-delivery precision: see
[Timers § latency floor](usage/timers.md#latency-floor).
-- You're on a non-Postgres database → this package is Postgres-only
+- You're on a non-Postgres database: this package is Postgres-only
at v0. CDC / Debezium may be a better fit (see
[Comparison](concepts/comparison.md)).
-- You're modelling ad-hoc background jobs rather than events tied to a
- DB transaction → see [Comparison § vs Celery](concepts/comparison.md#vs-celery-or-rq-dramatiq-with-a-db-backend).
+- You're modelling ad-hoc background jobs, not events tied to a
+ DB transaction: see [Comparison § vs Celery](concepts/comparison.md#vs-celery-or-rq-dramatiq-with-a-db-backend).
## Start where you're going
| If you want to… | Start at |
|---|---|
-| Install and write the first publisher / subscriber | [Installation](introduction/installation.md) → [Tutorial: Your first outbox app](tutorials/first-outbox-app.md) |
+| Install and write the first publisher / subscriber | [Installation](introduction/installation.md), then [Tutorial: Your first outbox app](tutorials/first-outbox-app.md) |
| See it work end-to-end on a FastAPI app | [FastAPI integration](usage/fastapi.md) |
| Relay outbox rows to Kafka / RabbitMQ / NATS / Redis | [Relay to Kafka / RabbitMQ / NATS](usage/relay.md) |
| Understand the architecture before adopting | [How it works](introduction/how-it-works.md) |
@@ -64,66 +64,66 @@ one-line summary.
### Getting started
-- [Installation](introduction/installation.md) — install, optional
+- [Installation](introduction/installation.md): install, optional
extras (`asyncpg`, `fastapi`, `validate`, `prometheus`,
`opentelemetry`), Postgres setup.
-- [Basic usage](usage/basic.md) — declare the table, create the
+- [Basic usage](usage/basic.md): declare the table, create the
broker, publish a row, register a subscriber.
-- [Tutorial: Your first outbox app](tutorials/first-outbox-app.md) —
+- [Tutorial: Your first outbox app](tutorials/first-outbox-app.md):
build a working publisher / subscriber from scratch.
-- [Tutorial: Add a Kafka relay](tutorials/add-kafka-relay.md) — extend
+- [Tutorial: Add a Kafka relay](tutorials/add-kafka-relay.md): extend
the first app to forward outbox rows to Kafka.
### Concepts
-- [How it works](introduction/how-it-works.md) — two-loop subscriber,
+- [How it works](introduction/how-it-works.md): two-loop subscriber,
lease-token invariant, at-least-once semantics, opt-in DLQ on terminal
failure.
- [Performance](concepts/performance.md): round-trips and table churn,
and the levers that address each.
-- [Comparison](concepts/comparison.md) — vs writing your own, vs CDC,
+- [Comparison](concepts/comparison.md): vs writing your own, vs CDC,
vs Kafka transactions, vs `LISTEN/NOTIFY`, vs Celery, vs FastStream
foreign-broker direct.
-- [Instrumentation seams](concepts/instrumentation-seams.md) — *concept:*
- the recorder seam vs native middleware, and why both exist. **Read this
- first** if you're deciding what to wire.
+- [Instrumentation seams](concepts/instrumentation-seams.md): the
+ recorder seam vs native middleware, and why both exist. Read this
+ first if you're deciding what to wire.
### Guides
-- [FastAPI integration](usage/fastapi.md) — the canonical use case:
+- [FastAPI integration](usage/fastapi.md): the canonical use case:
HTTP routes and outbox subscribers share one `AsyncSession`.
-- [Relay to Kafka / RabbitMQ / NATS](usage/relay.md) — forward outbox
+- [Relay to Kafka / RabbitMQ / NATS](usage/relay.md): forward outbox
rows to a real bus with one decorator; at-least-once preserved.
-- [Timers](usage/timers.md) — `activate_in` / `activate_at`,
+- [Timers](usage/timers.md): `activate_in` / `activate_at`,
`timer_id` dedup, `cancel_timer`.
-- [Testing](usage/testing.md) — `TestOutboxBroker` sync and
+- [Testing](usage/testing.md): `TestOutboxBroker` sync and
loop-driven modes.
-- [Schema validation](usage/schema-validation.md) — opt-in
+- [Schema validation](usage/schema-validation.md): opt-in
Alembic-driven check for `/health` and CI.
-- [Setup Prometheus and OpenTelemetry](usage/setup-prometheus-opentelemetry.md)
- — *step-by-step:* wire the native middleware and recorder adapters
+- [Setup Prometheus and OpenTelemetry](usage/setup-prometheus-opentelemetry.md):
+ step-by-step wiring of the native middleware and recorder adapters
end-to-end.
-- [A messaging service, end-to-end](usage/messaging-service.md) — the relay,
+- [A messaging service, end-to-end](usage/messaging-service.md): the relay,
timer, and testing guides composed in one service.
### Reference
-- [Subscriber](usage/subscriber.md) — options, ack policies, retry
+- [Subscriber](usage/subscriber.md): options, ack policies, retry
strategies, connection budget, slow-handler queue segregation.
-- [Publisher](usage/publisher.md) — `publish`, `publish_batch`,
+- [Publisher](usage/publisher.md): `publish`, `publish_batch`,
`OutboxPublisher`, chained publishing via `OutboxResponse`.
-- [Router](usage/router.md) — `OutboxRouter`, `OutboxRoute`,
+- [Router](usage/router.md): `OutboxRouter`, `OutboxRoute`,
walking every subscriber via `broker.subscribers`.
-- [Dead-letter queue](usage/dlq.md) — opt-in audit table, atomicity
+- [Dead-letter queue](usage/dlq.md): opt-in audit table, atomicity
via a single CTE, `dlq_written` metric, retention patterns.
-- [Observability](usage/observability.md) — *reference:* the recorder-seam
+- [Observability](usage/observability.md): the recorder-seam
API, the full event/tag catalog, and the operator PromQL playbook.
### Operations
-- [Production checklist](operations/checklist.md) — connection budget,
+- [Production checklist](operations/checklist.md): connection budget,
lease TTL sizing, and deploy-safety items before going live.
-- [Troubleshooting](operations/troubleshooting.md) — common symptoms
+- [Troubleshooting](operations/troubleshooting.md): common symptoms
(idle latency, `lease_lost` spikes, connection exhaustion) and fixes.
-- [Alembic migrations](operations/alembic.md) — autogenerate the outbox
+- [Alembic migrations](operations/alembic.md): autogenerate the outbox
table and its partial indexes.
diff --git a/docs/introduction/how-it-works.md b/docs/introduction/how-it-works.md
index 0dba312..5c86e0e 100644
--- a/docs/introduction/how-it-works.md
+++ b/docs/introduction/how-it-works.md
@@ -1,7 +1,7 @@
# How it works
`faststream-outbox` is a FastStream broker integration whose transport is
-**Postgres rows**, not a message bus. A producer writes an outbox row in the
+Postgres rows, not a message bus. A producer writes an outbox row in the
same SQLAlchemy transaction as its domain entity; a subscriber polls the
table, claims rows with `FOR UPDATE SKIP LOCKED`, runs the handler, and
deletes the row on success.
@@ -34,8 +34,8 @@ transactions are the better fit.*
## Producer side
`broker.publish(body, *, queue, session, ...)` inserts an outbox row through
-the caller's `AsyncSession`. It does **not** flush, commit, or open its own
-transaction — the row must commit with the caller's domain writes:
+the caller's `AsyncSession`. It does not flush, commit, or open its own
+transaction; the row must commit with the caller's domain writes:
```python
async with session_factory() as session, session.begin():
@@ -48,11 +48,11 @@ async with session_factory() as session, session.begin():
round-trip for many rows.
The producer also emits `SELECT pg_notify('outbox_', queue)` on the
-caller's session right after the INSERT, **except** when the row is
-genuinely future-dated (a future `activate_in` / `activate_at` — a *past*
+caller's session right after the INSERT, except when the row is
+genuinely future-dated (a future `activate_in` / `activate_at`; a *past*
`activate_at`, e.g. a recovered idempotency token, still notifies) or a
`timer_id` conflict made the insert a no-op. NOTIFY is transactional, so listeners only see it
-after the user's transaction commits — atomicity with the row insert is
+after the user's transaction commits, so atomicity with the row insert is
automatic.
Repeated publishes to the same queue within one transaction emit a single
@@ -63,7 +63,7 @@ bulk publish costs one NOTIFY, not one per row.
Per subscriber, two loops run concurrently:
-**1. Fetch loop.** Owns a long-lived `AsyncConnection` for the fetch CTE and
+The fetch loop owns a long-lived `AsyncConnection` for the fetch CTE and
a separate raw asyncpg connection for `LISTEN outbox_`. A single CTE
claims rows:
@@ -87,20 +87,20 @@ WHERE id IN (SELECT id FROM claimed)
RETURNING *
```
-This is **simplified for illustration**. The real query writes each `OR`
+This is simplified for illustration. The real query writes each `OR`
disjunct with its own partial-index predicate spelled out as a conjunct, so
Postgres can use the `outbox_pending_idx` / `outbox_lease_idx` partial
-indexes instead of a seq-scan — the naive `OR` above is the exact shape the
+indexes instead of a seq-scan. The naive `OR` above is the exact shape the
code avoids.
The CTE reclaims both unleased rows AND rows whose lease has expired
(`acquired_at < now() - lease_ttl_seconds`), so there is no separate stuck-row
-reaper. The idle-sleep is short-circuited by NOTIFY via an `asyncio.Event` —
+reaper. The idle-sleep is short-circuited by NOTIFY via an `asyncio.Event`, so
idle dispatch latency drops from up to `max_fetch_interval` (default 10s) to
~10ms. If LISTEN setup fails (asyncpg missing, non-asyncpg driver, permission
error), the loop logs once and falls back to polling.
-**2. Worker loop** (× `max_workers`). Pulls from an in-process
+The worker loop (× `max_workers`) pulls from an in-process
`asyncio.Queue(maxsize=fetch_batch_size)`, dispatches each row via
`OutboxSubscriber.dispatch_one` (which runs the handler), then
flushes the row's terminal state (`DELETE` on success, `UPDATE
@@ -117,12 +117,12 @@ DELETE FROM outbox WHERE id = :id AND acquired_token = :token
If a slow handler's lease expired and another worker reclaimed the row with
a fresh token, the slow handler's `DELETE` finds `rowcount == 0` and is
-silently dropped — preventing it from clobbering the new lease holder. This
+silently dropped, which keeps it from clobbering the new lease holder. This
is the load-bearing invariant; any new fetch or terminal path must preserve
it.
-`lease_ttl_seconds` (default `60.0`, per-subscriber) **must exceed the P99
-handler duration with margin**, otherwise healthy in-flight handlers race
+`lease_ttl_seconds` (default `60.0`, per-subscriber) must exceed the P99
+handler duration with margin; otherwise healthy in-flight handlers race
their own lease expiry and trigger duplicate deliveries. The lease cutoff is
computed server-side via `make_interval(secs => :lease_ttl)`, so it's
immune to worker / DB clock skew.
@@ -133,8 +133,8 @@ When the invariant fires, the broker emits a WARNING with structured fields:
extra = {"event": "lease_lost", "phase": "terminal" | "retry", "row_id": ..., "queue": ..., "deliveries_count": ...}
```
-Recurring `event=lease_lost` records mean `lease_ttl_seconds < handler P99`
-— that's the operator playbook signal. Log-pipeline aggregators can alert
+Recurring `event=lease_lost` records mean `lease_ttl_seconds < handler P99`;
+that's the operator playbook signal. Log-pipeline aggregators can alert
on the `event` field without parsing the message.
## At-least-once delivery
@@ -144,12 +144,12 @@ successfully. If the worker dies mid-handler, the lease expires and another
worker re-claims the row. The same applies if the handler ran but the
worker crashed before the terminal `DELETE` landed.
-The trade-off: handlers must be **idempotent**. A handler that succeeded
+The trade-off is that handlers must be idempotent. A handler that succeeded
but whose `DELETE` failed to land will be retried.
## Opt-in DLQ on terminal failure
-By default, terminal failures `DELETE` the row — no archive table, no
+By default, terminal failures `DELETE` the row, with no archive table and no
dead-letter queue. Pass `dlq_table=make_dlq_table(metadata)` to the broker
and terminal-by-failure rows are copied into a sibling audit table in the
same Postgres statement as the `DELETE`:
@@ -166,12 +166,12 @@ engine = create_async_engine("postgresql+asyncpg://outbox:outbox@localhost:5432/
broker = OutboxBroker(engine, outbox_table=outbox_table, dlq_table=dlq_table)
```
-Successful rows are never archived — the success path stays a plain
+Successful rows are never archived; the success path stays a plain
`DELETE`. Three failure paths land in the DLQ with a `failure_reason`
column: `max_deliveries`, `retry_terminal`, `rejected`. Atomicity is via a
single CTE (`DELETE … RETURNING` → `INSERT INTO `), so DLQ-write
-failures roll back the `DELETE` — misconfiguration surfaces as outbox
-growth plus `lease_lost` spikes rather than silent audit loss. When
+failures roll back the `DELETE`. Misconfiguration surfaces as outbox
+growth plus `lease_lost` spikes, not silent audit loss. When
`dlq_table` is configured, `broker.validate_schema()` checks both tables
in one call and reports drift on either one. See the
[Dead-letter queue](../usage/dlq.md) page for the schema, atomicity, and
@@ -184,9 +184,9 @@ you add).
## Failure modes
-- **Handlers must be idempotent.** A crash between the handler's side effect and the broker's `DELETE` re-delivers the message — see [At-least-once delivery](#at-least-once-delivery) above.
-- **Best-effort ordering only.** `FOR UPDATE SKIP LOCKED` does not preserve strict order under concurrent workers. If you need strict per-aggregate ordering, route to a single subscriber and run a single worker.
-- **DLQ is opt-in.** Without `dlq_table=`, terminal failures `DELETE` the row.
+- Handlers must be idempotent. A crash between the handler's side effect and the broker's `DELETE` re-delivers the message; see [At-least-once delivery](#at-least-once-delivery) above.
+- Ordering is best-effort only. `FOR UPDATE SKIP LOCKED` does not preserve strict order under concurrent workers. If you need strict per-aggregate ordering, route to a single subscriber and run a single worker.
+- The DLQ is opt-in. Without `dlq_table=`, terminal failures `DELETE` the row.
## Relay to Kafka / RabbitMQ / NATS / Redis
@@ -194,23 +194,23 @@ An `OutboxSubscriber` can source a FastStream-native cross-broker chain:
stack a foreign-broker publisher decorator on the subscriber
(`@kafka_pub @broker_outbox.subscriber("q")`) and the handler's return
value is forwarded to the real bus. The outbox row stays the durability
-boundary — the row commits with the domain write, and the relay carries
+boundary: the row commits with the domain write, and the relay carries
at-least-once end to end. Recovery comes from two tiers, so a bus outage
-never loses the row: a **transient blip** is absorbed by the client library
-(e.g. `aiokafka`) — the in-handler publish blocks until the broker returns,
-which the subscriber sees as one slow *successful* publish, no nack; a
-**sustained outage** eventually raises into the handler, which nacks the row
+never loses the row: a transient blip is absorbed by the client library
+(e.g. `aiokafka`), where the in-handler publish blocks until the broker returns,
+which the subscriber sees as one slow *successful* publish with no nack; a
+sustained outage eventually raises into the handler, which nacks the row
and hands it to the configured `retry_strategy` to reschedule. (The one path
that recovers via lease expiry rather than `retry_strategy` is a
-mis-composed publisher chain — see [relay guardrails](../usage/relay.md#what-not-to-do).)
+mis-composed publisher chain; see [relay guardrails](../usage/relay.md#what-not-to-do).)
-> **Worked end-to-end example → [Relay tutorial](../tutorials/add-kafka-relay.md).**
+For a worked end-to-end example, see the [Relay tutorial](../tutorials/add-kafka-relay.md).
## Acknowledgements
The architecture of this package is heavily informed by Arseniy Popov's
[PR #2704](https://github.com/ag2ai/faststream/pull/2704) (`feat: add sqla
-broker`) on upstream FastStream — the FastStream broker/registrator/subscriber
+broker`) on upstream FastStream. The FastStream broker/registrator/subscriber
wiring, the `SELECT … FOR UPDATE SKIP LOCKED` fetch-and-claim CTE, the retry
strategy hierarchy, and the in-transaction publish contract all originate
from there. This package is a Postgres-only reimplementation that diverges in
diff --git a/docs/introduction/installation.md b/docs/introduction/installation.md
index 60aa19e..43efa63 100644
--- a/docs/introduction/installation.md
+++ b/docs/introduction/installation.md
@@ -23,7 +23,7 @@
## Requirements
- Python 3.11+
-- PostgreSQL 12+ — the features used (partial indexes, `FOR UPDATE SKIP
+- PostgreSQL 12+. The features used (partial indexes, `FOR UPDATE SKIP
LOCKED`, `make_interval`, `pg_notify`) all predate 12. The examples and
CI run on 17; that is what's exercised, so 17 is the safest choice if
you're starting fresh.
@@ -32,7 +32,7 @@
## Free-threaded Python
`faststream-outbox` supports free-threaded (no-GIL) CPython. CI runs the full
-test suite on the **3.14t** interpreter, and a separate CI step asserts that
+test suite on the 3.14t interpreter, and a separate CI step asserts that
importing the outbox and its runtime dependencies keeps the GIL disabled. The
package is pure-Python
asyncio, so nothing in it depends on the GIL; installing on a `python3.14t`
@@ -49,13 +49,12 @@ interpreter resolves the free-threaded wheels of the compiled dependencies
It still applies to any foreign-broker client you install for the
[relay feature](../usage/relay.md): if it hasn't declared free-thread safety
- (for example `aiokafka`), importing it re-enables the GIL. That is the
- client library's limitation, not the outbox's — your outbox code still runs
- correctly.
+ (for example `aiokafka`), importing it re-enables the GIL. The limitation
+ is in the client library; your outbox code still runs correctly.
-What this does **not** change: the subscriber runs a single event loop by
+The subscriber runs a single event loop by
design, so free-threading does not add cross-core parallelism within one
-process. To use more cores, run more subscriber processes — the same scaling
+process. To use more cores, run more subscriber processes, the same scaling
lever as on a GIL build.
Free-threaded wheels for the compiled dependencies currently exist for 3.14t
@@ -75,8 +74,8 @@ docker run -d -p 5432:5432 \
## Optional extras
-The base install ships only `faststream` + `sqlalchemy[asyncio]` — **no
-async Postgres driver**, so you must install one (the `asyncpg` extra below,
+The base install ships only `faststream` + `sqlalchemy[asyncio]`, with no
+async Postgres driver, so you must install one (the `asyncpg` extra below,
or another async driver such as `psycopg`); the `postgresql+asyncpg://` DSNs
in the examples need `asyncpg` specifically. Each optional extra unlocks one
feature; nothing else changes if you omit them.
@@ -84,10 +83,10 @@ feature; nothing else changes if you omit them.
| Extra | Install | What it enables |
|---|---|---|
| `asyncpg` | `pip install 'faststream-outbox[asyncpg]'` | The `asyncpg` SQLAlchemy driver. Required to get `LISTEN/NOTIFY` short-circuit wakeups in the subscriber's fetch loop. With a *different* async driver (e.g. `psycopg`) but no `asyncpg`, the broker still works but the loop falls back to plain polling, which adds up to `max_fetch_interval` (default 10s) of idle latency between an INSERT and a dispatch; with no async driver at all the engine can't connect. |
-| `fastapi` | `pip install 'faststream-outbox[fastapi]'` | The `faststream_outbox.fastapi.OutboxRouter` — see [FastAPI integration](../usage/fastapi.md). |
-| `validate` | `pip install 'faststream-outbox[validate]'` | Alembic, for `broker.validate_schema()` — see [Schema validation](../usage/schema-validation.md). Calling `validate_schema()` without this extra raises `ImportError`; every other code path works. |
-| `prometheus` | `pip install 'faststream-outbox[prometheus]'` | The `PrometheusRecorder` metrics adapter and native `OutboxPrometheusMiddleware` — see [Observability](../usage/observability.md). |
-| `opentelemetry` | `pip install 'faststream-outbox[opentelemetry]'` | The `OpenTelemetryRecorder` metrics adapter and native `OutboxTelemetryMiddleware` — see [Observability](../usage/observability.md). |
+| `fastapi` | `pip install 'faststream-outbox[fastapi]'` | The `faststream_outbox.fastapi.OutboxRouter`. See [FastAPI integration](../usage/fastapi.md). |
+| `validate` | `pip install 'faststream-outbox[validate]'` | Alembic, for `broker.validate_schema()`. See [Schema validation](../usage/schema-validation.md). Calling `validate_schema()` without this extra raises `ImportError`; every other code path works. |
+| `prometheus` | `pip install 'faststream-outbox[prometheus]'` | The `PrometheusRecorder` metrics adapter and native `OutboxPrometheusMiddleware`. See [Observability](../usage/observability.md). |
+| `opentelemetry` | `pip install 'faststream-outbox[opentelemetry]'` | The `OpenTelemetryRecorder` metrics adapter and native `OutboxTelemetryMiddleware`. See [Observability](../usage/observability.md). |
Combine extras with commas:
diff --git a/docs/operations/alembic.md b/docs/operations/alembic.md
index 38ce059..505b282 100644
--- a/docs/operations/alembic.md
+++ b/docs/operations/alembic.md
@@ -1,6 +1,6 @@
# Alembic migrations
-The package never creates or migrates your schema — that's
+The package never creates or migrates your schema; that's
[Alembic](https://alembic.sqlalchemy.org/)'s job. This page shows what
Alembic produces against `make_outbox_table()` and
`make_dlq_table()`, and gives recipes for drift detection and DLQ
@@ -54,17 +54,17 @@ op.create_index(
# ### end Alembic commands ###
```
-The three indexes carry **load-bearing partial predicates**:
+The three indexes carry load-bearing partial predicates:
-- `outbox_pending_idx` — `(queue, next_attempt_at) WHERE acquired_token IS NULL`.
+- `outbox_pending_idx`: `(queue, next_attempt_at) WHERE acquired_token IS NULL`.
Branch A of the fetch CTE (unleased rows) is served by this index. The
WHERE clause is written so Postgres' planner recognizes the implied
predicate and uses the partial index; drop the predicate and the planner
falls back to a seq-scan as the table grows.
-- `outbox_lease_idx` — `(queue, acquired_at) WHERE acquired_token IS NOT NULL`.
+- `outbox_lease_idx`: `(queue, acquired_at) WHERE acquired_token IS NOT NULL`.
Branch B of the fetch CTE (expired-lease reclaim) is served by this
- index. Same story: predicate is load-bearing for fetch performance.
-- `outbox_timer_id_uq` — unique `(queue, timer_id) WHERE timer_id IS NOT NULL`.
+ index. Here too, the predicate is load-bearing for fetch performance.
+- `outbox_timer_id_uq`: unique `(queue, timer_id) WHERE timer_id IS NOT NULL`.
Backs the `timer_id` dedup contract via
`pg_insert(...).on_conflict_do_nothing(...)`. Without the partial
predicate, the unique constraint applies to all rows and breaks
@@ -79,17 +79,17 @@ maintained by the broker; no application code touches them.
The `outbox_lease_ck` CHECK (`(acquired_token IS NULL) = (acquired_at
IS NULL)`) is equally load-bearing: it enforces that a row's lease token
and lease timestamp are always set or cleared together, so a half-written
-lease can never exist. Autogenerate renders it here **only because this is
-a fresh `create_table`** — on an incremental migration onto a pre-existing
-table Alembic has no check-constraint comparator and would ship a missing
+lease can never exist. Autogenerate renders it here only because this is
+a fresh `create_table`. On an incremental migration onto a pre-existing
+table, Alembic has no check-constraint comparator and would ship a missing
or drifted CHECK silently (that gap is what
[`validate_schema()`](../usage/schema-validation.md) backstops).
-The `# please adjust!` comment from Alembic is misleading here —
-**don't adjust**. The column types, the CHECK, the predicates, and the
-indexes are exactly what the broker depends on. The
-[`validate_schema()`](../usage/schema-validation.md) check — when you wire
-it into a `/health` probe or CI gate — fails when the live DB drifts from
+The `# please adjust!` comment from Alembic is misleading here:
+don't adjust. The column types, the CHECK, the predicates, and the
+indexes are exactly what the broker depends on. When you wire the
+[`validate_schema()`](../usage/schema-validation.md) check into a
+`/health` probe or CI gate, it fails when the live DB drifts from
this declaration. (It is opt-in; it never runs at `broker.start()`.)
## Adding the DLQ after the fact
@@ -120,7 +120,7 @@ op.create_index("outbox_dlq_queue_failed_idx", "outbox_dlq", ["queue", "failed_a
# ### end Alembic commands ###
```
-This is **purely additive**: no `op.alter_table` against the outbox
+This migration is purely additive: no `op.alter_table` against the outbox
table itself, no column add, no constraint flip. The runtime change
that activates the DLQ is the broker's
[atomicity CTE](../usage/dlq.md#atomicity) (DELETE … RETURNING → INSERT
@@ -161,8 +161,8 @@ asyncio.run(main())
Non-zero exit on drift; CI fails before "deploy".
-This check is **opt-in for `/health`** and not always-on at
-`broker.start()`. The reason: a running migration plus an always-on
+This check is opt-in for `/health` and does not run at
+`broker.start()`, because a running migration plus an always-on
validator would race. Operators must be able to roll forward a new
schema version without spinning every pod into a crash loop. The drift
check belongs *between* `alembic upgrade head` and the deploy step,
@@ -173,8 +173,8 @@ constructing the broker, `validate_schema()` checks both tables in
one call and surfaces drift on either one.
Both autogenerate and `validate_schema()` run with
-`compare_server_default=False`, so **server-default drift is not
-detected** — neither the autogenerated migration nor the drift gate will
+`compare_server_default=False`, so server-default drift is not
+detected: neither the autogenerated migration nor the drift gate will
flag a column whose `server_default` is missing or wrong. The
consequence that bites is a missing `server_default=now()` on
`next_attempt_at`; see the server-defaults caveat in
@@ -183,19 +183,19 @@ consequence that bites is a missing `server_default=now()` on
## Autovacuum tuning (recommended)
The outbox is a high-churn queue table: every message is one `INSERT`, one lease
-`UPDATE`, and one terminal `DELETE`, so dead tuples accumulate at roughly **twice
-the message rate**, and autovacuum has to reclaim them. Two independent levers
+`UPDATE`, and one terminal `DELETE`, so dead tuples accumulate at roughly twice
+the message rate, and autovacuum has to reclaim them. Two independent levers
matter, and they do different things.
For how this fits with the outbox's other performance levers, see the
[Performance](../concepts/performance.md) guide.
-### Eligibility — when autovacuum fires
+### Eligibility: when autovacuum fires
Postgres' default `autovacuum_vacuum_scale_factor = 0.2` fires vacuum only after a
*fraction of the table* is dead. On a queue table whose `reltuples` estimate can go
stale (a table that backed up once keeps a high estimate), that fraction is a high,
-size-dependent bar that fires rarely — the classic queue-table death-spiral.
+size-dependent bar that fires rarely. This is the classic queue-table death spiral.
Setting the scale factor to `0` with a constant threshold makes vacuum eligible on a
fixed dead-tuple count instead, independent of table size and of a stale estimate:
@@ -210,19 +210,19 @@ def upgrade() -> None:
This sets `autovacuum_vacuum_scale_factor = 0` and `autovacuum_vacuum_threshold =
1000` (plus the insert-triggered pair, Postgres 13+). Tune the thresholds for your
-message rate — `outbox_autovacuum_ddl("outbox", vacuum_threshold=5000,
+message rate, for example `outbox_autovacuum_ddl("outbox", vacuum_threshold=5000,
insert_threshold=5000)`.
This is standard queue-table hygiene: it keeps vacuum behavior predictable and
-size-independent, and it matters most under the **default 60-second autovacuum
-daemon** with variable or bursty backlogs, where being eligible at *every* wake
+size-independent, and it matters most under the default 60-second autovacuum
+daemon with variable or bursty backlogs, where being eligible at *every* wake
(rather than rarely) bounds how many dead tuples pile up between vacuums. Treat it as
insurance against a size-dependent bar going stale, not a guaranteed bloat
reduction.
-### Throughput — how fast autovacuum reclaims
+### Throughput: how fast autovacuum reclaims
-Eligibility is necessary but not sufficient. Under **heavy sustained churn** the
+Eligibility is necessary but not sufficient. Under heavy sustained churn, the
binding constraint is vacuum *throughput*: a throttled autovacuum (Postgres default
`autovacuum_vacuum_cost_delay`) falls behind the dead-tuple rate no matter how
eagerly it is eligible, and the table bloats anyway. The `vacuum_cost_delay` /
@@ -233,8 +233,8 @@ eagerly it is eligible, and the table bloats anyway. The `vacuum_cost_delay` /
op.execute(outbox_autovacuum_ddl("outbox", vacuum_cost_delay=0))
```
-`vacuum_cost_delay=0` removes autovacuum's I/O throttle for this table — the lever
-that actually bounds bloat under heavy churn. But it is **I/O-heavy**: on a shared
+`vacuum_cost_delay=0` removes autovacuum's I/O throttle for this table, and it is the lever
+that bounds bloat under heavy churn. It is also I/O-heavy: on a shared
cluster, an unthrottled vacuum of a large table can spike disk load, so raise it
deliberately and measure. Both cost params default to unset (the cluster default),
so a plain `outbox_autovacuum_ddl("outbox")` changes only eligibility.
@@ -243,7 +243,7 @@ so a plain `outbox_autovacuum_ddl("outbox")` changes only eligibility.
HOT updates are impossible and `fillfactor` buys almost nothing here.
If your outbox table lives in a non-default `MetaData(schema=...)`, pass the same
-schema — `outbox_autovacuum_ddl("outbox", schema="app")` — so the `ALTER TABLE`
+schema (`outbox_autovacuum_ddl("outbox", schema="app")`) so the `ALTER TABLE`
targets that table rather than an unqualified name resolved via `search_path`.
To catch a table that never had the eligibility settings applied, pass
@@ -259,21 +259,21 @@ await broker.validate_schema(check_autovacuum=True)
## Fixing drift autogenerate can't see { #fixing-drift-autogenerate-cant-see }
Two kinds of drift that
-[`validate_schema()`](../usage/schema-validation.md) reports **cannot** be
-remediated by `alembic revision --autogenerate` — the same blindness that
+[`validate_schema()`](../usage/schema-validation.md) reports cannot be
+remediated by `alembic revision --autogenerate`. The same blindness that
let them drift in also stops autogenerate from emitting a fix:
-- **The `outbox_lease_ck` CHECK constraint.** Alembic's `compare_metadata`
+- The `outbox_lease_ck` CHECK constraint: Alembic's `compare_metadata`
has no check-constraint comparator, so a missing or altered CHECK never
appears in an autogenerated migration.
-- **Partial-index predicates.** Alembic's index comparator ignores
+- Partial-index predicates: Alembic's index comparator ignores
`postgresql_where`, so an index that exists but was created non-partial,
with the wrong `WHERE`, or (for `outbox_timer_id_uq`) non-unique is
invisible to the diff.
When `validate_schema()` raises for one of these, its error ends with a
pointer to this section. Re-running autogenerate produces an empty
-`upgrade()` — hand-write the migration instead, then re-run
+`upgrade()`, so hand-write the migration instead, then re-run
`validate_schema()` to confirm the drift is cleared.
### Restore the lease CHECK
@@ -314,14 +314,14 @@ op.create_index(
Substitute the columns / `unique` / predicate from the table above for
`outbox_pending_idx` and `outbox_lease_idx`.
-The recipes pass literal names (`'outbox_lease_ck'`, `'outbox_timer_id_uq'`) —
-the exact names the package emits **with no `naming_convention`**.
+The recipes pass literal names (`'outbox_lease_ck'`, `'outbox_timer_id_uq'`),
+which are the exact names the package emits when no `naming_convention` is set.
### Naming conventions: the CHECK name doesn't matter
-`validate_schema()`'s CHECK probe matches the lease constraint **by predicate,
-not name**. So if your `MetaData` carries a SQLAlchemy `naming_convention` with a
-`ck` key, you don't need to match any particular name — create the constraint
+`validate_schema()`'s CHECK probe matches the lease constraint by predicate,
+not by name. So if your `MetaData` carries a SQLAlchemy `naming_convention` with a
+`ck` key, you don't need to match any particular name. Create the constraint
under whatever name your migration produces and it will validate, as long as its
predicate is `(acquired_token IS NULL) = (acquired_at IS NULL)`. The literal
`op.create_check_constraint('outbox_lease_ck', ...)` recipe above is fine even
@@ -330,21 +330,21 @@ under a convention.
(Why the probe ignores the name: a `ck` convention re-templates the in-memory
`CheckConstraint.name` to e.g. `ck_outbox_outbox_lease_ck`, but a hand-written
`op.create_check_constraint('outbox_lease_ck', ...)` creates the literal name
-verbatim — Alembic op functions don't apply the convention. The live name is
+verbatim, because Alembic op functions don't apply the convention. The live name is
therefore unpredictable, so the probe keys off the stable predicate instead.)
-The **index** recipes still use literal names, because the explicitly-named
-indexes (`outbox_pending_idx` etc.) are **not** re-templated by the `ix`/`uq`
-convention keys — those only rename auto-named indexes.
+The index recipes still use literal names, because the explicitly named
+indexes (`outbox_pending_idx` etc.) are not re-templated by the `ix`/`uq`
+convention keys, which only rename auto-named indexes.
## DLQ retention via partition drop { #dlq-retention-via-partition-drop }
Plain `DELETE FROM outbox_dlq WHERE failed_at < now() - interval '90
days'` works fine for low-volume DLQs (< ~1 GB / month) and needs no
schema change. For higher volume, converting the DLQ to a
-range-partitioned table by `failed_at` lets you **drop entire
-partitions** instead of deleting row by row — orders-of-magnitude
-faster, no vacuum debt.
+range-partitioned table by `failed_at` lets you drop entire
+partitions instead of deleting row by row, which is orders of magnitude
+faster and leaves no vacuum debt.
### One-time migration to partitioned DLQ
@@ -397,7 +397,7 @@ op.execute("""
op.drop_table("outbox_dlq_old")
```
-`validate_schema()` continues to work against the partitioned table —
+`validate_schema()` works against the partitioned table because
Alembic's autogenerate ignores partition boundaries when comparing
column shape.
diff --git a/docs/operations/checklist.md b/docs/operations/checklist.md
index 5554459..c3bc215 100644
--- a/docs/operations/checklist.md
+++ b/docs/operations/checklist.md
@@ -6,84 +6,84 @@ story.
## Sizing
-- [ ] **Engine pool ≥ `Σ subs × (max_workers + 1)`** — every
+- [ ] Engine pool ≥ `Σ subs × (max_workers + 1)`. Every
subscriber holds `max_workers + 1` SQLAlchemy pool connections (one
writer per worker + one fetch) plus one raw asyncpg connection for
`LISTEN`. Sub-budget formula in [Subscriber § Connection
budget](../usage/subscriber.md#connection-budget).
-- [ ] **Postgres `max_connections` ≥ `replicas × Σ subs × (max_workers + 2)`**
- — `max_workers + 1` pool connections **plus** the raw asyncpg `LISTEN`
- connection per subscriber; the formula is per-process and rolling deploys
- multiply it. Failure mode: pods refuse with `FATAL: too many connections`.
+- [ ] Postgres `max_connections` ≥ `replicas × Σ subs × (max_workers + 2)`.
+ That covers `max_workers + 1` pool connections plus the raw asyncpg `LISTEN`
+ connection per subscriber. The formula is per-process, and rolling deploys
+ multiply it. If it is too low, pods refuse with `FATAL: too many connections`.
## Subscribers
-- [ ] **`lease_ttl_seconds` > handler P99 with margin** — otherwise
+- [ ] `lease_ttl_seconds` > handler P99 with margin. Otherwise
healthy in-flight handlers race their own lease expiry. The lease
cutoff is server-side `make_interval(...)`, immune to clock skew.
- Tuning: [Subscriber § Slow handlers — dedicated
+ Tuning: [Subscriber § Slow handlers: dedicated
queue](../usage/subscriber.md#slow-handlers-dedicated-queue).
-- [ ] **Slow handlers segregated** onto their own subscriber with a
- taller `lease_ttl_seconds`. Don't raise it globally — that delays
+- [ ] Slow handlers segregated onto their own subscriber with a
+ taller `lease_ttl_seconds`. Don't raise it globally; that delays
reclaim of *actually* stuck rows everywhere.
-- [ ] **`max_deliveries` set** (or knowingly unbounded). Default is
+- [ ] `max_deliveries` set (or knowingly unbounded). Default is
unbounded; pair with a non-`NoRetry()` retry strategy or
wedge-prone handlers can replay forever.
-- [ ] **Retry strategy chosen.** Default
+- [ ] Retry strategy chosen. The default
`ExponentialRetry(initial_delay_seconds=1.0, multiplier=2.0,
max_delay_seconds=300.0, max_attempts=10, jitter_factor=0.2)` is fine
for most. Opt into `NoRetry()` explicitly for an audit feed.
## DLQ
-- [ ] **`dlq_table=` configured** — opt-in but recommended for any
+- [ ] `dlq_table=` configured. It is opt-in but recommended for any
service where terminal failures need forensic recovery. See
[Dead-letter queue](../usage/dlq.md).
-- [ ] **Alert on `nacked_terminal` rate vs `dlq_written` divergence**
- — persistent divergence means either DLQ schema drift (CTE rolls
+- [ ] Alert on `nacked_terminal` rate vs `dlq_written` divergence.
+ Persistent divergence means either DLQ schema drift (CTE rolls
back) or `lease_ttl_seconds` too low. See [DLQ § Metric:
dlq_written](../usage/dlq.md#metric-dlq_written).
-- [ ] **DLQ retention plan.** Partition by `failed_at` + cron-drop old
+- [ ] DLQ retention plan. Partition by `failed_at` + cron-drop old
partitions, or a simple `DELETE … WHERE failed_at < interval` cron
for low volume. Walk-through: [Alembic migrations § DLQ retention via
partition drop](./alembic.md#dlq-retention-via-partition-drop).
## Drain & lifecycle
-- [ ] **`graceful_timeout` ≥ handler P99 + margin** — otherwise
+- [ ] `graceful_timeout` ≥ handler P99 + margin. Otherwise
`OutboxSubscriber.stop()` cancels in-flight work and rows are
reclaimed mid-handler.
-- [ ] **Kubernetes `terminationGracePeriodSeconds` ≥ broker
- `graceful_timeout`** with margin for the parallel-subscriber drain.
+- [ ] Kubernetes `terminationGracePeriodSeconds` ≥ broker
+ `graceful_timeout`, with margin for the parallel-subscriber drain.
The broker gathers subscriber drains in parallel, but k8s
`SIGKILL`s after the grace period regardless.
## Schema
-- [ ] **`/health` calls `validate_schema()`** — opt-in; requires the
- `[validate]` extra. Do **not** call at `broker.start()` — that
+- [ ] `/health` calls `validate_schema()`. This is opt-in and requires the
+ `[validate]` extra. Do not call it at `broker.start()`, because that
would crash-loop on a pending migration. See [Schema validation §
Where to call it](../usage/schema-validation.md#where-to-call-it).
-- [ ] **Outbox `table_name` short enough for every derived identifier** —
+- [ ] Outbox `table_name` short enough for every derived identifier.
`make_outbox_table` raises `ValueError` at table-build time when the
- *longest* identifier it derives exceeds Postgres' 63-**byte** limit. That
+ *longest* identifier it derives exceeds Postgres' 63-byte limit. That
longest identifier is usually an index/constraint name
(`_pending_idx`, `_timer_id_uq`), which is longer
- than the NOTIFY channel `outbox_` — so a name that fits the
+ than the NOTIFY channel `outbox_`, so a name that fits the
channel can still overflow an index name. There is no silent truncation or
polling fallback; the guard makes an over-long name impossible to ship.
## Observability
-- [ ] **`metrics_recorder` set, native middleware registered, or
- both** — the recommended setup is both. See [Instrumentation seams §
+- [ ] `metrics_recorder` set, native middleware registered, or
+ both. The recommended setup is both. See [Instrumentation seams §
Layering](../concepts/instrumentation-seams.md#layering-middleware-seam-vs-recorder-seam).
-- [ ] **Alert on `lease_lost` rate** — non-zero means
+- [ ] Alert on `lease_lost` rate. Non-zero means
`lease_ttl_seconds < handler P99` for at least one subscriber. See
[Troubleshooting § `event=lease_lost`](./troubleshooting.md#event-lease_lost-recurring-in-logs).
-- [ ] **`LISTEN/NOTIFY` fallback understood** — a *connection* or
+- [ ] `LISTEN/NOTIFY` fallback understood. A *connection* or
*permission* failure (`asyncpg.connect` / `add_listener` raising) logs a
- WARNING once and falls back to polling. A **missing asyncpg driver or a
- non-asyncpg engine URL falls back silently** (no log) — diagnose those
+ WARNING once and falls back to polling. A missing asyncpg driver or a
+ non-asyncpg engine URL falls back silently (no log), so diagnose those
from the engine URL, not the logs. Either way the subscriber lives with
up-to-`max_fetch_interval` idle latency.
diff --git a/docs/operations/troubleshooting.md b/docs/operations/troubleshooting.md
index 60ff4c6..28adf78 100644
--- a/docs/operations/troubleshooting.md
+++ b/docs/operations/troubleshooting.md
@@ -1,15 +1,15 @@
# Troubleshooting
-Symptom → likely cause → fix. Each section below is the same shape:
-what you see, what's probably wrong, how to confirm, what to change,
-and a link into the reference page that owns the underlying design.
+Each entry below starts with what you see, then gives the probable cause,
+how to confirm it, what to change, and a link to the reference page that
+owns the underlying design.
| Symptom | Likely cause |
|---|---|
| [`event=lease_lost` recurring in logs](#event-lease_lost-recurring-in-logs) | Handler P99 > `lease_ttl_seconds` |
| [Outbox row count grows + `lease_lost` spike](#outbox-row-count-grows-lease_lost-spike) | DLQ CTE failing (DLQ schema drift) |
| [Outbox row count grows, no `lease_lost`](#outbox-row-count-grows-no-lease_lost) | Fetch loop not running, or rows future-dated |
-| [Idle dispatch latency > `max_fetch_interval`](#idle-dispatch-latency-max_fetch_interval) | LISTEN setup failed → polling fallback |
+| [Idle dispatch latency > `max_fetch_interval`](#idle-dispatch-latency-max_fetch_interval) | LISTEN setup failed, so the subscriber falls back to polling |
| [Subscriber dispatch never starts; rows pile up](#subscriber-blocks-at-brokerstart) | Engine pool exhausted on writer-connection checkout |
| [Duplicate handler invocations](#duplicate-handler-invocations) | Lease expired before handler returned, or handler not idempotent |
| [Rolling deploy leaks rows](#rolling-deploy-leaks-rows) | `graceful_timeout` < handler P99, or k8s grace too short |
@@ -21,298 +21,294 @@ and a link into the reference page that owns the underlying design.
## `event=lease_lost` recurring in logs { #event-lease_lost-recurring-in-logs }
-**Symptom.** WARNING-level logs with the message text
+You see WARNING-level logs with the message text
`lease expired before terminal write` or `lease expired before retry write`,
one per affected row. The record also carries `event=lease_lost` and
`phase=terminal` / `phase=retry` as structured extras, visible when your log
formatter renders extras (for example a JSON formatter).
-**Likely cause.** The subscriber's `lease_ttl_seconds` is shorter than
-the handler's P99 duration. A handler took longer than the lease,
-another fetch reclaimed the row mid-flight, and the original handler's
-terminal `DELETE` / `UPDATE` matched zero rows.
+The likely cause is a subscriber `lease_ttl_seconds` shorter than the
+handler's P99 duration. A handler took longer than the lease, another
+fetch reclaimed the row mid-flight, and the original handler's terminal
+`DELETE` / `UPDATE` matched zero rows.
-**Diagnose.** Grep for `lease expired before` (or, with an
+To confirm, grep for `lease expired before` (or, with an
extras-rendering formatter, `event=lease_lost`) over the last hour and
-compare the rate against `dispatched`. A non-zero baseline rate
-(rather than occasional spikes) confirms TTL is the issue.
+compare the rate against `dispatched`. A steady non-zero rate, as
+opposed to occasional spikes, confirms that the TTL is the issue.
-**Fix.** Raise `lease_ttl_seconds` for the affected subscriber, OR
-segregate slow work onto its own subscriber with a taller TTL
-(recommended — keeps the fast queue's reclaim tight). TTL must exceed
-handler P99 with margin.
+To fix it, raise `lease_ttl_seconds` for the affected subscriber, or move
+slow work onto its own subscriber with a taller TTL. The second option is
+recommended because it keeps the fast queue's reclaim tight. Either way,
+the TTL must exceed handler P99 with margin.
-**Reference.** [Subscriber § Slow handlers — dedicated
+See [Subscriber § Slow handlers: dedicated
queue](../usage/subscriber.md#slow-handlers-dedicated-queue).
## Outbox row count grows + `lease_lost` spike { #outbox-row-count-grows-lease_lost-spike }
-**Symptom.** Two things at once: row count in the outbox table grows
-without bound, *and* `event=lease_lost` log rate spikes.
+Two things happen at once: the row count in the outbox table grows
+without bound, *and* the `event=lease_lost` log rate spikes.
-**Likely cause.** The DLQ CTE is failing on every terminal flush —
-DLQ schema drift means the `INSERT INTO ` clause inside the
-`WITH deleted AS (DELETE … RETURNING …)` statement rolls back the
-DELETE too. Rows stay in the outbox, leases keep expiring, the
+The likely cause is a DLQ CTE that fails on every terminal flush. With
+DLQ schema drift, the `INSERT INTO ` clause inside the
+`WITH deleted AS (DELETE … RETURNING …)` statement fails and rolls back
+the DELETE too. Rows stay in the outbox, leases keep expiring, and the
pattern compounds.
-**Diagnose.** Run `await broker.validate_schema()` against the live
-DB (the `[validate]` extra is required). It will surface missing
-columns / indexes on the DLQ table. A frequent cause on older
-deployments is a hand-written DLQ migration missing the `timer_id`
-column — `validate_schema()` reports it as a missing column on the DLQ
+To confirm, run `await broker.validate_schema()` against the live
+DB (the `[validate]` extra is required). It reports missing
+columns and indexes on the DLQ table. A frequent cause on older
+deployments is a hand-written DLQ migration without the `timer_id`
+column, which `validate_schema()` reports as a missing column on the DLQ
table. The [Alembic guide](../operations/alembic.md#adding-the-dlq-after-the-fact)
includes it.
-**Fix.** Bring the DLQ schema up to spec (apply the missing migration,
-or rename / drop the drifted column / index). After the schema is
+To fix it, bring the DLQ schema up to spec: apply the missing migration,
+or rename or drop the drifted column or index. Once the schema is
correct, the next claim of each stuck row flushes through the CTE
-and the outbox drains naturally.
+and the outbox drains on its own.
-**Recommended alerts.** A persistent DLQ misconfiguration (or a permanent
-relay config error) is the one way a config bug degrades into a
-storage-exhaustion outage — the affected rows cycle through fetch/fail
-forever while new rows accumulate. There is no built-in circuit breaker,
-so **alert on outbox row count (trend / absolute ceiling) and on the
-`lease_lost` rate**, and watch `dlq_written` vs `nacked_terminal`
-divergence (a gap means terminal failures aren't reaching the DLQ).
+A persistent DLQ misconfiguration (or a permanent relay config error) is
+the one way a config bug degrades into a storage-exhaustion outage: the
+affected rows cycle through fetch and fail forever while new rows
+accumulate. There is no built-in circuit breaker, so alert on the outbox
+row count (trend and absolute ceiling) and on the `lease_lost` rate. Also
+watch for divergence between `dlq_written` and `nacked_terminal`; a gap
+means terminal failures aren't reaching the DLQ.
-**Reference.** [DLQ § Atomicity](../usage/dlq.md#atomicity), [Schema
+See [DLQ § Atomicity](../usage/dlq.md#atomicity) and [Schema
validation](../usage/schema-validation.md).
## Outbox row count grows, no `lease_lost` { #outbox-row-count-grows-no-lease_lost }
-**Symptom.** Outbox rows accumulate, but logs are clean — no
-`lease_lost`, no exceptions.
+Outbox rows accumulate, but the logs are clean, with no `lease_lost` and
+no exceptions.
-**Likely cause.** Either no subscriber is registered for that queue,
-or the rows are future-dated (`activate_in` / `activate_at` set) and
-genuinely waiting to fire.
+Either no subscriber is registered for that queue, or the rows are
+future-dated (`activate_in` / `activate_at` set) and are waiting to fire.
-**Diagnose.** Inspect a stuck row's `next_attempt_at` — if it's in the
-future, the row is correctly waiting. Otherwise check whether a
-subscriber is registered: walk `broker.subscribers`, which covers
+To tell the two apart, inspect a stuck row's `next_attempt_at`. If it's
+in the future, the row is correctly waiting. Otherwise check whether a
+subscriber is registered by walking `broker.subscribers`, which covers
router-attached subscribers too.
-**Fix.** Register the subscriber, or adjust the producer's `activate_*`
-arg if the future date was unintentional.
+To fix it, register the subscriber, or adjust the producer's `activate_*`
+argument if the future date was unintentional.
-**Reference.** [Subscriber](../usage/subscriber.md), [Router § Gotcha:
+See [Subscriber](../usage/subscriber.md), [Router § Gotcha:
walking every subscriber](../usage/router.md#gotcha-walking-every-subscriber),
-[Timers](../usage/timers.md).
+and [Timers](../usage/timers.md).
## Idle dispatch latency > `max_fetch_interval` { #idle-dispatch-latency-max_fetch_interval }
-**Symptom.** Rows arrive but take up to `max_fetch_interval` (default
-10 s) to dispatch, even though no other rows are in flight. NOTIFY
-should short-circuit the idle wait to ~10 ms.
+Rows arrive but take up to `max_fetch_interval` (default 10 s) to
+dispatch, even though no other rows are in flight. NOTIFY should cut the
+idle wait short to about 10 ms.
-**Likely cause.** `LISTEN` setup failed at subscriber start. The raw
+The likely cause is a `LISTEN` setup failure at subscriber start. The raw
asyncpg connection that owns `LISTEN outbox_` is separate from
-the SQLAlchemy fetch connection; common failure modes are: the
-asyncpg driver isn't installed (no `[asyncpg]` extra), the engine URL
-is not asyncpg, or Postgres user lacks `LISTEN` permission.
-
-**Diagnose.** A connection or permission failure (`asyncpg.connect` /
-`add_listener` raising) logs a WARNING once at startup noting the NOTIFY
-fallback to polling. A **missing asyncpg driver or a non-asyncpg engine URL
-falls back silently** — there is no log line, so check the engine URL
-(`drivername` must be `postgresql+asyncpg`) and that the `[asyncpg]` extra
-is installed.
-
-**Fix.** Install the `[asyncpg]` extra and use an asyncpg-driven
-engine URL (`postgresql+asyncpg://...`). Restart the subscriber.
-
-**Reference.** [Installation § Optional extras
-](../introduction/installation.md#optional-extras), [How it works §
+the SQLAlchemy fetch connection. Common failure modes are a missing
+asyncpg driver (no `[asyncpg]` extra), an engine URL that is not asyncpg,
+and a Postgres user without `LISTEN` permission.
+
+A connection or permission failure (`asyncpg.connect` or `add_listener`
+raising) logs a WARNING once at startup noting the NOTIFY fallback to
+polling. A missing asyncpg driver or a non-asyncpg engine URL falls back
+silently, with no log line. In that case, check that the engine URL's
+`drivername` is `postgresql+asyncpg` and that the `[asyncpg]` extra is
+installed.
+
+To fix it, install the `[asyncpg]` extra, use an asyncpg-driven engine
+URL (`postgresql+asyncpg://...`), and restart the subscriber.
+
+See [Installation § Optional extras
+](../introduction/installation.md#optional-extras) and [How it works §
Fetch loop](../introduction/how-it-works.md#subscriber-two-async-loops).
## Subscriber dispatch never starts; rows pile up { #subscriber-blocks-at-brokerstart }
-**Symptom.** Rows are published but never dispatched (the table grows) and
-the subscriber's loops emit repeating reconnect ERROR logs. `broker.start()`
-(or the FastAPI `include_router` lifespan) itself returns normally — it only
-schedules the loop tasks, so the failure shows up *after* startup, not as a
-hang.
+Rows are published but never dispatched (the table grows), and the
+subscriber's loops emit repeating reconnect ERROR logs. `broker.start()`
+(or the FastAPI `include_router` lifespan) returns normally because it
+only schedules the loop tasks, so the failure shows up *after* startup
+and does not look like a hang.
-**Likely cause.** SQLAlchemy pool exhausted on the per-worker writer
-connection checkout — the fetch/worker loops can't acquire their
+The likely cause is an exhausted SQLAlchemy pool on the per-worker writer
+connection checkout. The fetch and worker loops can't acquire their
connections, so each cycle errors and backs off. Each subscriber needs
-`max_workers + 1` pool connections; the default pool is `pool_size=5,
-max_overflow=10`. A handful of single-worker subscribers fits, but a fleet
+`max_workers + 1` pool connections, and the default pool is `pool_size=5,
+max_overflow=10`. A handful of single-worker subscribers fits; a fleet
of high-`max_workers` subscribers does not.
-**Diagnose.** Inspect the engine pool. Compute `Σ subs × (max_workers
-+ 1)` from your subscriber registrations and compare to
+To confirm, inspect the engine pool. Compute `Σ subs × (max_workers
++ 1)` from your subscriber registrations and compare it to
`pool_size + max_overflow`.
-**Fix.** Raise `pool_size` / `max_overflow` on the engine, OR lower
-`max_workers` per subscriber. Also confirm Postgres
+To fix it, raise `pool_size` / `max_overflow` on the engine, or lower
+`max_workers` per subscriber. Also confirm that Postgres has
`max_connections ≥ replicas × Σ subs × (max_workers + 2)` (the pool's
-`max_workers + 1` plus the raw `LISTEN` connection) — rolling
-deploys multiply the demand.
+`max_workers + 1` plus the raw `LISTEN` connection). Rolling deploys
+multiply the demand.
-**Reference.** [Subscriber § Connection
-budget](../usage/subscriber.md#connection-budget), [Production
+See [Subscriber § Connection
+budget](../usage/subscriber.md#connection-budget) and [Production
checklist § Sizing](./checklist.md#sizing).
## Duplicate handler invocations
-**Symptom.** The same outbox row's handler runs more than once. Side
-effects double up if the handler isn't idempotent.
+The same outbox row's handler runs more than once, and side effects
+double up if the handler isn't idempotent.
-**Likely cause.** Either the handler's wall-clock duration exceeded
-`lease_ttl_seconds` and another fetch reclaimed the row mid-flight,
-or the worker crashed between the handler's external side effect and
-the terminal `DELETE`. Both are at-least-once-delivery edge cases.
+There are two likely causes, both edge cases of at-least-once delivery.
+Either the handler's wall-clock duration exceeded `lease_ttl_seconds`
+and another fetch reclaimed the row mid-flight, or the worker crashed
+between the handler's external side effect and the terminal `DELETE`.
-**Diagnose.** Cross-reference handler-side logs (the side effect)
+To tell them apart, cross-reference handler-side logs (the side effect)
with `lease expired before` warnings (`event=lease_lost` with an
-extras-rendering formatter). Matching row IDs confirm TTL is too
-short. Crash-induced duplicates correlate with worker-process
-restarts.
+extras-rendering formatter). Matching row IDs confirm that the TTL is too
+short. Crash-induced duplicates correlate with worker-process restarts.
-**Fix.** Handlers must be idempotent; delivery is at-least-once. Also
-tune `lease_ttl_seconds` above handler P99 so healthy handlers don't
-race their lease.
+Delivery is at-least-once, so handlers must be idempotent. Also tune
+`lease_ttl_seconds` above handler P99 so healthy handlers don't race
+their lease.
-**Reference.** [How it works § At-least-once
-delivery](../introduction/how-it-works.md#at-least-once-delivery),
-[Subscriber § Slow handlers — dedicated
+See [How it works § At-least-once
+delivery](../introduction/how-it-works.md#at-least-once-delivery) and
+[Subscriber § Slow handlers: dedicated
queue](../usage/subscriber.md#slow-handlers-dedicated-queue).
## Rolling deploy leaks rows
-**Symptom.** During a rolling restart, outbox rows are left in the
-"acquired" state until lease expiry, even though handlers were
-nominally healthy. Drain duration appears longer than expected.
+During a rolling restart, outbox rows stay in the "acquired" state until
+lease expiry, even though handlers were nominally healthy. Draining takes
+longer than expected.
-**Likely cause.** Either the broker's `graceful_timeout` is shorter
-than the in-flight handler's remaining work, or Kubernetes
-`terminationGracePeriodSeconds` is shorter than the broker's
-`graceful_timeout` (subscribers drain concurrently, so a clean shutdown
-takes about one `graceful_timeout`), and `SIGKILL` arrives mid-drain.
+Either the broker's `graceful_timeout` is shorter than the in-flight
+handler's remaining work, or Kubernetes `terminationGracePeriodSeconds`
+is shorter than the broker's `graceful_timeout` and `SIGKILL` arrives
+mid-drain. Subscribers drain concurrently, so a clean shutdown takes
+about one `graceful_timeout`.
-**Diagnose.** Time a clean shutdown locally (`docker compose kill -s
-SIGTERM application`) and compare to your k8s grace period. Look for
+To confirm, time a clean shutdown locally (`docker compose kill -s
+SIGTERM application`) and compare it to your k8s grace period. Look for
log lines indicating drain abandonment.
-**Fix.** Raise `graceful_timeout` past handler P99 + margin. Raise
-`terminationGracePeriodSeconds` past `graceful_timeout` plus a
-buffer. The `dispatch_one` shutdown-race guard is
-always on; you don't need to opt into it.
+To fix it, raise `graceful_timeout` past handler P99 plus margin, and
+raise `terminationGracePeriodSeconds` past `graceful_timeout` plus a
+buffer. The `dispatch_one` shutdown-race guard is always on; you don't
+need to opt into it.
-**Reference.** [Production checklist § Drain &
+See [Production checklist § Drain &
lifecycle](./checklist.md#drain-lifecycle).
## `activate_in` / `activate_at` fires immediately in tests { #activate_in-activate_at-fires-immediately-in-tests }
-**Symptom.** A unit test publishes a row with `activate_in=30s` and
-the handler runs synchronously inside `await broker.publish(...)`.
+A unit test publishes a row with `activate_in=30s`, and the handler runs
+synchronously inside `await broker.publish(...)`.
-**Likely cause.** By design. `TestOutboxBroker(run_loops=False)`
-(the default) drives handlers synchronously through `dispatch_one`,
-which ignores `next_attempt_at`. This is the documented test-broker
-contract — trades production parity for test ergonomics.
+This is by design. `TestOutboxBroker(run_loops=False)` (the default)
+drives handlers synchronously through `dispatch_one`, which ignores
+`next_attempt_at`. That is the documented test-broker contract: it
+trades production parity for test ergonomics.
-**Diagnose.** Check the call site: `TestOutboxBroker(broker)` →
-sync mode, expected immediate firing.
+Check the call site. `TestOutboxBroker(broker)` runs in sync mode, where
+immediate firing is expected.
-**Fix.** Opt into `TestOutboxBroker(broker, run_loops=True)` for
-tests that need scheduled delivery to actually wait. Loop mode runs
-the real fetch and worker loops against the fake client.
+For tests that need scheduled delivery to wait, opt into
+`TestOutboxBroker(broker, run_loops=True)`. Loop mode runs the real
+fetch and worker loops against the fake client.
-**Reference.** [Testing § Loop-driven
-mode](../usage/testing.md#loop-driven-mode), [Timers § Test broker
+See [Testing § Loop-driven
+mode](../usage/testing.md#loop-driven-mode) and [Timers § Test broker
note](../usage/timers.md#test-broker-note).
## `AckPolicy.ACK_FIRST` raises `ValueError` at registration { #ackpolicyack_first-raises-valueerror-at-registration }
-**Symptom.** `@broker.subscriber("q", ack_policy=AckPolicy.ACK_FIRST)`
-fails with `ValueError` at decoration time.
+`@broker.subscriber("q", ack_policy=AckPolicy.ACK_FIRST)` fails with
+`ValueError` at decoration time.
-**Likely cause.** By design. `ACK_FIRST` would delete the outbox row
-*before* the handler runs, so a handler crash would silently drop
-the message — exactly the failure mode the outbox pattern exists to
-prevent.
+This is by design. `ACK_FIRST` would delete the outbox row *before* the
+handler runs, so a handler crash would silently drop the message, which
+is exactly the failure the outbox pattern exists to prevent. The error
+message names the policy, so there is nothing else to diagnose.
-**Diagnose.** None needed; the message identifies the policy.
+Use the default `AckPolicy.NACK_ON_ERROR` (retry on handler exception
+via the configured retry strategy), `AckPolicy.REJECT_ON_ERROR` (delete
+on first failure), or `AckPolicy.MANUAL` (the handler calls `ack` /
+`nack` / `reject`).
-**Fix.** Use the default `AckPolicy.NACK_ON_ERROR` (retry on handler
-exception via the configured retry strategy), or
-`AckPolicy.REJECT_ON_ERROR` (delete on first failure), or
-`AckPolicy.MANUAL` (handler calls `ack` / `nack` / `reject`).
-
-**Reference.** [Subscriber § Ack
-policy](../usage/subscriber.md#ack-policy).
+See [Subscriber § Ack policy](../usage/subscriber.md#ack-policy).
## `OutboxResponse(...)` + foreign-publisher decorator logs a configuration error { #outboxresponse-foreign-publisher-decorator-config-error }
-**Symptom.** A handler with both `@kafka_pub` and an
-`OutboxResponse(...)` return value logs an ERROR on every dispatch:
+A handler with both `@kafka_pub` and an `OutboxResponse(...)` return
+value logs an ERROR on every dispatch:
`Outbox configuration error (fix required; row left to lease-expiry retry)`.
-**Likely cause.** By design. The combination would both insert a row
-into the outbox *and* publish to Kafka, a dual-fire that doubles
-delivery. The subscriber refuses the chain composition after the
-handler returns. The worker logs the error and moves on without
-nacking, so the retry strategy is not consulted. The row's lease expires, a later fetch reclaims
-it, and the cycle repeats until the configuration is fixed.
+This is by design. The combination would both insert a row into the
+outbox *and* publish to Kafka, a dual-fire that doubles delivery. The
+subscriber refuses the chain composition after the handler returns. The
+worker logs the error and moves on without nacking, so the retry
+strategy is not consulted. The row's lease expires, a later fetch
+reclaims it, and the cycle repeats until the configuration is fixed.
-**Diagnose.** Inspect the handler decorator stack and return type.
+To confirm, inspect the handler's decorator stack and return type.
-**Fix.** Pick one path. Either `return body` plain (the foreign
+To fix it, pick one path: either `return body` plain (the foreign
publisher picks it up) or `return OutboxResponse(body, queue="...",
-session=...)` (an outbox-internal chain) but not both.
+session=...)` (an outbox-internal chain), but not both.
-**Reference.** [Relay § What not to do](../usage/relay.md#what-not-to-do),
+See [Relay § What not to do](../usage/relay.md#what-not-to-do) and
[Publisher § Chained
publishing](../usage/publisher.md#chained-publishing).
## A chained `OutboxResponse` row's handler keeps retrying after the handler "succeeded" { #outboxresponse-relay-publish-failure }
-**Symptom.** A handler that returns `OutboxResponse(...)` completes its
-own logic, yet the inbound row keeps nacking/retrying (and may DLQ as
-`retry_terminal`), with an exception that's about the *publish*, not the
-handler's work.
-
-**Likely cause.** The follow-on `OutboxResponse` row is published **after**
-the handler returns, inside the same consume scope — so a failure there
-(e.g. a DB error on the follow-on insert) unwinds through the
-`AcknowledgementMiddleware` and nacks the inbound row. There is
-**no distinct signal** separating "handler OK, relay-publish failed" from
-an ordinary handler exception: the metric reads as a normal
+A handler that returns `OutboxResponse(...)` completes its own logic,
+yet the inbound row keeps nacking and retrying (and may go to the DLQ as
+`retry_terminal`), with an exception about the *publish* and not about
+the handler's work.
+
+The follow-on `OutboxResponse` row is published after the handler
+returns, inside the same consume scope. A failure there (for example a
+DB error on the follow-on insert) unwinds through the
+`AcknowledgementMiddleware` and nacks the inbound row. No distinct
+signal separates "handler OK, relay-publish failed" from an ordinary
+handler exception: the metric reads as a normal
`nacked_retried`/`retry_terminal`, and the ERROR log shows the publish
-exception rather than a handler one.
+exception, not a handler one.
-**Diagnose.** Read the logged exception: a `sqlalchemy`/`asyncpg` error or
+To confirm, read the logged exception. A `sqlalchemy`/`asyncpg` error or
an envelope `ValueError` naming `content-type`/`correlation_id` points at
the relay publish, not the handler body.
-**Fix.** Resolve the underlying publish failure (schema/connection for the
-follow-on insert; drop conflicting headers). For non-idempotent chains,
-pass a deterministic `timer_id` so a redelivery's insert is a no-op.
+To fix it, resolve the underlying publish failure: the schema or
+connection for the follow-on insert, or conflicting headers that need
+dropping. For non-idempotent chains, pass a deterministic `timer_id` so
+a redelivery's insert is a no-op.
-**Reference.** [Publisher § Chained publishing](../usage/publisher.md#chained-publishing).
+See [Publisher § Chained publishing](../usage/publisher.md#chained-publishing).
## `validate_schema()` raises `ImportError` { #validate_schema-raises-importerror }
-**Symptom.** Calling `await broker.validate_schema()` raises:
+Calling `await broker.validate_schema()` raises:
```text
ImportError: validate_schema() requires alembic. Install with `pip install 'faststream-outbox[validate]'`.
```
-**Likely cause.** The `[validate]` extra isn't installed. Alembic is
-an optional dependency by design — every other code path works
-without it, but the schema validator delegates to Alembic's
-`autogenerate.compare_metadata` and so requires it.
+The `[validate]` extra isn't installed. Alembic is an optional
+dependency by design. Every other code path works without it, but the
+schema validator delegates to Alembic's `autogenerate.compare_metadata`
+and so requires it.
-**Diagnose.** `pip show alembic` returns nothing, or
-`pip list | grep alembic` is empty.
+To confirm, run `pip show alembic` (it returns nothing) or
+`pip list | grep alembic` (empty output).
-**Fix.** `pip install 'faststream-outbox[validate]'`. The validator
-runs unchanged after that; nothing else in the package needs to
+To fix it, run `pip install 'faststream-outbox[validate]'`. The
+validator works after that, and nothing else in the package needs to
change.
-**Reference.** [Schema validation](../usage/schema-validation.md).
+See [Schema validation](../usage/schema-validation.md).
diff --git a/docs/tutorials/add-kafka-relay.md b/docs/tutorials/add-kafka-relay.md
index eff8181..191ea24 100644
--- a/docs/tutorials/add-kafka-relay.md
+++ b/docs/tutorials/add-kafka-relay.md
@@ -4,7 +4,7 @@
In [Tutorial: Your first outbox app](./first-outbox-app.md) the handler
printed the row and that was the end of it. Real outbox systems usually
-*relay* the row to a real message bus — Kafka, RabbitMQ, NATS — so
+*relay* the row to a real message bus such as Kafka, RabbitMQ, or NATS so
downstream services can consume it. In this tutorial you'll add a Kafka
broker, stack a single decorator above the existing subscriber, and watch
a row written inside a Postgres transaction land on a Kafka topic.
@@ -16,9 +16,9 @@ relay and seen the row arrive at a `kafka-console-consumer`.
- You finished [Tutorial: Your first outbox app](./first-outbox-app.md).
This tutorial extends that same `app.py`, the same `outbox-postgres`
- container, and the same project directory. **If you ran Tutorial 1's
- final cleanup**, its `--rm` Postgres container (and its data) is gone —
- re-run Tutorial 1's Postgres-start and schema-creation steps first; this
+ container, and the same project directory. If you ran Tutorial 1's
+ final cleanup, its `--rm` Postgres container (and its data) is gone, so
+ re-run Tutorial 1's Postgres-start and schema-creation steps first. This
tutorial assumes `outbox-postgres` is up with the `outbox` table.
- Docker Compose (the `docker compose` CLI) for the Kafka container.
- Another ten minutes.
@@ -26,15 +26,15 @@ relay and seen the row arrive at a `kafka-console-consumer`.
## Step 1: Add Kafka via docker-compose
Postgres should still be running from Tutorial 1 (see the note above if you
-ran its cleanup). Add Kafka via a small `docker-compose.yml`. Single-broker
-[KRaft mode](https://kafka.apache.org/documentation/#kraft) — no separate
+ran its cleanup). Add Kafka via a small `docker-compose.yml`. It runs a single broker in
+[KRaft mode](https://kafka.apache.org/documentation/#kraft), so there is no separate
ZooKeeper service, and Confluent's `cp-kafka:7.6.0` image is known to
-run well on Apple Silicon. Two listeners: one on the host at `localhost:9092`
+run well on Apple Silicon. There are two listeners: one on the host at `localhost:9092`
(for your `faststream run` process) and one inside the Docker network at
`kafka:29092` (inter-broker traffic). The Step 5 console consumer runs
*inside* the broker container via `docker compose exec`, so it bootstraps
-against the host listener at its advertised address `localhost:9092` —
-inside the container the loopback reaches the same `0.0.0.0:9092` listener,
+against the host listener at its advertised address `localhost:9092`.
+Inside the container the loopback reaches the same `0.0.0.0:9092` listener,
so no separate in-network client listener is needed for it.
```yaml title="docker-compose.yml"
@@ -130,7 +130,7 @@ Stack `@kafka_publisher` above the existing
`@broker_outbox.subscriber("orders")` and change the handler to `return
order_id`. The stacked decorator picks up the return value and publishes
it to `orders.kafka`. The outbox subscriber is still the one driving
-delivery — Kafka becomes the *destination*, not a second subscriber.
+delivery; Kafka becomes the *destination*, not a second subscriber.
```python title="app.py (edits)"
@kafka_publisher
@@ -198,7 +198,7 @@ got order 1
The `Topic orders.kafka not found in cluster metadata` line is
`aiokafka` noticing a brand-new topic and asking the broker to
-auto-create it — first-run only.
+auto-create it. It appears on the first run only.
In a second terminal, attach a console consumer to the topic:
@@ -224,10 +224,10 @@ If Kafka were unavailable when the outbox subscriber dispatched a row,
the foreign publish would raise, the outbox row would be nacked, and
the configured `retry_strategy` would reschedule it. The next dispatch
re-runs the handler and re-attempts the foreign publish. The net effect
-is **at-least-once delivery to the foreign broker** — the outbox row is
+is at-least-once delivery to the foreign broker. The outbox row is
the durability boundary, and it stays in the table for the duration of the
retry budget (the default `ExponentialRetry` allows 10 attempts). Once the
-budget is exhausted the row is deleted — the default configures no DLQ — so
+budget is exhausted the row is deleted (the default configures no DLQ), so
configure a longer `retry_strategy` or a `dlq_table` to survive outages
beyond that (the default schedule spans roughly 8-9 minutes: nine backoffs
of 1, 2, 4, … 256 seconds sum to ~8.5 minutes before the 10th attempt is
@@ -248,20 +248,20 @@ contract in full.
- A two-broker app: an `OutboxBroker` over Postgres and a `KafkaBroker`
over a local Kafka container.
- A single subscriber whose return value is forwarded to a Kafka topic
- via a stacked publisher decorator — no second handler, no manual
- client code.
+ via a stacked publisher decorator, with no second handler and no
+ manual client code.
- An at-least-once relay: the row is durable in Postgres until the
Kafka publish succeeds.
The interesting property is the *transactional* part of the publish.
The `broker_outbox.publish(1, ...)` call in `publish_one` ran inside a
-session that committed atomically — the row reached the outbox table
+session that committed atomically: the row reached the outbox table
as part of the same `COMMIT` that any sibling domain writes would have
committed. There is no window in which the row exists but a sibling
domain write doesn't, or vice versa. The Kafka delivery happens *after*
that boundary, asynchronously, with its own retry safety net. The
-outbox is what makes those two halves — transactional domain write and
-non-transactional bus publish — survive a process crash together.
+outbox is what makes those two halves (transactional domain write and
+non-transactional bus publish) survive a process crash together.
## Clean up
@@ -275,13 +275,13 @@ stops the Postgres container from Tutorial 1.
## What's next
-- [Relay reference](../usage/relay.md) — the full contract: header
+- [Relay reference](../usage/relay.md): the full contract, including header
propagation, two-broker lifecycle, other foreign brokers
- (RabbitMQ / NATS / Redis), what *not* to do.
-- [Subscriber retry strategies](../usage/subscriber.md#retry-strategies)
- — `ExponentialRetry`, `LinearRetry`, `ConstantRetry`, `NoRetry`, and
+ (RabbitMQ / NATS / Redis), and what *not* to do.
+- [Subscriber retry strategies](../usage/subscriber.md#retry-strategies):
+ `ExponentialRetry`, `LinearRetry`, `ConstantRetry`, `NoRetry`, and
"retry only on transient errors."
-- [Comparison](../concepts/comparison.md) — see the section *"vs.
+- [Comparison](../concepts/comparison.md): see the section *"vs.
FastStream + `KafkaBroker` / `RabbitBroker` directly"* for the
pattern's trade-offs vs. just publishing to Kafka straight from
your request handler.
diff --git a/docs/tutorials/first-outbox-app.md b/docs/tutorials/first-outbox-app.md
index b6097c5..2db7d7a 100644
--- a/docs/tutorials/first-outbox-app.md
+++ b/docs/tutorials/first-outbox-app.md
@@ -3,7 +3,7 @@
## What you'll build
A tiny app where calling `broker.publish` inside a database transaction
-triggers a handler — no message bus required, just Postgres. By the end
+triggers a handler. It needs no message bus, just Postgres. By the end
you will have run a single message end-to-end and seen the handler
print it.
@@ -118,7 +118,7 @@ session_factory = async_sessionmaker(engine, expire_on_commit=False)
`make_outbox_table` returns a `sqlalchemy.Table` attached to your
`MetaData`. The package never creates or migrates the schema on its
-own — Step 4 is where we run that.
+own; Step 4 is where we run that.
## Step 4: Create the schema
@@ -158,7 +158,7 @@ Run it:
uv run python create_schema.py
```
-You should see no output — that is success.
+You should see no output, which means it succeeded.
Verify the table landed:
@@ -195,7 +195,7 @@ Check constraints:
```
Three partial indexes show up alongside the columns, plus a check
-constraint that keeps a lease either fully set or fully unset — the
+constraint that keeps a lease either fully set or fully unset. The
broker uses both at runtime; you don't need to think about them.
## Step 5: Define a handler
@@ -209,14 +209,14 @@ async def handle(order_id: int) -> None:
print(f"got order {order_id}")
```
-No command yet — the handler runs once we publish a row and start the
+There is nothing to run yet; the handler runs once we publish a row and start the
app.
## Step 6: Publish a row
Add an `@app.after_startup` hook to the bottom of `app.py` that publishes
one row right after the app boots. `broker.publish` inserts an outbox
-row through the session you give it — the row commits with the
+row through the session you give it, and the row commits with the
surrounding transaction. There is no separate "send" step; the commit
is the send.
@@ -293,7 +293,7 @@ Press `Ctrl-C`:
- An outbox table inside your own Postgres database, owned by your
schema.
-- A FastStream app whose "transport" is rows in that table — no
+- A FastStream app whose "transport" is rows in that table, with no
external broker.
- A handler that ran exactly once, in-process, against a row committed
by your own session.
@@ -301,9 +301,9 @@ Press `Ctrl-C`:
The interesting property is what happened *inside* `publish_one`: the
`broker.publish` call inserted a row into the outbox table through the
session you opened. `session.begin()` committed it. If that commit had
-rolled back — say, because a domain write on the same session
-failed — the outbox row would have rolled back with it. The row and
-the domain write commit or roll back together. That atomicity is the whole point.
+rolled back (say, because a domain write on the same session
+failed), the outbox row would have rolled back with it. The row and
+the domain write commit or roll back together.
## Clean up
@@ -313,16 +313,16 @@ docker stop outbox-postgres
## What's next
-- [Subscriber reference](../usage/subscriber.md) — tuning, worker
+- [Subscriber reference](../usage/subscriber.md): tuning, worker
counts, retry strategies.
-- [Publisher reference](../usage/publisher.md) — `publish_batch`, the
+- [Publisher reference](../usage/publisher.md): `publish_batch`, the
`OutboxPublisher` decorator, chained publishing.
-- [FastAPI integration](../usage/fastapi.md) — wire the outbox into
+- [FastAPI integration](../usage/fastapi.md): wire the outbox into
a real HTTP service with `Depends(get_session)`.
-- [Schema validation](../usage/schema-validation.md) — this tutorial
- installed the `validate` extra; call `validate_schema()` from a
+- [Schema validation](../usage/schema-validation.md): this tutorial
+ installed the `validate` extra, so you can call `validate_schema()` from a
startup hook or `/health` check to catch a table that drifted from
what the broker expects (e.g. a missing partial index after a
migration).
-- [Tutorial: Add a Kafka relay](./add-kafka-relay.md) — extend this
+- [Tutorial: Add a Kafka relay](./add-kafka-relay.md): extend this
app to forward each row into Kafka with one stacked decorator.
diff --git a/docs/usage/basic.md b/docs/usage/basic.md
index 82ddf22..b553647 100644
--- a/docs/usage/basic.md
+++ b/docs/usage/basic.md
@@ -2,7 +2,7 @@
## 1. Declare the outbox table
-The package never creates or migrates your schema — that's Alembic's job.
+The package never creates or migrates your schema; Alembic does that.
`make_outbox_table(metadata, table_name="outbox")` returns a
`sqlalchemy.Table` you attach to your own `MetaData`:
@@ -14,10 +14,10 @@ metadata = MetaData()
outbox_table = make_outbox_table(metadata, table_name="outbox")
```
-The returned `Table` carries three indexes the broker needs at runtime — a
+The returned `Table` carries three indexes the broker needs at runtime: a
partial index for the fetch CTE's unleased branch, a partial index for the
expired-lease reclaim branch, and a partial unique index for `timer_id`
-deduplication — plus a `CHECK ((acquired_token IS NULL) = (acquired_at IS
+deduplication. It also carries a `CHECK ((acquired_token IS NULL) = (acquired_at IS
NULL))` constraint that makes a half-set lease unrepresentable. Alembic
autogenerate picks them all up alongside the table itself.
@@ -50,8 +50,8 @@ via `max_workers`, tuning, retry strategies).
## 4. Publish a message
`broker.publish(body, *, queue, session, ...)` inserts an outbox row through
-the caller's `AsyncSession`. It does **not** flush, commit, or open its own
-transaction — the row commits with the caller's domain writes:
+the caller's `AsyncSession`. It does not flush, commit, or open its own
+transaction; the row commits with the caller's domain writes:
```python
from sqlalchemy.ext.asyncio import async_sessionmaker
@@ -111,8 +111,8 @@ pip install 'faststream-outbox[asyncpg]' 'faststream[cli]'
See [Installation](../introduction/installation.md) for the full extras
list and a one-line Postgres container.
-The package **declares** the outbox table but never creates it. Create it
-once before the first run — for local dev, a one-shot `metadata.create_all`
+The package declares the outbox table but never creates it. Create it
+once before the first run. For local dev, use a one-shot `metadata.create_all`
against the same `metadata`/`engine` from above:
```python
@@ -127,16 +127,16 @@ full `create_schema.py` script. Once the table exists, save the module as
## Connection ownership
-`OutboxBroker` does **not** close the `AsyncEngine` you pass in — the
+`OutboxBroker` does not close the `AsyncEngine` you pass in; the
caller owns its lifecycle. The same engine can be shared with other
SQLAlchemy users (your FastAPI app, an Alembic upgrade, etc.); closing it
from the broker would surprise them. Manage the engine with `try/finally`
-or — when running under FastAPI — let the framework's lifespan handle it
+or, when running under FastAPI, let the framework's lifespan handle it
(see [FastAPI integration](./fastapi.md)).
## Where to read next
-- [How it works](../introduction/how-it-works.md) — architecture, lease invariant, at-least-once semantics
-- [Subscriber](./subscriber.md) — tuning, retry strategies, slow-handler queue segregation
-- [Publisher](./publisher.md) — `publish_batch`, `OutboxPublisher`, chained publishing
-- [FastAPI integration](./fastapi.md) — `OutboxRouter`, `Depends(get_session)` pattern
+- [How it works](../introduction/how-it-works.md): architecture, lease invariant, at-least-once semantics
+- [Subscriber](./subscriber.md): tuning, retry strategies, slow-handler queue segregation
+- [Publisher](./publisher.md): `publish_batch`, `OutboxPublisher`, chained publishing
+- [FastAPI integration](./fastapi.md): `OutboxRouter`, `Depends(get_session)` pattern
diff --git a/docs/usage/dlq.md b/docs/usage/dlq.md
index d5cd077..854fc51 100644
--- a/docs/usage/dlq.md
+++ b/docs/usage/dlq.md
@@ -2,8 +2,8 @@
Opt-in audit for terminal failures. Pass `dlq_table=make_dlq_table(metadata)`
to the broker and every row that fails terminally is copied into the DLQ in
-the same Postgres statement as the outbox `DELETE`. Default behavior is
-unchanged when `dlq_table` is omitted — no audit table, no new code paths.
+the same Postgres statement as the outbox `DELETE`. When `dlq_table` is
+omitted, there is no audit table and none of the DLQ code paths run.
## Quickstart
@@ -25,20 +25,20 @@ engine = create_async_engine("postgresql+asyncpg://outbox:outbox@localhost:5432/
broker = OutboxBroker(engine, outbox_table=outbox_table, dlq_table=dlq_table)
```
-The package does not create or migrate the table — run `metadata.create_all`
+The package does not create or migrate the table. Run `metadata.create_all`
(or your Alembic migration) once both tables are declared. Subscribers and
publishers need no further configuration; the broker reads `dlq_table` from
its own config when it builds the terminal-flush SQL.
## What gets archived
-A row lands in the DLQ when it is **terminal-by-failure**, i.e. the
+A row lands in the DLQ when it is terminal-by-failure, i.e. the
subscriber's terminal flush would otherwise `DELETE` the row because of a
failure (not a clean ack). Three paths produce that:
| `failure_reason` | Trigger |
|---|---|
-| `max_deliveries` | `deliveries_count > max_deliveries` — handler is never invoked for this attempt. |
+| `max_deliveries` | `deliveries_count > max_deliveries`; the handler is never invoked for this attempt. |
| `retry_terminal` | Handler raised; the retry strategy returned `None` (attempts / total-delay exhausted, or `NoRetry()`). |
| `rejected` | Handler called `await msg.reject()` directly, or `AckPolicy.REJECT_ON_ERROR` rejected the row on an exception. |
@@ -55,23 +55,23 @@ behaviors that decide which reason fires.
| Column | Type | Notes |
|---|---|---|
| `id` | `BigInteger`, PK, autoincrement | DLQ row identity. |
-| `original_id` | `BigInteger`, not null | The outbox row's id, for operator forensics. Not unique — a re-delivered `timer_id` row could legitimately land here twice. |
+| `original_id` | `BigInteger`, not null | The outbox row's id, for operator forensics. Not unique: a re-delivered `timer_id` row could legitimately land here twice. |
| `queue` | `String(255)`, not null | Source queue name. |
| `payload` | `LargeBinary`, not null | Verbatim copy of the outbox payload bytes. |
| `headers` | `JSONB`, nullable | Verbatim copy, including the inherited `correlation_id`. |
| `deliveries_count` | `BigInteger`, not null | Attempt count at the moment of failure. |
-| `created_at` | `DateTime(timezone=True)`, not null | The outbox row's original `created_at` — measures time-to-terminal-failure. |
+| `created_at` | `DateTime(timezone=True)`, not null | The outbox row's original `created_at`, which measures time-to-terminal-failure. |
| `failed_at` | `DateTime(timezone=True)`, not null, default `now()` | When the audit row was written. |
| `failure_reason` | `String(64)`, not null | One of the three values in the table above. |
| `last_exception` | `String`, nullable | `repr()` of the raised exception, bounded at 8 KiB (see below). `None` on manual `reject()` without an exception. |
| `timer_id` | `String(255)`, nullable | The originating single-publish dedup key, carried into the audit trail so a terminally-failed timer keeps its business key. `None` for non-timer rows. |
-Index: `(queue, failed_at)` (btree, non-unique) — supports "show me recent
+Index: `(queue, failed_at)` (btree, non-unique). It supports "show me recent
failures for queue X" queries without a sequential scan as the DLQ grows.
No foreign key references the outbox table: the source row is gone in the
same transaction, so the constraint would be unsatisfiable. There is also no
-`LISTEN/NOTIFY` channel — nobody polls the DLQ.
+`LISTEN/NOTIFY` channel, since nobody polls the DLQ.
## Atomicity
@@ -92,37 +92,37 @@ FROM deleted;
Two operator-visible properties fall out of this shape:
-- **Lease-lost is a transparent no-op.** If another worker reclaimed the
+- A lost lease is a transparent no-op. If another worker reclaimed the
row after a lease expiry, `WHERE acquired_token = :token` matches
nothing, `deleted` is empty, the INSERT inserts zero rows, and the
- caller sees `rowcount == 0` — same observable as the no-DLQ path. The
+ caller sees `rowcount == 0`, the same observable result as the no-DLQ path. The
lease-token guard documented in [Subscriber](./subscriber.md) is
preserved.
-- **DLQ-write failure rolls back the DELETE.** If the INSERT fails
+- A failed DLQ write rolls back the DELETE. If the INSERT fails
(column mismatch, disk full, a violated constraint), the whole statement
rolls back. The outbox row stays leased and is reclaimed when the
lease expires. Misconfiguration surfaces as outbox growth plus
`lease_lost` spikes rather than silent audit loss.
-The statement runs on the worker's autocommit writer connection — one
-round-trip per terminal flush, same cost as the no-DLQ path.
+The statement runs on the worker's autocommit writer connection: one
+round-trip per terminal flush, the same cost as the no-DLQ path.
## `last_exception` truncation
The serialized exception (`repr(exc)`) is bounded at 8 KiB.
Anything longer is truncated and `…[truncated]` appended.
-Rationale: some exceptions carry MB-scale payloads — pydantic validation
-errors with the rejected request body, asyncpg `DataError` with the full
-row, etc. An unbounded `repr` would extend the writer round-trip on a
+The cap exists because some exceptions carry MB-scale payloads, such as pydantic
+validation errors with the rejected request body or asyncpg `DataError` with the
+full row. An unbounded `repr` would extend the writer round-trip on a
poison row by hundreds of milliseconds and bloat the DLQ table. 8 KiB
preserves the traceback and any structured detail while bounding worst
case.
## Redacting `last_exception` (PII / secrets)
-That same `repr` is exactly why a poison-message exception can embed a
-request body, a rejected row, or a credential — and the DLQ persists it.
+Because of that same `repr`, a poison-message exception can embed a
+request body, a rejected row, or a credential, and the DLQ persists it.
For deployments handling sensitive data, pass `last_exception_renderer` to
transform (or drop) the stored text:
@@ -142,7 +142,7 @@ The same kwarg is available on the FastAPI `OutboxRouter`.
When `dlq_table` is set, `await broker.validate_schema()` checks both
tables and surfaces missing columns / indexes on either one. The DLQ
-table is validated independently — drift in one table does not mask drift
+table is validated independently, so drift in one table does not mask drift
in the other. See [Schema validation](./schema-validation.md) for the
opt-in install + `/health` pattern.
@@ -160,12 +160,12 @@ Tags:
| `subscriber` | Subscriber handler name (`call_name`). |
| `deliveries_count` | Attempt count at terminal flush. |
| `failure_reason` | Same value set as the schema column. |
-| `exception_type` | The exception class name. **Omitted** (not set to `None`) for terminals with no exception — `max_deliveries`, or a manual `reject()` without one — so a custom recorder should treat the key as optional. |
+| `exception_type` | The exception class name. Omitted (not set to `None`) for terminals with no exception (`max_deliveries`, or a manual `reject()` without one), so a custom recorder should treat the key as optional. |
The bundled adapters surface the event without further wiring:
-- **Prometheus**: counter `faststream_outbox_dlq_written_total{reason}`.
-- **OpenTelemetry**: counter `messaging.outbox.dlq_written` with the
+- Prometheus records the counter `faststream_outbox_dlq_written_total{reason}`.
+- OpenTelemetry records the counter `messaging.outbox.dlq_written` with the
`messaging.outbox.dlq_reason` attribute and the standard
`error.type` attribute when present.
@@ -173,8 +173,8 @@ Pair with `nacked_terminal` to alert on DLQ misconfiguration: every
terminal-failure row should produce one `nacked_terminal` *and* one
`dlq_written`. A persistent divergence (terminal rate > DLQ rate) means
either the CTE keeps rolling back (DLQ schema drift) or the lease keeps
-expiring before flush (`lease_ttl_seconds` too low for handler P99) —
-both are operator-actionable signals. See
+expiring before flush (`lease_ttl_seconds` too low for handler P99).
+Both are operator-actionable signals. See
[Observability](./observability.md) for the broader recorder + middleware
story.
@@ -185,7 +185,7 @@ story.
There is no built-in pruning. Operators are responsible for archival or
expiry.
-Recommended pattern: partition the DLQ by `failed_at` (monthly or
+The recommended pattern is to partition the DLQ by `failed_at` (monthly or
weekly) and drop old partitions via a cron job. The `(queue, failed_at)`
index already supports partition pruning in operator queries; convert it
to a partitioned table at create time if you expect a steady DLQ
diff --git a/docs/usage/fastapi.md b/docs/usage/fastapi.md
index e4176aa..6001d8d 100644
--- a/docs/usage/fastapi.md
+++ b/docs/usage/fastapi.md
@@ -1,9 +1,9 @@
# FastAPI integration
-The outbox + FastAPI is the **canonical use case**: HTTP routes and outbox
+FastAPI is the canonical use case for the outbox. HTTP routes and outbox
subscribers share the same `AsyncSession` via FastAPI's dependency
-injection, and the outbox row commits with the caller's domain writes —
-same transaction, same `session.commit()`.
+injection, and the outbox row commits with the caller's domain writes in
+the same transaction and the same `session.commit()`.
`faststream_outbox.fastapi.OutboxRouter` subclasses FastStream's
`StreamRouter` (which itself subclasses FastAPI's `APIRouter`), so HTTP
@@ -82,8 +82,8 @@ app = FastAPI()
app.include_router(router)
```
-Mounting the router auto-starts the inner broker via FastAPI's lifespan —
-**you do not call `broker.start()`**. HTTP routes (`@router.get`,
+Mounting the router auto-starts the inner broker via FastAPI's lifespan,
+so you do not call `broker.start()`. HTTP routes (`@router.get`,
`@router.post`, …) and outbox subscribers coexist on one router.
## Why this works
@@ -96,7 +96,7 @@ delivery, opened in a `session.begin()` block, committed on handler return,
rolled back on exception.
A handler's `AsyncSession` is therefore resolved exactly as in an HTTP
-route — a fresh session per delivery, not a shared instance — and
+route (a fresh session per delivery, not a shared instance), and
`OutboxResponse(session=...)` commits the follow-on row with the handler's
domain writes. See [Chained
publishing](./publisher.md#chained-publishing).
@@ -123,30 +123,30 @@ dependency resolver, so `Depends(...)` and these shortcuts can be mixed
freely.
These shortcuts resolve through FastStream's subscriber-dispatch
-machinery, so they work **only inside `@router.subscriber` handlers** — not
+machinery, so they work only inside `@router.subscriber` handlers, not
in HTTP routes. In an HTTP route, reach the broker via `router.broker` (as
the quickstart's `create_order` does); a `broker: OutboxBroker` annotation
there resolves as a request field and fails with a 422.
## What's intentionally not exposed
-Several `OutboxBroker.__init__` arguments are intentionally **not exposed**
+Several `OutboxBroker.__init__` arguments are intentionally not exposed
on `OutboxRouter.__init__`:
-- `apply_types` — `StreamRouter` forces `apply_types=False` because
+- `apply_types`: `StreamRouter` forces `apply_types=False` because
FastAPI's FastDepends takes over the parameter resolution. Letting the
user flip it would produce weird half-resolved handlers.
-- `dependencies` — on the router signature this means FastAPI
+- `dependencies`: on the router signature this means FastAPI
`Depends(...)` only; the broker's FastStream `Dependant` list is the
wrong shape for this flow.
-- `routers` — not forwarded through the router; its semantics through the
+- `routers`: not forwarded through the router; its semantics through the
FastAPI lifespan are unsettled. Register subscribers directly on the
`OutboxRouter` instead.
The [DLQ](./dlq.md) and the [metrics-recorder seam](./observability.md)
-**are** available through the router: pass `dlq_table=` and
+are also available through the router: pass `dlq_table=` and
`metrics_recorder=` to `OutboxRouter(...)` exactly as you would to
-`OutboxBroker(...)` — they forward to the inner broker.
+`OutboxBroker(...)`, and they forward to the inner broker.
```python
from faststream_outbox import make_dlq_table # alongside make_outbox_table
@@ -165,7 +165,7 @@ in handlers for dependencies.
## Engine ownership
-The caller owns the `AsyncEngine`. `OutboxBroker` does **not** close it.
+The caller owns the `AsyncEngine`. `OutboxBroker` does not close it.
Dispose it in your app's lifespan:
```python
diff --git a/docs/usage/messaging-service.md b/docs/usage/messaging-service.md
index 6092556..504bc32 100644
--- a/docs/usage/messaging-service.md
+++ b/docs/usage/messaging-service.md
@@ -2,19 +2,19 @@
The [tutorials](../tutorials/first-outbox-app.md) build a greenfield app one
primitive at a time, and each guide documents one feature on its own. This page
-is different: it walks a single service that **composes** three outbox
-primitives — a transactional event relay, a fire-unless-cancelled timer, and an
+walks through a single service that composes three outbox
+primitives: a transactional event relay, a fire-unless-cancelled timer, and an
in-process test of the whole chain.
The service is a generic chat / notifications backend. Users post messages into
-chats. It has two obligations, and both must be **atomic with the database
-write** — they commit with the domain row and must never fire if the
+chats. It has two obligations, and both must be atomic with the database
+write: they commit with the domain row and must never fire if the
transaction rolls back:
-1. **Broadcast events.** Every message created, read, or deleted is published to
+1. Every message created, read, or deleted is broadcast as an event to
downstream consumers over Kafka.
-2. **Unread notifications.** If a message is still unread `N` seconds after it
- arrives, notify the recipient — *unless they read it first*.
+2. If a message is still unread `N` seconds after it arrives, an unread
+ notification goes to the recipient, *unless they read it first*.
A plain message bus can't give you "commits with the domain row": publishing to
Kafka and committing to Postgres are two systems, so a crash between them either
@@ -32,11 +32,11 @@ two obligations: `chat-events` and `unread-timers`.
ReadMessageUseCase ── cancel_timer("unread-timers", timer_id) ┘
```
-- **Use cases** write domain rows and outbox rows in one transaction.
-- **The broker** is an `OutboxBroker` over the application's `AsyncEngine`; the
+- Use cases write domain rows and outbox rows in one transaction.
+- The broker is an `OutboxBroker` over the application's `AsyncEngine`; the
outbox table lives on the app's own `MetaData` via `make_outbox_table`, so
Alembic owns its migrations.
-- **Subscribers** poll each queue and relay the row onward to Kafka.
+- Subscribers poll each queue and relay the row onward to Kafka.
```python title="tables.py"
from faststream_outbox import make_outbox_table
@@ -74,11 +74,11 @@ class Resources(Group):
)
```
-## Pattern 1 — Transactional event relay
+## Pattern 1: transactional event relay
-A thin producer wraps `broker.publish`. Note the contract: `publish` inserts the
-outbox row through the caller's `AsyncSession` but **does not flush, commit, or
-open its own transaction** — the row commits with your domain writes.
+A thin producer wraps `broker.publish`. `publish` inserts the
+outbox row through the caller's `AsyncSession` but does not flush, commit, or
+open its own transaction; the row commits with your domain writes.
```python title="producers.py"
import dataclasses
@@ -107,7 +107,7 @@ class OutboxEventProducer:
)
```
-The use case calls the producer **inside** its transaction, beside the domain
+The use case calls the producer inside its transaction, beside the domain
write, and commits once. If the commit fails, no event row exists; if it
succeeds, the event is guaranteed durable:
@@ -136,9 +136,9 @@ class CreateMessageUseCase:
```
This atomicity is load-bearing on one wiring detail: the `producer`,
-`messages_repository`, and `transaction` must all resolve the **same
-request-scoped `AsyncSession`**. `publish` inserts through whatever session the
-producer holds — if that is a different session from the one the repository
+`messages_repository`, and `transaction` must all resolve the same
+request-scoped `AsyncSession`. `publish` inserts through whatever session the
+producer holds. If that is a different session from the one the repository
writes through, the outbox row commits on its own and the "commits with the
domain row" guarantee silently breaks, with no error. Scope the session per
request in your DI container so all three share it.
@@ -172,9 +172,9 @@ Register the router on the broker (the `OutboxBroker` built in `ioc.py`) with
> single decorator over the subscriber, see
> [Relay to Kafka / RabbitMQ / NATS](relay.md).
-## Pattern 2 — Fire-unless-cancelled timer
+## Pattern 2: fire-unless-cancelled timer
-The unread notification is a **delayed** outbox row, armed in the same create
+The unread notification is a delayed outbox row, armed in the same create
transaction. `timer_id` deduplicates while a row is live (at most one live
row per `(queue, timer_id)`, not a global idempotency key); `activate_in`
defers it:
@@ -206,7 +206,7 @@ class OutboxEventProducer: # ... continued from Pattern 1
)
```
-When the recipient reads the message, a second use case **cancels** the timer in
+When the recipient reads the message, a second use case cancels the timer in
its own transaction:
```python title="use_cases.py (continued)"
@@ -228,22 +228,22 @@ class ReadMessageUseCase:
Two properties make this safe, and one is a limit worth knowing:
-- **At-most-one-live.** `timer_id` deduplicates per `(queue, timer_id)`. Arming
+- At most one timer row is live: `timer_id` deduplicates per `(queue, timer_id)`. Arming
the same id twice while a row is in flight is a no-op, so retries don't
produce two notifications.
-- **Cancel is lease-guarded.** `cancel_timer` only deletes a row that is not yet
+- Cancel is lease-guarded. `cancel_timer` only deletes a row that is not yet
being delivered (it filters on an unheld lease) and returns `False` otherwise.
-- **The race window is real.** Once the timer is leased for delivery, a read can
- no longer cancel it — the notification fires. Downstream consumers should
+- The race window is real. Once the timer is leased for delivery, a read can
+ no longer cancel it, and the notification fires. Downstream consumers should
tolerate the occasional already-read notification.
More on scheduling semantics: [Timers](timers.md).
-## Pattern 3 — Testing the composed app
+## Pattern 3: testing the composed app
-Nest `TestOutboxBroker` and `TestKafkaBroker`. In the default **sync mode**,
+Nest `TestOutboxBroker` and `TestKafkaBroker`. In the default sync mode,
`broker.publish` drives the subscriber in-process, so one call to a use case
-runs the whole chain — outbox row → relay handler → Kafka — and you assert on
+runs the whole chain (outbox row → relay handler → Kafka), and you assert on
the Kafka test broker without any background loop:
```python title="test_messaging.py"
@@ -266,12 +266,12 @@ async def test_create_message_relays_event_to_kafka(
Two caveats specific to this composition:
-- **Future-dated rows fire immediately in sync mode.** The 30-second
+- Future-dated rows fire immediately in sync mode. The 30-second
`unread-timers` row is dispatched at once, so a sync-mode test sees the
notification without waiting. To test the *delay* and the cancel race for
- real, construct `TestOutboxBroker(outbox_broker, run_loops=True)` — that runs
+ real, construct `TestOutboxBroker(outbox_broker, run_loops=True)`, which runs
the real fetch/worker loops against the in-memory store.
-- **`validate_schema()` needs a real engine.** The fake client raises
+- `validate_schema()` needs a real engine. The fake client raises
`NotImplementedError`, so put the schema check in its own test against a real
`OutboxBroker`:
@@ -287,9 +287,9 @@ More on the test broker's two modes: [Testing](testing.md).
## See also
-- [Relay to Kafka / RabbitMQ / NATS](relay.md) — the native relay
+- [Relay to Kafka / RabbitMQ / NATS](relay.md): the native relay
decorator, an alternative to the hand-rolled hop above.
-- [Dead-letter queue](dlq.md) — archive terminal failures instead of
+- [Dead-letter queue](dlq.md): archive terminal failures instead of
deleting them.
-- [Observability](observability.md) — the metrics recorder and the
+- [Observability](observability.md): the metrics recorder and the
Prometheus / OpenTelemetry middleware.
diff --git a/docs/usage/observability.md b/docs/usage/observability.md
index 357f0a6..375df3b 100644
--- a/docs/usage/observability.md
+++ b/docs/usage/observability.md
@@ -3,7 +3,7 @@
*Setting it up: [Setup Prometheus and OpenTelemetry](./setup-prometheus-opentelemetry.md).
Why two seams: [Concepts § Instrumentation seams](../concepts/instrumentation-seams.md).*
-This page is the **Reference**: the recorder-seam API, the event
+This page is the reference for the recorder-seam API, the event
catalog, and the operator PromQL playbook.
## The recorder seam
@@ -57,16 +57,16 @@ broker = OutboxBroker(
```
`OpenTelemetryRecorder` (`faststream_outbox.metrics.opentelemetry`) is the
-OTel equivalent. Full wiring — including running the recorder seam and the
-native middleware together — is in
+OTel equivalent. Full wiring, including running the recorder seam and the
+native middleware together, is in
[Setup Prometheus and OpenTelemetry](./setup-prometheus-opentelemetry.md).
### Recorder must not block
-The recorder is called from the event loop. **Do not block in it.**
+The recorder is called from the event loop. Do not block in it.
Synchronous `prometheus_client.Counter.inc()` is fine (microseconds); a
blocking HTTP / StatsD call is not. The library does not wrap recorders in
-`asyncio.to_thread` — that would destroy ordering and explode the task
+`asyncio.to_thread`, because that would destroy ordering and explode the task
graph.
Every call site wraps the recorder in `try/except` and logs at DEBUG, so a
@@ -76,15 +76,15 @@ broken recorder never poisons the dispatch loop.
| Event | Tags (always present) | Tags (situational) | Fired by |
|---|---|---|---|
-| `fetched` | `queue`, `subscriber`, `count` | | Fetch loop, once per fetch attempt (`count=0` on an empty fetch) — **skipped** when the in-flight queue is full (no fetch is issued). `queue` is tagged with the subscriber's **first** queue only; multi-queue subscribers should break down by queue using the row-level events instead |
+| `fetched` | `queue`, `subscriber`, `count` | | Fetch loop, once per fetch attempt (`count=0` on an empty fetch). Skipped when the in-flight queue is full (no fetch is issued). `queue` is tagged with the subscriber's first queue only; multi-queue subscribers should break down by queue using the row-level events instead |
| `dispatched` | `queue`, `subscriber`, `deliveries_count`, `size_bytes` | | Worker loop, before handler runs |
| `acked` | `queue`, `subscriber`, `deliveries_count`, `duration_seconds` | | Handler returned successfully |
| `nacked_retried` | `queue`, `subscriber`, `deliveries_count`, `duration_seconds`, `next_delay_seconds` | `exception_type` | Retry scheduled |
| `nacked_terminal` | `queue`, `subscriber`, `deliveries_count`, `reason` | `duration_seconds`, `exception_type` | Row terminally failed (`duration_seconds` absent for `max_deliveries`, which never ran the handler) |
| `lease_lost` | `queue`, `subscriber`, `phase`, `row_id`, `deliveries_count` | | Terminal or retry write found `rowcount == 0` (`phase` = `terminal` \| `retry`) |
| `published` | `queue`, `status`, `count`, `size_bytes`, `duration_seconds` | `exception_type` | Producer, after the INSERT executes (pre-commit; also fires on error with `status="error"`) |
-| `dlq_written` | `queue`, `subscriber`, `deliveries_count`, `failure_reason` | `exception_type` | DLQ CTE wrote an audit row. `exception_type` is **omitted** — not set to `None` — when the terminal had no exception (`max_deliveries`, or a manual `reject()` without one) |
-| `drain_timeout` | `queue`, `subscriber`, `drain_timeout_seconds` | | A `stop()` drain exceeded `graceful_timeout`; in-flight rows were abandoned to lease-expiry retry. `queue` is the subscriber's **first** queue |
+| `dlq_written` | `queue`, `subscriber`, `deliveries_count`, `failure_reason` | `exception_type` | DLQ CTE wrote an audit row. `exception_type` is omitted (not set to `None`) when the terminal had no exception (`max_deliveries`, or a manual `reject()` without one) |
+| `drain_timeout` | `queue`, `subscriber`, `drain_timeout_seconds` | | A `stop()` drain exceeded `graceful_timeout`; in-flight rows were abandoned to lease-expiry retry. `queue` is the subscriber's first queue |
`reason` on `nacked_terminal` is one of `max_deliveries`,
`retry_terminal`, `rejected`. The same value lands in the DLQ
@@ -95,9 +95,9 @@ broken recorder never poisons the dispatch loop.
Operator queries that key off the recorder-side metrics. The
`faststream_outbox_*` series below (`_lease_lost_total`,
`_terminal_total`, `_dlq_written_total`) are emitted by
-**`PrometheusRecorder`** (`faststream_outbox.metrics.prometheus`), wired
-via `metrics_recorder=…` — see [Setup](./setup-prometheus-opentelemetry.md);
-the native `OutboxPrometheusMiddleware` does **not** emit them. The
+`PrometheusRecorder` (`faststream_outbox.metrics.prometheus`), wired
+via `metrics_recorder=…` (see [Setup](./setup-prometheus-opentelemetry.md)).
+The native `OutboxPrometheusMiddleware` does not emit them. The
`broker` label is always `"outbox"`; add the filter to disambiguate from
upstream FastStream services.
@@ -148,8 +148,8 @@ dlq_written](./dlq.md#metric-dlq_written).
`TestOutboxBroker` replaces `broker.publish` with a patched version that
skips the middleware publish path, so middleware-registered
-**publish-scope** metrics do **not** fire in test mode. Middleware
-**consume-scope** metrics still fire, because handlers are still consumed
+publish-scope metrics do not fire in test mode. Middleware
+consume-scope metrics still fire, because handlers are still consumed
through the normal middleware stack.
The recorder-seam `published` event provides synthetic publish-side
@@ -157,4 +157,4 @@ coverage in test mode via `FakeOutboxProducer`. The synthetic events use
`duration_seconds=0.0` since the in-memory client has no real write to
time.
-Mirrors `TestKafkaBroker` / `TestRabbitBroker` — same posture, same reason.
+`TestKafkaBroker` and `TestRabbitBroker` behave the same way, for the same reason.
diff --git a/docs/usage/publisher.md b/docs/usage/publisher.md
index b56db6b..3060e16 100644
--- a/docs/usage/publisher.md
+++ b/docs/usage/publisher.md
@@ -2,13 +2,13 @@
There are three ways to write an outbox row:
-1. **`broker.publish(...)`** — inline call, one row.
-2. **`broker.publish_batch(...)`** — inline call, many rows in one INSERT.
-3. **`broker.publisher(queue, ...)`** — a typed, queue-scoped wrapper for
+1. `broker.publish(...)`: inline call, one row.
+2. `broker.publish_batch(...)`: inline call, many rows in one INSERT.
+3. `broker.publisher(queue, ...)`: a typed, queue-scoped wrapper for
per-queue config and AsyncAPI spec coverage.
All three share the same transactional contract: the caller supplies an
-`AsyncSession`, and the row commits with the caller's domain writes — the
+`AsyncSession`, and the row commits with the caller's domain writes. The
broker does not flush, commit, or open its own transaction.
For "consume from A → enqueue to B" relay flows, a fourth path is
@@ -80,15 +80,15 @@ await broker.publish_batch(
) -> None
```
-`publish_batch` returns nothing and does **not** accept `timer_id` —
-per-row dedup makes no sense in a batch. It also accepts `activate_in` /
+`publish_batch` returns nothing and does not accept `timer_id`,
+because per-row dedup makes no sense in a batch. It also accepts `activate_in` /
`activate_at` to schedule every row in the batch identically; the schedule
is applied client-side rather than server-side (a few-ms drift vs. the
single-`publish` path).
## `broker.publisher(queue, ...)`
-`broker.publisher(queue, ...)` returns an `OutboxPublisher` — a typed,
+`broker.publisher(queue, ...)` returns an `OutboxPublisher`: a typed,
queue-scoped wrapper around `broker.publish` with the same transactional
contract:
@@ -120,12 +120,12 @@ broker.publisher(
```
The publisher exists primarily for AsyncAPI spec coverage and to
-encapsulate per-queue config — hence the `title` / `description` / `schema`
-/ `include_in_schema` knobs above, alongside the static `headers`.
+encapsulate per-queue config, which is why it has the `title` / `description` / `schema`
+/ `include_in_schema` knobs above alongside the static `headers`.
### Not a relay decorator
-It is **standalone-only**: stacking it as a relay decorator on a
+It is standalone-only: stacking it as a relay decorator on a
subscriber (`@orders_pub @broker.subscriber("inbox", ...)`) raises
`NotImplementedError` at decoration time, because the dispatch loop has
no reachable `AsyncSession` without breaking the outbox transactional
@@ -133,7 +133,7 @@ contract.
For "consume from queue A → enqueue to queue B" relays, either call
`broker.publish(value, queue="B", session=session)` directly inside your
-handler — on the same session that holds your domain writes — or
+handler, on the same session that holds your domain writes, or
`return OutboxResponse(...)` (see below). (The inbound row's own terminal
DELETE runs separately, on the worker's autocommit connection, not this
session.)
@@ -148,15 +148,15 @@ transactional contract applies (you provide the session, the row commits
with your domain writes):
!!! note "The `session` must outlive the handler return"
- The returned `OutboxResponse` is published **after** the handler
+ The returned `OutboxResponse` is published after the handler
returns, so its `session` must still be open at that point. The
requirement is about session *lifetime*, not any particular framework:
provide the session through a dependency that the framework tears down
- *after* the response flow — FastAPI's `Depends(get_session)` or
+ *after* the response flow. FastAPI's `Depends(get_session)` or
FastStream's own `Depends` / `Context` session both do this. Opening
your own `async with session_factory() as session:` inside the handler
- does **not** work here: that session closes on `return`, before the row
- is inserted — in that case call `broker.publish(..., session=session)`
+ does not work here: that session closes on `return`, before the row
+ is inserted. In that case, call `broker.publish(..., session=session)`
directly inside the `async with` instead (see
[§ Not a relay decorator](#not-a-relay-decorator)).
@@ -186,23 +186,23 @@ async def handle(
```
`correlation_id` propagates from the inbound message if you don't set one
-explicitly — useful for trace stitching. Plain returns (`None`, `dict`,
+explicitly, which is useful for trace stitching. Plain returns (`None`, `dict`,
etc.) are silently skipped, so handlers that don't want to chain just
return normally.
!!! warning "Duplicate delivery on crash"
The chained `downstream` row commits with the handler's transaction,
- but the inbound `orders` row's terminal `DELETE` runs **after** the
+ but the inbound `orders` row's terminal `DELETE` runs after the
handler returns, on the worker's separate autocommit connection. A
crash between those two points leaves the inbound row undeleted, so it
- is redelivered — producing a **second** chained row. For non-idempotent
+ is redelivered and produces a second chained row. For non-idempotent
chains, pass a deterministic `timer_id` derived from the inbound message
so the duplicate insert is a no-op (see [Timers](./timers.md)).
## Annotated handler params
`faststream_outbox.annotations` exports `Annotated[..., Context(...)]`
-shortcuts for the broker, producer, and client — useful when you want to
+shortcuts for the broker, producer, and client, useful when you want to
publish from inside a handler:
```python
@@ -215,7 +215,7 @@ async def handle(msg: OutboxMessage, broker: OutboxBroker) -> None:
await broker.publish({"chained": True}, queue="downstream", session=session)
```
-For FastAPI handlers, import the same names from `faststream_outbox.fastapi`
-— they resolve via the same `Context()` paths but go through FastAPI's
-dependency resolver so `Depends(...)` and these shortcuts can be mixed
+For FastAPI handlers, import the same names from `faststream_outbox.fastapi`.
+They resolve via the same `Context()` paths but go through FastAPI's
+dependency resolver, so `Depends(...)` and these shortcuts can be mixed
freely. See [FastAPI integration](./fastapi.md).
diff --git a/docs/usage/relay.md b/docs/usage/relay.md
index ea90b0b..eeeced8 100644
--- a/docs/usage/relay.md
+++ b/docs/usage/relay.md
@@ -3,20 +3,20 @@
> Want a worked end-to-end example? See
> [Tutorial: Add a Kafka relay](../tutorials/add-kafka-relay.md).
-The outbox pattern's payoff line: domain code writes a row to the outbox in
-the same DB transaction as its other writes, and a separate worker relays
-those rows to a real bus (Kafka, RabbitMQ, NATS, Redis…). `faststream-outbox`
-supports this directly via FastStream's cross-broker chain — stack a
-foreign-broker publisher decorator on an outbox subscriber and you're done.
+In the outbox pattern, domain code writes a row to the outbox in the same
+DB transaction as its other writes, and a separate worker relays those rows
+to a real bus (Kafka, RabbitMQ, NATS, Redis…). `faststream-outbox` supports
+this directly via FastStream's cross-broker chain: stack a foreign-broker
+publisher decorator on an outbox subscriber.
*If you don't have a database write to atomically commit alongside, use
-the foreign broker directly — see
+the foreign broker directly. See
[Comparison](../concepts/comparison.md).*
## Why an outbox relay
When a request must (a) update your database and (b) emit an event onto a
-message bus, the naive shape — DB commit, then bus publish — leaks events on
+message bus, the naive shape (DB commit, then bus publish) leaks events on
crashes between the two steps. The outbox pattern fixes this by writing the
event as a row in the same transaction as the domain update; a separate
worker reads the row and publishes to the bus. The row is the durability
@@ -43,7 +43,7 @@ async def relay(body: dict) -> dict:
return body
```
-That's the whole thing. `await broker_outbox.publish(body, queue="outbox_queue", session=session)`
+`await broker_outbox.publish(body, queue="outbox_queue", session=session)`
in your domain transaction writes a row; the subscriber dispatches it; the
handler returns it; the Kafka publisher decorator picks it up and publishes
to `kafka_topic`. Failure handling, retries, and DLQ are unchanged from
@@ -51,11 +51,11 @@ the rest of the outbox subscriber's behavior.
## Two-broker lifecycle
-Both brokers must be started for the relay to work. There's a built-in
-safety net: at `start()` the outbox broker logs a WARNING (one per unstarted
-foreign broker) naming the affected queue(s), and a relay to an unstarted
-foreign broker simply fails-and-retries until that broker is started — the
-row is never lost. Two idiomatic shapes:
+Both brokers must be started for the relay to work. As a safety net, at
+`start()` the outbox broker logs a WARNING (one per unstarted foreign broker)
+naming the affected queue(s), and a relay to an unstarted foreign broker
+fails and retries until that broker is started, so the row is never lost.
+There are two idiomatic shapes:
### FastAPI (recommended)
@@ -114,7 +114,7 @@ If the foreign publish raises (Kafka down, partition unavailable, etc.),
the exception propagates through FastStream's `AcknowledgementMiddleware`,
the outbox row is nacked, and the configured `retry_strategy` reschedules
it. The next dispatch re-runs the handler and re-attempts the foreign
-publish. **Net effect: at-least-once delivery to the foreign broker.**
+publish. The net effect is at-least-once delivery to the foreign broker.
Downstream consumers should handle duplicates idempotently, the same way
they would behind any at-least-once bus.
@@ -123,9 +123,9 @@ they would behind any at-least-once bus.
By default, FastStream's `Response(value)` ships with empty headers, so
the inbound outbox row's headers (`content-type`, custom trace keys, etc.)
-are **not** forwarded to the foreign publish. Two ways to override:
+are not forwarded to the foreign publish. Two ways to override:
-**Explicit (per handler):**
+Set them explicitly, per handler:
```python
from faststream.response import Response
@@ -138,7 +138,7 @@ async def relay(body: dict, msg: OutboxMessage) -> Response:
return Response(body, headers=msg.headers)
```
-**Opt-in (per subscriber):**
+Or opt in per subscriber:
```python
@publisher_kafka
@@ -153,7 +153,7 @@ from the inbound `OutboxMessage.headers` *unless* the handler returned a
## Using routers
-Both halves of the chain can live on routers — the FastAPI shape above
+Both halves of the chain can live on routers. The FastAPI shape above
already does this with `KafkaRouter` and `OutboxRouter`. The constraint is
that `broker.include_router(router)` must happen *before* the brokers
start. Inside `FastAPI(..., lifespan=...)` the include happens during app
@@ -186,7 +186,7 @@ broker_outbox.include_router(outbox_router)
## What not to do
-**Do not** combine `OutboxResponse(...)` and a foreign-publisher decorator.
+Do not combine `OutboxResponse(...)` and a foreign-publisher decorator:
```python
from faststream_outbox import OutboxResponse
@@ -200,13 +200,13 @@ async def relay(body: dict) -> OutboxResponse:
This would both insert a row into the outbox AND publish to Kafka. The
subscriber raises `RuntimeError` at dispatch time when it detects the
-combination — pick one path. The worker catches that error and logs it at
-ERROR; it does **not** flush a nack and does **not** route the row through
-the `retry_strategy`. The row's lease simply expires and a later fetch
+combination, so pick one path. The worker catches that error and logs it at
+ERROR; it does not flush a nack and does not route the row through
+the `retry_strategy`. The row's lease expires and a later fetch
reclaims it, so the row keeps being retried (not lost) until you fix the
configuration.
-**Do not** stack an outbox publisher on a foreign subscriber.
+Do not stack an outbox publisher on a foreign subscriber:
```python
@broker_outbox.publisher("outbox_queue") # NotImplementedError at decoration
@@ -216,15 +216,15 @@ async def relay(body: dict) -> dict:
```
This direction would need the Kafka subscriber's dispatch loop to provide
-an `AsyncSession` for the outbox insert — there isn't one without breaking
+an `AsyncSession` for the outbox insert, and there isn't one without breaking
the transactional contract. `OutboxPublisher.__call__` raises
`NotImplementedError` at decoration time. Call `await broker_outbox.publish(...)`
inside the handler instead, on a session you opened yourself.
## Other foreign brokers
-The same pattern works for Confluent, RabbitMQ, NATS, and Redis — the only
-change is the `publisher` line:
+The same pattern works for Confluent, RabbitMQ, NATS, and Redis; only the
+`publisher` line changes:
| Foreign broker | Publisher line |
|---|---|
@@ -235,5 +235,5 @@ change is the `publisher` line:
| Redis | `broker_redis.publisher("channel")` |
Any FastStream broker whose publisher's `_publish` accepts a generic
-`PublishCommand` works as a relay destination — that is the FastStream
+`PublishCommand` works as a relay destination. That is the FastStream
cross-broker contract, not an outbox-specific feature.
diff --git a/docs/usage/router.md b/docs/usage/router.md
index c3c1ccb..e25e14b 100644
--- a/docs/usage/router.md
+++ b/docs/usage/router.md
@@ -21,7 +21,7 @@ async def handle_order(order_id: int) -> None:
## No `prefix`
-Unlike some FastStream routers, `OutboxRouter` does **not** accept a
+Unlike some FastStream routers, `OutboxRouter` does not accept a
`prefix` argument. Queues are routed by their literal name, so producers
and consumers must agree on the exact string. If you want namespacing
(e.g., one Postgres instance shared across services), put it in the queue
@@ -32,7 +32,7 @@ name itself:
async def handle_order(...): ...
```
-The reason is simple: the outbox row's `queue` column is what the fetch
+Prefixes are left out because the outbox row's `queue` column is what the fetch
CTE filters on, and adding an implicit prefix would mean producers need to
know which router published the subscriber. Explicit queue names keep that
contract local.
@@ -73,7 +73,7 @@ router = OutboxRouter(
All `@broker.subscriber` options (`max_workers`, `retry_strategy`,
`fetch_batch_size`, `lease_ttl_seconds`, `max_deliveries`, `ack_policy`,
-…) are accepted by `OutboxRoute` and `router.subscriber` — see the
+…) are accepted by `OutboxRoute` and `router.subscriber`. See the
[subscriber page](./subscriber.md) for the full list.
## Gotcha: walking every subscriber
diff --git a/docs/usage/schema-validation.md b/docs/usage/schema-validation.md
index d8d623b..4fb30af 100644
--- a/docs/usage/schema-validation.md
+++ b/docs/usage/schema-validation.md
@@ -1,7 +1,7 @@
# Schema validation
-The package never creates or migrates your schema — that's Alembic's job
-— but it does provide an opt-in helper that verifies the live table has
+The package never creates or migrates your schema (that's Alembic's job),
+but it does provide an opt-in helper that verifies the live table has
everything the broker needs at runtime.
`broker.validate_schema()` delegates to Alembic's
@@ -18,7 +18,7 @@ predicate or lease check constraint.
## Install
-Alembic is an **optional dependency**:
+Alembic is an optional dependency:
```bash
pip install 'faststream-outbox[validate]'
@@ -34,7 +34,7 @@ works without it.
await broker.validate_schema()
```
-Raises `RuntimeError` if the live table is missing what the broker needs —
+Raises `RuntimeError` if the live table is missing what the broker needs:
absent table, missing columns, mismatched column types, flipped
nullability, missing partial indexes.
@@ -49,33 +49,33 @@ default (`False`) skips this check.
await broker.validate_schema(check_autovacuum=True)
```
-Extras are intentionally ignored: the validator only flags **missing**
+Extras are intentionally ignored: the validator only flags missing
schema (`add_*` / `modify_*` ops). `remove_*` ops are silently dropped so
you can attach your own audit columns or additional indexes without the
validator complaining.
-Some drift cannot be fixed by re-running `alembic revision --autogenerate` — a
+Some drift cannot be fixed by re-running `alembic revision --autogenerate`: a
missing/altered `outbox_lease_ck` CHECK or a drifted partial-index predicate.
For those, the `RuntimeError` ends with a pointer to
[Alembic migrations § Fixing drift autogenerate can't see](../operations/alembic.md#fixing-drift-autogenerate-cant-see),
which holds the hand-written migration recipe.
!!! warning "Server defaults are not checked"
- The diff runs with `compare_server_default=False` — Alembic's
+ The diff runs with `compare_server_default=False`. Alembic's
server-default comparison is flaky against Postgres' normalized
expressions (`now()` vs `CURRENT_TIMESTAMP`), so it is disabled to avoid
- false positives. A **green** `validate_schema()` therefore does **not**
+ false positives. A green `validate_schema()` therefore does not
prove your server defaults exist. The load-bearing case: a table missing
`server_default=now()` on `next_attempt_at` leaves fresh rows with NULL
`next_attempt_at`, which the fetch CTE's `next_attempt_at <= now()`
- predicate silently filters out — a silent broker outage that validation
- will not catch. Generate your migration from `make_outbox_table(...)` so
+ predicate silently filters out. The result is a broker outage that
+ validation will not catch. Generate your migration from `make_outbox_table(...)` so
the defaults are in place to begin with.
## Where to call it
-Call it from a `/health` endpoint or startup hook — **not** at
-`broker.start()`. The reason: if `validate_schema()` ran at startup and
+Call it from a `/health` endpoint or startup hook, not at
+`broker.start()`. If `validate_schema()` ran at startup and
your migration hadn't been applied yet, the broker would crash-loop
itself. Operators need to be able to roll out a new schema version and
have Alembic catch up against the same DB without a startup loop.
@@ -121,7 +121,7 @@ asyncio.run(main())
## In tests
-`FakeOutboxClient.validate_schema()` raises `NotImplementedError` — there
+`FakeOutboxClient.validate_schema()` raises `NotImplementedError`: there
is no real DB to validate against, and a silent pass would let users ship
broken schemas while their `TestOutboxBroker`-backed tests stay green.
diff --git a/docs/usage/setup-prometheus-opentelemetry.md b/docs/usage/setup-prometheus-opentelemetry.md
index 291626f..13affcb 100644
--- a/docs/usage/setup-prometheus-opentelemetry.md
+++ b/docs/usage/setup-prometheus-opentelemetry.md
@@ -1,6 +1,6 @@
# Setup Prometheus and OpenTelemetry
-You've decided to wire metrics. This page is the recipe. For the *why
+This page is the recipe for wiring metrics. For the *why
two instrumentation seams*, see [Concepts § Instrumentation
seams](../concepts/instrumentation-seams.md); for the event catalog
and operator PromQL playbook, see [Reference §
@@ -54,10 +54,10 @@ app = AsgiFastStream(
`AsgiFastStream` accepts any ASGI sub-app under `asgi_routes`; mount
`make_asgi_app(REGISTRY)` to expose Prometheus exposition without
pulling FastAPI in. `make_ping_asgi(broker)` is FastStream's built-in
-liveness probe — handy for Kubernetes.
+liveness probe, handy for Kubernetes.
-The `broker` label is always `"outbox"`; existing FastStream Grafana
-dashboards keep working — add `broker="outbox"` to the PromQL filter.
+The `broker` label is always `"outbox"`. Existing FastStream Grafana
+dashboards keep working; add `broker="outbox"` to the PromQL filter.
### Consume vs publish label set
@@ -72,8 +72,8 @@ operator query catalog.
## OpenTelemetry adapter
-Drop-in compatible with FastStream's `TelemetryMiddleware`, **meter
-only — no spans** (use the [native middleware](#native-middleware-spans--bus-parity)
+Drop-in compatible with FastStream's `TelemetryMiddleware`, but it records
+meters only, no spans (use the [native middleware](#native-middleware-spans--bus-parity)
section below if you need spans).
```bash
@@ -124,10 +124,10 @@ swap the reader for `PeriodicExportingMetricReader(OTLPMetricExporter(...))`
and drop the `/metrics` route.
Instrument names match `faststream.opentelemetry.TelemetryMiddleware` for
-the bus-scope metrics — `messaging.process.duration`,
+the bus-scope metrics: `messaging.process.duration`,
`messaging.publish.duration`, and (when `include_messages_counters=True`)
-`messaging.process.messages` / `messaging.publish.messages` — plus four
-outbox-specific counters the middleware can't emit:
+`messaging.process.messages` / `messaging.publish.messages`. The adapter
+adds four outbox-specific counters the middleware can't emit:
`messaging.outbox.fetch.batches`, `messaging.outbox.lease_lost`,
`messaging.outbox.dlq_written`, and `messaging.outbox.drain_timeout`. Units
and constructor args
@@ -135,7 +135,7 @@ and constructor args
The `messaging.system="outbox"` attribute disambiguates outbox traffic
from Kafka / Rabbit data on the same instruments.
-**Tracing (spans) is not modelled by this adapter** — the callable
+This adapter does not model tracing (spans), because the callable
seam can't bracket a span lifecycle. For spans, use the [native
middleware](#native-middleware-spans--bus-parity) integration below.
@@ -143,14 +143,14 @@ middleware](#native-middleware-spans--bus-parity) integration below.
For OTel spans wrapping `consume_scope` / `publish_scope` and the
exact upstream label / instrument schema, register the native
-middleware subclasses via `middlewares=[...]` — same
+middleware subclasses via `middlewares=[...]`, using the same
registration pattern as `KafkaPrometheusMiddleware` /
`RabbitTelemetryMiddleware`.
## Both seams together { #both-seams-together }
The recommended setup pairs middleware with the recorder so every
-event the bus emits **and** every outbox-internal event lands in one
+event the bus emits and every outbox-internal event lands in one
observability stack:
```bash
@@ -158,8 +158,8 @@ pip install 'faststream-outbox[opentelemetry,prometheus]' \
opentelemetry-exporter-otlp uvicorn
```
-Here OpenTelemetry supplies **spans** (exported to OTLP) and Prometheus
-supplies **all metrics** (two registries scraped over HTTP). The OTel
+Here OpenTelemetry supplies spans (exported to OTLP) and Prometheus
+supplies all metrics (two registries scraped over HTTP). The OTel
middleware runs span-only: it gets no `meter_provider`, so its meters go to
the global OpenTelemetry meter provider, which is a no-op unless you set one.
Neither endpoint below exposes them.
@@ -228,15 +228,15 @@ app = AsgiFastStream(
```
Traces flow to OTLP (Jaeger / Tempo / Honeycomb / collector); the
-Prometheus **middleware's** consume/publish meters land on `/metrics` and the
-**recorder's** outbox-internal counters on `/metrics/outbox` for Prometheus to
-scrape — two scrape targets, one process. (The OTel middleware here
+Prometheus middleware's consume/publish meters land on `/metrics` and the
+recorder's outbox-internal counters on `/metrics/outbox` for Prometheus to
+scrape: two scrape targets, one process. (The OTel middleware here
contributes spans only, per the note above.)
-**The two seams overlap on consume/publish series.** Both the middleware
+The two seams overlap on consume/publish series. Both the middleware
and the recorder emit the same `faststream_received_*` / `faststream_published_*`
-collectors, which is why they must live on **separate registries** (above) —
-sharing one raises `Duplicated timeseries in CollectorRegistry` as soon as
+collectors, which is why they must live on separate registries (above).
+Sharing one raises `Duplicated timeseries in CollectorRegistry` as soon as
the second of them is created, and summing across both double-counts every consume and
publish. Treat the middleware as the source of truth for consume/publish;
the recorder's unique value is the outbox-internal events the middleware
diff --git a/docs/usage/subscriber.md b/docs/usage/subscriber.md
index a8683f7..d61bd7d 100644
--- a/docs/usage/subscriber.md
+++ b/docs/usage/subscriber.md
@@ -28,14 +28,14 @@ async def handle(body: dict) -> None: ...
```
The subscriber claims rows from any of its queues in a single fetch. Its
-[connection budget](#connection-budget) is unchanged — `max_workers + 1`
+[connection budget](#connection-budget) is unchanged: `max_workers + 1`
pool connections regardless of how many queues it serves.
In the AsyncAPI document it appears as one channel per queue
(`orders:Handle`, `refunds:Handle`), each addressed by that queue, rather
than one channel for the subscriber.
-Do **not** register two subscribers on the **same** queue: they compete for
+Do not register two subscribers on the same queue: they compete for
the same rows, and registration emits a warning to that effect. To run more
than one handler over a queue, attach them to a single subscriber; to scale
throughput, raise `max_workers`.
@@ -75,7 +75,7 @@ async def handle(msg: OutboxMessage, broker: OutboxBroker) -> None: ...
`OutboxMessage`, `OutboxBroker`, `OutboxProducer`, and `OutboxClient` are
all available. For FastAPI handlers, import the same names from
-`faststream_outbox.fastapi` — they resolve via the same `Context()` paths
+`faststream_outbox.fastapi`. They resolve via the same `Context()` paths
but go through FastAPI's dependency resolver so `Depends(...)` and these
shortcuts can be mixed freely.
@@ -89,7 +89,7 @@ Per-subscriber knobs, passed to `@broker.subscriber("…", …)`:
| `fetch_batch_size` | `10` | Rows claimed per fetch cycle |
| `min_fetch_interval` | `1.0` s | Base for the adaptive idle backoff (jittered ±50%, so an actual wait can land below it) and the wait when the inflight queue is full; no sleep at all while fetches keep returning rows |
| `max_fetch_interval` | `10.0` s | Ceiling for the adaptive idle backoff (with jitter) |
-| `lease_ttl_seconds` | `60.0` s | How long a claim is valid before another fetch may reclaim it. **Must exceed your handler's P99 with margin.** |
+| `lease_ttl_seconds` | `60.0` s | How long a claim is valid before another fetch may reclaim it. Must exceed your handler's P99 with margin. |
| `max_deliveries` | `None` (unbounded) | Total claims (including lease-expiry re-claims) after which the row is dropped without invoking the handler. Defends against handlers that consistently wedge. |
| `terminal_flush_batch_size` | `1` (off) | Coalesce completed terminal `DELETE`s into one `DELETE … RETURNING` per N rows. `1` is one round-trip per message (unchanged). Higher trades a wider crash-redelivery window for far fewer round-trips. See [Batching terminal deletes](#batching-terminal-deletes). |
| `ack_policy` | `AckPolicy.NACK_ON_ERROR` | See [Ack policy](#ack-policy) |
@@ -111,7 +111,7 @@ async def handle_urgent(body: dict) -> None: ...
Subscriber options are validated when the subscriber is created: likely-wrong
combinations (`lease_ttl_seconds <= max_fetch_interval`, `max_deliveries`
without retry, `min_fetch_interval > max_fetch_interval`, etc.) warn or
-raise. **Every** construction path (`@broker.subscriber`, `@router.subscriber`,
+raise. Every construction path (`@broker.subscriber`, `@router.subscriber`,
`OutboxRoute`) is checked.
The table above lists the outbox-specific knobs. The standard FastStream
@@ -122,12 +122,12 @@ subscriber kwargs pass through unchanged too: `dependencies`, `parser`,
spanning several queues it prefixes each channel (`Ingest:orders`,
`Ingest:refunds`), since one title cannot name several channels on its own.
-## Slow handlers — dedicated queue
+## Slow handlers: dedicated queue
When a handler's tail latency exceeds the subscriber's `lease_ttl_seconds`,
-the row's lease expires mid-flight and another fetch reclaims it →
-duplicate delivery. Don't hike `lease_ttl_seconds` globally — that delays
-reclaim of *actually* stuck rows everywhere. Instead, segregate slow work
+the row's lease expires mid-flight and another fetch reclaims it,
+causing duplicate delivery. Don't hike `lease_ttl_seconds` globally, because
+that delays reclaim of *actually* stuck rows everywhere. Instead, segregate slow work
onto its own subscriber with a longer TTL:
```python
@@ -148,7 +148,7 @@ to the appropriate queue at `publish` time.
!!! note "Account for queue depth, not just per-row latency"
A fetch claims up to `fetch_batch_size` rows at once and each takes its
lease at fetch time, then they wait their turn in an in-memory queue drained
- by `max_workers` handlers. A row's lease clock runs **while it waits**, so the
+ by `max_workers` handlers. A row's lease clock runs while it waits, so the
relevant bound is the *serialized* time to reach it, not one handler's P99:
```
@@ -165,7 +165,7 @@ to the appropriate queue at `publish` time.
## Batching terminal deletes
-By default each processed row is deleted with its own `DELETE` — one round-trip
+By default each processed row is deleted with its own `DELETE`, one round-trip
per message. At `max_workers=1` those deletes serialise, and the round-trip
(not the database work) is the throughput ceiling. Set
`terminal_flush_batch_size` above `1` to coalesce completed rows and flush them
@@ -177,12 +177,12 @@ async def handle(order: dict) -> None: ...
```
A worker buffers completed rows and flushes when the buffer reaches
-`terminal_flush_batch_size` **or** its inflight queue empties — so a
+`terminal_flush_batch_size` or its inflight queue empties, so a
lightly-loaded queue still flushes immediately and batching adds no latency;
batching only engages under sustained load.
-**What it buys.** In the benchmark (5 000 messages, `fetch_batch_size=100`), the
-terminal round-trips drop from one per message to one per batch — a **100×**
+In the benchmark (5 000 messages, `fetch_batch_size=100`), the
+terminal round-trips drop from one per message to one per batch, a 100×
reduction in terminal `DELETE`s (5 000 → 50) with the same rows deleted. Because
the terminal write stops being the bottleneck, a single batched worker
out-throughputs a four-worker per-row subscriber, so you reach high throughput
@@ -190,12 +190,12 @@ without spending the extra [connection budget](#connection-budget) that more
workers cost. The win is largest at low `max_workers` (where per-row deletes
serialise) and narrows as worker parallelism rises.
-**The tradeoff — read before enabling.** Batching holds completed-but-undeleted
-rows in memory until the flush. On a **graceful** stop the buffer is flushed
-(no redelivery). But on an **ungraceful** crash (SIGKILL / OOM / power loss),
+Read the tradeoff before enabling it. Batching holds completed-but-undeleted
+rows in memory until the flush. On a graceful stop the buffer is flushed
+(no redelivery). But on an ungraceful crash (SIGKILL / OOM / power loss),
up to `terminal_flush_batch_size` rows that already ran their handler are
redelivered when another replica reclaims them. The outbox is *already*
-at-least-once — **handlers must be idempotent** — so this is not a new failure
+at-least-once (handlers must be idempotent), so this is not a new failure
class, only a wider window: from at most one at-risk row (per-row) to up to a
full batch. Two further effects to size for:
@@ -206,7 +206,7 @@ full batch. Two further effects to size for:
`fetch_batch_size + max_workers × (terminal_flush_batch_size + 1)`; keep
`lease_ttl_seconds` sized against that.
-It is **off by default** (`terminal_flush_batch_size=1` is byte-for-byte the
+It is off by default (`terminal_flush_batch_size=1` is byte-for-byte the
per-row path). Enable it per subscriber when the queue is high-throughput and
its handler is idempotent; leave it off for low-volume or
exactly-once-sensitive queues.
@@ -225,10 +225,10 @@ row.
| `AckPolicy.NACK_ON_ERROR` (default) | Consult the retry strategy on handler exceptions |
| `AckPolicy.REJECT_ON_ERROR` | Delete on the first failure (the retry strategy is ignored) |
| `AckPolicy.MANUAL` | Handler must call `await msg.ack()` / `nack()` / `reject()` itself |
-| `AckPolicy.ACK_FIRST` | **Not supported.** Passing it raises `ValueError` at registration |
+| `AckPolicy.ACK_FIRST` | Not supported. Passing it raises `ValueError` at registration |
`ACK_FIRST` would delete the row *before* the handler runs, so a handler
-crash silently drops the message — defeating the outbox reliability
+crash silently drops the message, which defeats the outbox reliability
guarantee. The factory rejects it at registration.
```python
@@ -248,9 +248,9 @@ async def handle(msg: OutboxMessage, body: dict) -> None:
```
!!! warning "MANUAL: returning without acking is a terminal reject"
- Under `AckPolicy.MANUAL`, a handler that returns **without** calling
+ Under `AckPolicy.MANUAL`, a handler that returns without calling
`ack()` / `nack()` / `reject()` (and without raising) is treated as a
- terminal **reject** — the row is **deleted** (or written to the DLQ with
+ terminal reject: the row is deleted (or written to the DLQ with
`failure_reason="rejected"` if a `dlq_table` is configured), not retried.
A handler that *raises* is nacked through the retry strategy instead, so
only the silent-return path is destructive. Always ack/nack/reject on
@@ -293,7 +293,7 @@ U(-jitter_factor/2, +jitter_factor/2)` to spread out retries, matching
| Strategy | Required | Optional (default) |
|---|---|---|
-| `NoRetry` | — | — |
+| `NoRetry` | - | - |
| `ConstantRetry` | `delay_seconds` | `jitter_factor` (`0.0`) |
| `LinearRetry` | `initial_delay_seconds`, `step_seconds` | `jitter_factor` (`0.0`) |
| `ExponentialRetry` | `initial_delay_seconds` | `multiplier` (`2.0`), `max_delay_seconds` (`None`), `jitter_factor` (`0.0`) |
@@ -335,17 +335,17 @@ The base strategy also enforces `max_attempts` and
Each subscriber holds `max_workers + 1` long-lived SQLAlchemy pool
connections (one writer per worker + one fetch), plus one raw asyncpg
-connection for `LISTEN` when available. Size your **engine pool** for
-`Σ subscribers × (max_workers + 1)`. An undersized pool does **not** block
-`broker.start()` — `start()` only schedules the loop tasks and returns;
-instead the fetch/worker loops stall on pool checkout and surface as
+connection for `LISTEN` when available. Size your engine pool for
+`Σ subscribers × (max_workers + 1)`. An undersized pool does not block
+`broker.start()`, because `start()` only schedules the loop tasks and returns.
+Instead the fetch/worker loops stall on pool checkout and surface as
repeating reconnect ERROR logs with dispatch silently starved. SQLAlchemy's
default `pool_size=5, max_overflow=10` covers a handful of single-worker
subscribers; raise it for larger fleets.
Server-side, the footprint is one larger: the raw asyncpg `LISTEN`
-connection lives **outside** the pool, so each subscriber consumes
-`max_workers + 2` Postgres connections. The budget is **per process** —
+connection lives outside the pool, so each subscriber consumes
+`max_workers + 2` Postgres connections. The budget is per process:
each replica opens its own pool and LISTEN connections, so your Postgres
`max_connections` needs to cover `replicas × Σ subscribers × (max_workers +
2)`, otherwise additional replicas (or rolling deployments) are refused at
@@ -355,8 +355,8 @@ startup with `FATAL: too many connections`.
## Read-only inspection
-`subscriber.get_one()` and `async for msg in subscriber:` are **not
-supported** on `OutboxSubscriber` — both raise `NotImplementedError`.
+`subscriber.get_one()` and `async for msg in subscriber:` are not
+supported on `OutboxSubscriber`; both raise `NotImplementedError`.
They would acquire a lease and bump `deliveries_count`, surprising
semantics for a peek API. Use
`broker.fetch_unprocessed(session=..., queue=...)` for lease-free reads of
diff --git a/docs/usage/testing.md b/docs/usage/testing.md
index 190cc12..55d68d8 100644
--- a/docs/usage/testing.md
+++ b/docs/usage/testing.md
@@ -1,11 +1,11 @@
# Testing
-`faststream-outbox` ships `TestOutboxBroker` — a test context manager that
+`faststream-outbox` ships `TestOutboxBroker`, a test context manager that
swaps the SQLAlchemy-backed client for an in-memory `FakeOutboxClient` so
unit tests don't need Postgres.
-By default it dispatches handlers **synchronously inside `publish`** —
-matching `TestKafkaBroker` / `TestRabbitBroker`. No `_wait_until`, no
+By default it dispatches handlers synchronously inside `publish`, matching
+`TestKafkaBroker` / `TestRabbitBroker`, so tests need no `_wait_until` or
`sleep`.
## Basic test
@@ -37,13 +37,13 @@ async def test_handler() -> None:
assert received == [1]
```
-On `broker.publish` (and `publish_batch`), `session=` is optional in tests —
-the test broker patches those methods to ignore it. This does **not** extend
+On `broker.publish` (and `publish_batch`), `session=` is optional in tests:
+the test broker patches those methods to ignore it. This does not extend
to `broker.publisher("q").publish(...)`, which still requires a `session`
(see [Testing publishers](#testing-publishers) below).
The fake client keeps an in-memory list of rows you can inspect via
-`fake_client.rows` — but `fake_client` is an attribute of the
+`fake_client.rows`. `fake_client` is an attribute of the
`TestOutboxBroker` harness, not the broker, so bind the harness to a name:
```python
@@ -100,9 +100,9 @@ shown.
## Loop-driven mode
-For tests that exercise real polling semantics — retry rescheduling, lease
+For tests that exercise real polling semantics (retry rescheduling, lease
expiry / reclaim, fetch-loop error recovery, or honoring `activate_in`
-delays — opt in with `run_loops=True`:
+delays), opt in with `run_loops=True`:
```python
import asyncio
@@ -139,22 +139,22 @@ registered handlers are not started, matching production.
## Notes
-- **`activate_in` / `activate_at` are ignored in sync mode.** Timers fire
+- `activate_in` / `activate_at` are ignored in sync mode. Timers fire
immediately. The intended firing time is preserved on the harness's
`fake_client.rows[i].next_attempt_at` for assertions. Use
`run_loops=True` if you need scheduled delivery to actually wait.
-- **`cancel_timer` and `fetch_unprocessed` run against the fake client**
+- `cancel_timer` and `fetch_unprocessed` run against the fake client
(the test broker swaps its client in, so `broker.` reaches it).
The `session` argument is still required but is ignored in tests.
-- **The fake producer uses the same envelope format as the real one**, so
+- The fake producer uses the same envelope format as the real one, so
all serialization paths are exercised.
-- **`lease_ttl_seconds` and re-delivery are not simulated** in sync mode —
- handlers that exceed the configured TTL in production may be re-delivered
+- `lease_ttl_seconds` and re-delivery are not simulated in sync mode.
+ Handlers that exceed the configured TTL in production may be re-delivered
to another worker, but tests will only invoke the handler once.
Idempotency must be verified separately. Use `run_loops=True` for tests
that need to observe lease-expiry behavior.
-- **`broker.validate_schema()` raises `NotImplementedError` under
- `TestOutboxBroker`**: there is no real DB to validate against, and a
+- `broker.validate_schema()` raises `NotImplementedError` under
+ `TestOutboxBroker`: there is no real DB to validate against, and a
silent pass would let users ship broken schemas while their tests stay
green. Tests that need real schema validation must call
`validate_schema()` on an `OutboxBroker(real_engine, outbox_table=...)`
diff --git a/docs/usage/timers.md b/docs/usage/timers.md
index 20cd4f8..e31cb8f 100644
--- a/docs/usage/timers.md
+++ b/docs/usage/timers.md
@@ -1,7 +1,7 @@
# Timers
-Schedule an outbox row to fire later by passing `activate_in` (relative)
-or `activate_at` (absolute, tz-aware) — exactly one. Pass `timer_id` to
+Schedule an outbox row to fire later by passing exactly one of
+`activate_in` (relative) or `activate_at` (absolute, tz-aware). Pass `timer_id` to
deduplicate per `(queue, timer_id)`; cancel a not-yet-leased timer with
`broker.cancel_timer(...)`.
@@ -36,7 +36,7 @@ same `(queue, timer_id)` already exists.
### Mutually exclusive
Passing both `activate_in` and `activate_at` raises `ValueError`. They are
-two ways to say the same thing — "make this row invisible to fetch until
+two ways to say the same thing: "make this row invisible to fetch until
the given moment".
### Timezone-aware `activate_at`
@@ -49,7 +49,7 @@ explicit `ValueError` rather than guessing your intended zone.
For `publish` with `activate_in`, `next_attempt_at` is computed server-side
via `now() + make_interval(secs => :s)` to stay clock-skew-safe. With
`activate_at`, you supply an absolute instant, so it is stored verbatim and
-compared against the *worker's* clock at fetch time — only `activate_in` is
+compared against the *worker's* clock at fetch time. Only `activate_in` is
skew-safe; `activate_at` is as accurate as your producers' and workers'
clocks agree. For `publish_batch`, `activate_in` is also client-side
(`datetime.now(UTC) + activate_in`) because executemany doesn't compose
@@ -85,9 +85,8 @@ assert second is None
```
NOTIFY is skipped when the row is genuinely future-dated (a *future*
-`activate_in` / `activate_at`) OR the conflict path returned no row — both
-cases would either wake listeners that find nothing, or wake them
-prematurely. A *past* `activate_at` is already eligible, so it still
+`activate_in` / `activate_at`) or the conflict path returned no row. In both cases a NOTIFY would either
+wake listeners that find nothing or wake them prematurely. A *past* `activate_at` is already eligible, so it still
notifies.
`timer_id` is only available on single `publish`, not on `publish_batch`
@@ -95,17 +94,17 @@ notifies.
### `timer_id` dedups only *live* rows
-The unique index is **partial** — `(queue, timer_id) WHERE timer_id IS NOT
+The unique index is partial: `(queue, timer_id) WHERE timer_id IS NOT
NULL`. It constrains only rows currently in the table. Once a timer fires
(the row is deleted) or is cancelled, the same `timer_id` can be inserted
-fresh. So `timer_id` is a dedup key for **in-flight / pending** timers, not
-a permanent idempotency key — it won't stop a value from being re-published
+fresh. So `timer_id` is a dedup key for in-flight / pending timers, not
+a permanent idempotency key: it won't stop a value from being re-published
after the original delivery has completed.
## Cancellation
`broker.cancel_timer(*, queue, timer_id, session)` issues a `DELETE` on
-the caller's session, but only if the row is **not yet leased**:
+the caller's session, but only if the row is not yet leased:
```python
deleted = await broker.cancel_timer(
@@ -128,7 +127,7 @@ returns `False`.
Timer firing latency is bounded by the subscriber's `max_fetch_interval`
(default `10` seconds) after `next_attempt_at` elapses. NOTIFY does not
-help here — listeners can't act on a future row, so the fetch loop has to
+help here: listeners can't act on a future row, so the fetch loop has to
poll for it.
Lower `max_fetch_interval` for sub-10s precision. Sub-second precision is
@@ -138,12 +137,12 @@ sleep inside the handler, or use a different scheduler.
## Test broker note
In tests using `TestOutboxBroker` (default `run_loops=False` mode),
-`activate_in` / `activate_at` are **ignored** and timers fire immediately
-— sync dispatch ignores `next_attempt_at`. This trades production parity
+`activate_in` / `activate_at` are ignored and timers fire immediately,
+because sync dispatch ignores `next_attempt_at`. This trades production parity
for test ergonomics: tests can assert handler effects without time travel.
-The schedule is still recorded on the fake row — bind the
-`TestOutboxBroker` to a name and read
-`tb.fake_client.rows[0].next_attempt_at` — if a test needs to assert on it.
+The schedule is still recorded on the fake row. If a test needs to assert
+on it, bind the `TestOutboxBroker` to a name and read
+`tb.fake_client.rows[0].next_attempt_at`.
Pass `run_loops=True` if you need scheduled delivery to actually
wait. See [Testing](./testing.md).