Skip to content

feat: Cassandra - ENABLING/DISABLING TTL states, non-blocking updates, durable backfill cursor - #349

Merged
robinnsc merged 2 commits into
feat/storage-cassandra-in-treefrom
feat/cassandra-ttl-lifecycle
Sep 19, 2026
Merged

robinnsc merged 2 commits into
feat/storage-cassandra-in-treefrom
feat/cassandra-ttl-lifecycle

Conversation

@robinnsc

Copy link
Copy Markdown
Collaborator

What

The last functional item from the TTL production-readiness list: observable lifecycle states and non-blocking UpdateTimeToLive.

TimeToLiveStatus gains ENABLING and DISABLING (DynamoDB's wire names). Until now UpdateTimeToLive blocked for the entire enable backfill — a full table scan, so the call took as long as the table was large and could outlive a client timeout — and blocked again on disable for the queue drain. Meanwhile DescribeTimeToLive reported ENABLED the moment the attribute was set, including for tables whose queue didn't yet cover every item and where nothing would expire.

The interesting part: the Cassandra catalog already encoded the full lifecycle (ttl_index_ready published by the backfill, ttl_cleanup_generation held until the drain finishes) — nothing read it. So the Cassandra change is mostly surfacing existing state:

  • Enable returns once the catalog flip is durable; the backfill runs detached and the table reports ENABLING until readiness is published. If the detached task dies, the TTL worker's pending-index pass retries.
  • A durable cursor makes that retry resume instead of rescan. New catalog column (V004 migration) written after every scanned page, keyed to the enable generation. A backfill that dies at page 400 of a large table costs one replayed page on retry, not 400.
  • Disable returns once the flip is durable; the worker's pending-cleanup pass finishes the drain and the table reports DISABLING until it does.
  • Updates are rejected mid-transition ("Time to live is being updated for this table"), matching DynamoDB. The gate is a pure function with its full 8-case truth table unit-tested.

PostgreSQL derives ENABLING from its own ttl_index_ready (its disable is synchronous, so DISABLING never appears there). SQLite and MongoDB keep their synchronous two-state behavior — see the design notes for why the handler treats them differently.

Why

This closes the "no ENABLING/DISABLING status" and "backfill has no durable cursor" rows from ADR-0010's known-gaps table — the two that were user-visible API behavior rather than internals. After this, the remaining gaps in that table are throughput scaling and async-GSI support, both explicitly deferred product decisions.

Design decisions worth reviewing

  • The cursor write is an LWT fenced on the live generation, not a plain write. This came out of adversarial review, and it's the most important line in the change: a plain UPDATE is an upsert in Cassandra, so a detached backfill racing DeleteTable would have re-created a partial catalog row for the dead table — blocking any same-name CreateTable afterward. With the fence, a missing row or a moved generation refuses the write and the backfill stops. Residual exposure is at most one page of queue entries for a dead table_id, which are unreachable garbage (nothing sweeps a table absent from the catalog), not resurrected state. One Paxos round per 1,000-item page is noise next to the page's inserts.
  • The handler only detaches the backfill when the backend reports ENABLING. Also from review. SQLite and MongoDB describe TTL as ENABLED straight from the attribute — they have no transitional state to report — so detaching their index creation would leave a table claiming ENABLED while its index doesn't exist yet, with nothing observable to say otherwise. The handler describes after the flip and spawns only on ENABLING; two-state backends keep the awaited call they had before. Behavior is derived from what each backend can truthfully report rather than assumed uniform.
  • Disable during ENABLING is rejected, which means a mistaken enable on a huge table can't be cancelled until its backfill lands. DynamoDB has the same rule, and allowing the cancel would mean tearing down a scan that's mid-flight under the control lease. Judged acceptable; flagging it because it's the one place the gate has real user-facing cost.
  • PostgreSQL hardening that review caught while looking at the spawn: a CREATE INDEX CONCURRENTLY that fails partway leaves an INVALID index which IF NOT EXISTS would silently keep on retry — so invalid leftovers are dropped before rebuilding — and the readiness publication is now fenced on the attribute the build indexed, so a stale detached task can't certify a later enable's readiness. Both were latent before this PR (the worker retry could hit them) but the detached path made them likelier.
  • DISABLING reports AttributeName as absent, unlike DynamoDB (the catalog nulls the attribute at flip time). Recorded in differences-from-dynamodb rather than worked around with an extra column.

Risks

  • The V004 migration is the first schema migration this feature ships. It's ADD IF NOT EXISTS (verified against 4.1) so a half-applied run is rerunnable. One trap documented in the file: the migration runner splits statements on every semicolon including inside comments — a semicolon in a comment broke the first local run of this migration.
  • Existing disable-path tests encoded the old synchronous-drain semantics; they now invoke the worker's drain pass explicitly. If any external caller relied on "disable returns ⇒ queue drained," that assumption breaks — I found none in-tree (the engine handler and delete/update table paths were checked during review).
  • The enable path does one extra describe_ttl per UpdateTimeToLive call (to decide spawn vs await). That's one catalog read on a rare control-plane operation.

Planned follow-ups (not in this PR)

  • Porting the plug-in repo's 135 direct integration tests (from the feat(cassandra): TTL production-readiness follow-ups #339 review thread) — inventoried, coming as its own PR since it's ~8,600 LOC of test-only change.
  • Load testing before production performance claims; per-shard sweep scaling; retiring the item_data fence fallback (all tracked in ADR-0010).

Testing done

Real Cassandra 4.1, single node, serial:

  • cargo test -p extenddb-storage-cassandra --test ttl_integration -- --test-threads=1 — 29 passed, 0 failed (26 prior + 3 new)
  • cargo test -p extenddb-storage-cassandra --test metadata_integration — 1 passed
  • cargo test --workspace --lib — 1,091 passed, 0 failed
  • cargo clippy --all-targets -- -D warnings — clean; cargo fmt --all -- --check — clean

New tests, each earning its place:

  • test_ttl_lifecycle_states: the four states in order, including that DISABLING is genuinely observable (the drain really is deferred) and that the worker's cleanup pass is what lands DISABLED.
  • test_backfill_cursor_resume: kills the backfill task mid-scan (2,500 items, aborted after the first page's cursor lands), simulates the control-lease expiry a real dead host would age out of, runs the worker's retry pass, and proves full coverage by asserting the audit repairs zero items afterward.
  • test_backfill_cursor_is_honored: the resume test alone can't distinguish resume from silent rescan (both yield full coverage), so this one plants a cursor at the scan-last item and asserts the audit finds exactly the ten skipped registrations — a rescan would find zero. Scan order is asked from the engine rather than assumed, since token order isn't insertion order.
  • transition_gate (engine unit test): all 8 state × request combinations.

The change went through an adversarial subagent review before commit; the two HIGH findings (cursor-upsert resurrection, two-state-backend detachment) are fixed above and called out in the design notes. Not tested here: PostgreSQL paths beyond compilation — covered by the postgres CI integration jobs.

Checklist

  • I have read CONTRIBUTING.md
  • Affected packages build and all applicable tests pass against real Cassandra
  • Code is formatted (cargo fmt --all -- --check)
  • This branch introduces no new Clippy warnings (workspace -D warnings clean)
  • Tests cover the new behavior
  • Documentation (ADR-0010, differences-from-dynamodb) updated to match the code
  • Wire-visible change: TimeToLiveStatus gains two variants (additive; serialized names match DynamoDB). No Storage trait breaking change; one new catalog column via migration.

ADR / RFC: ADR-0010 — Durable sharded expiration queue for Cassandra TTL

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache License 2.0 and I agree to the Developer Certificate of Origin (DCO). See CONTRIBUTING.md for details.

@robinnsc robinnsc changed the title feat: ENABLING/DISABLING TTL states, non-blocking updates, durable backfill cursor feat: Cassandra - ENABLING/DISABLING TTL states, non-blocking updates, durable backfill cursor Sep 15, 2026
@robinnsc

Copy link
Copy Markdown
Collaborator Author

The test port Joel asked about is up as its own PR: #351 (135 previously-unported tests + the 4 weak ones strengthened — one of which, the sync-GSI test, was masking a real semantic difference between the plug-in engine and the in-tree default async GSI propagation).

@robinnsc

Copy link
Copy Markdown
Collaborator Author

Correction to my previous comment: the test-port PR is #350, not #351.

@jcshepherd
jcshepherd force-pushed the feat/storage-cassandra-in-tree branch 2 times, most recently from 8151972 to 845f933 Compare September 15, 2026 21:05
@robinnsc
robinnsc force-pushed the feat/cassandra-ttl-lifecycle branch from 5dacde7 to d7d1f0a Compare September 17, 2026 19:59
…ckfill cursor

UpdateTimeToLive used to block for the whole enable backfill (a full
table scan) and the whole disable drain, so the call took as long as the
table was large and could outlive a client timeout; meanwhile
DescribeTimeToLive reported ENABLED for tables whose queue did not yet
cover every item. The Cassandra catalog already encoded the real
lifecycle in ttl_index_ready and ttl_cleanup_generation — nothing read
it.

TimeToLiveStatus gains Enabling and Disabling (DynamoDB's wire names).
Cassandra derives all four states from the catalog; PostgreSQL derives
Enabling from its own ttl_index_ready (its disable is synchronous, so
Disabling never appears); SQLite and MongoDB keep their two-state
synchronous behavior, and the handler only detaches the backfill when
the backend actually reports a transitional state — detaching on a
two-state backend would leave a table claiming Enabled with no index.
Updates are rejected while a transition settles, as DynamoDB does; the
gate is a pure function with the full 8-case truth table tested.

Enable returns once the catalog flip is durable and the backfill runs
detached, resuming after a failure from a new durable cursor: a JSON
resume point written after every scanned page as an LWT fenced on the
live ttl_generation. The fence is what makes the cursor safe, not just
useful — a plain write is an upsert, and a detached backfill racing
DeleteTable would resurrect a partial catalog row for the dead table,
blocking a same-name CreateTable (adversarial-review finding). A
refused fence means the lifecycle moved and the scan stops. Disable
returns once the flip is durable and the worker's pending-cleanup pass
finishes the drain; the table reports DISABLING until it does.

Also from review: a PostgreSQL CONCURRENTLY build that fails partway
leaves an INVALID index that IF NOT EXISTS would silently keep on
retry, so invalid leftovers are dropped before rebuilding, and the
readiness publication is fenced on the attribute the build indexed so a
stale detached task cannot certify a later lifecycle. The V004
migration uses ADD IF NOT EXISTS so a run that dies between applying
and recording is rerunnable. The migration runner splits statements on
every semicolon including in comments; V004 documents that trap.

Tests: the four lifecycle states in order; a backfill killed mid-scan
(cursor present, still ENABLING) resumed to completion by the worker
pass with an audit proving full coverage; a planted cursor at the
scan-last item proving the worker path resumes rather than rescans
(audit finds exactly the ten skipped registrations); the transition
gate truth table. Existing disable tests now invoke the worker drain
they previously got inline.
Enable returns during ENABLING now and updates are rejected until the
transition settles, so the enable-then-immediately-disable sequence in
the Python and Rust E2E suites hits the transition gate — the same
rejection real DynamoDB gives that sequence, which is why both tests
already carried a cooldown comment about it. Poll DescribeTimeToLive for
ENABLED (30s bound) before disabling, like a real client must. Both
tests' final assertions already accepted DISABLED or DISABLING.
@robinnsc
robinnsc force-pushed the feat/cassandra-ttl-lifecycle branch from d7d1f0a to 3223a92 Compare September 17, 2026 20:17
@robinnsc

Copy link
Copy Markdown
Collaborator Author

Rebuilt on the current base. Context for the diff churn: the branch reconstruction (845f933, "chore:rebase from main") re-applied the backend squash and the PR #338 OCC work but dropped merged PR #339 entirely — the audit worker, TTL metrics, transaction-recovery reconciliation, config cache, concurrent sweep, CI workflow, and test gating all vanished from the branch, which is why this PR's diff ballooned and conflicted. Restored directly on the base as 439ca5b (adapted to #338's OCC layer rather than re-applied verbatim — #338's claim redesign supersedes #339's put-path/claim-fence changes and stays). This PR is now just its own two commits again. One bonus from the test-port PR (#350): its delete suite caught a real #338 regression — the new OCC delete fast path skipped writing partition_max_delete_timestamp, disarming the new-item-PUT-vs-delete transaction check for the most common delete shape. Fixed on the base as 0aebbe4.

@robinnsc
robinnsc merged commit 69efe88 into feat/storage-cassandra-in-tree Sep 19, 2026
25 checks passed
@robinnsc
robinnsc deleted the feat/cassandra-ttl-lifecycle branch September 19, 2026 01:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants