Skip to content

fix(operator): index recovery waits for the interrupted PostgreSQL statement - #100

Merged
krzysztof-smartdataengines merged 2 commits into
mainfrom
fix/index-recovery-waits-for-the-interrupted-statement
Sep 27, 2026
Merged

krzysztof-smartdataengines merged 2 commits into
mainfrom
fix/index-recovery-waits-for-the-interrupted-statement

Conversation

@krzysztof-smartdataengines

Copy link
Copy Markdown
Contributor

Summary

Resume and abandonment of a PostgreSQL in-place index build now wait for the interrupted statement
before they read the catalogue. Before, they read the catalogue while a budget-interrupted or
killed operator's CREATE INDEX CONCURRENTLY still ran in the server. Recovery then raced that
statement's commit and failed with "tuple concurrently updated", or dropped the index it was
finishing. This is in 0.1.0.

How it was found

CI on #99 (Python 3.12) failed in
test_the_build_budget_bounds_a_held_build_and_recovery_finishes_it[postgres]:
DROP INDEX CONCURRENTLY "sde_i_…_000001" raised tuple concurrently updated, and recovery
reported index build recovery remains incomplete; preserve its state. The same suite passed on
the other three Pythons and locally. The test's exact sequence, repeated 20 times locally, did not
fail once.

What was measured (PostgreSQL 15.19)

A probe watched pg_stat_activity, pg_stat_progress_create_index and pg_index through a
budget-ended build:

  • After the budget and the reconnect, the build's backend is still active. It waits on
    virtualxid in waiting for old snapshots, and the index is not valid.
  • Right after the snapshot is released, the backend is active in IO/WALSync, and its
    progress row is gone.
  • About 0.1 s later, the index is valid and the backend has exited.

The server notices a vanished client only when it next talks to it. For a concurrent build, that
is once the index is valid. DefineIndex releases the table's session lock before that last
transaction commits. A DROP INDEX CONCURRENTLY that was waiting for the lock starts in that
window, and it updates the pg_index row the build has not committed yet.

docs/in-place-index.md already promised "recovery waits for it, and keeps what it finished or
drops what it left unfinished". The code read the catalogue first, then waited only through the
drop's lock.

The change

  • NativeIndexBuild.settle: on PostgreSQL it waits until no active statement names the index
    (pg_stat_activity, the quoted bound name, not its own backend). The operator's deadline bounds
    the wait.
  • build and drop settle before they inspect. After the wait, what the statement finished is
    kept, and what it left unfinished is dropped and built again.
  • remove does not settle, and says why. Two drops of one index are serialized by the exclusive
    lock the first holds on it until it commits, and
    test_a_resumed_removal_waits_for_the_stopped_drop_and_finishes pins that. Mutation M4 below shows
    a settle there would break it.
  • One limit is documented: pg_stat_activity shows a statement's text only to its own login.

Tests

Two new live tests use a killed operator whose build still runs in the server. Neither calls
pg_terminate_backend. The snapshot is released a second into recovery or abandonment:

  • test_recovery_waits_for_a_killed_operators_build_and_keeps_what_it_finished checks that no drop
    is active before the release, that the outcome is built, and that the index has the same oid,
    so it was kept, not rebuilt.
  • test_abandoning_waits_for_a_killed_operators_build_before_dropping_it checks that no drop is
    active before the release, that the outcome is abandoned, and that the index is absent.

Before the change, both failed deterministically with the CI's tuple concurrently updated. The
budget test also asserts that the recovered index is the one the interrupted statement finished.

Mutations, each alone, restored from a copy kept beside the file:

Mutation Test Result
M1: build does not settle recovery beside the killed build red, tuple concurrently updated and recovery incomplete
M2: drop does not settle abandonment beside the killed build red, tuple concurrently updated and abandonment incomplete
M3: settle waits for idle statements, not active ones both red, both, with the same error
M4: remove settles too test_a_resumed_removal_waits_for_the_stopped_drop_and_finishes red, the resume waited for the stopped drop past its 6 s budget
C1 (control): poll every 0.05 s instead of 0.2 s all three green

C1 is the one that must survive, and it does.

Results

make check with both live engines on 02d0b99:

  • ruff and mypy are clean;
  • Python: 2103 passed, which is the 2101 of main plus the two new tests, and 10 skipped (the
    orderbook slice);
  • TypeScript: 999 of 999.

The index operator and index change files alone passed 119 of 119 with the fix, including the
removal test that M4 breaks.

🤖 Generated with Claude Code

…atement

A deadline or a killed operator ends the client, not its CREATE INDEX
CONCURRENTLY: the server notices the vanished client only once the index is
valid (measured, PostgreSQL 15.19). The build also releases the table's session
lock before its last transaction commits. Recovery read the catalogue first,
found the index unfinished, and sent DROP INDEX CONCURRENTLY, which waited for
that lock, started in the window and failed with "tuple concurrently updated"
(SDK CI, 27 September). Without the collision it dropped an index the
interrupted statement was finishing, against docs/in-place-index.md.

build and drop now wait until no active statement names the index, bounded by
the operator's deadline, and read the catalogue after: what the statement
finished is kept, what it left is dropped and built again. remove needs no such
wait - two drops of one index are serialized by the exclusive lock the first
holds on it - and a test pins that.

Two tests kill the operator mid-build and release the snapshot a second into
recovery and into abandonment. Before this change both failed with the CI's
"tuple concurrently updated". The budget test now checks that the index it
recovers is the one the interrupted statement finished.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
docs/in-place-index.md said recovery waits for a PostgreSQL build interrupted
by the budget; it did not. It now says how it waits (pg_stat_activity, within
the same budget), what went wrong when it read the catalogue first, that only
the operator's own login is visible, and why a removal needs no such wait. The
changelog gets an Unreleased section with the fix.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@krzysztof-smartdataengines
krzysztof-smartdataengines merged commit d362ec8 into main Sep 27, 2026
14 checks passed
@krzysztof-smartdataengines
krzysztof-smartdataengines deleted the fix/index-recovery-waits-for-the-interrupted-statement branch September 27, 2026 22:41
krzysztof-smartdataengines added a commit that referenced this pull request Sep 28, 2026
…101)

A patch release of the Python library alone, for #100: resume and
abandonment of a PostgreSQL in-place index build wait for the interrupted
statement before reading the catalogue, instead of racing its commit or
dropping the index it was finishing.

@smart-data-engines/sde has no 0.1.1: of what it ships, only the version
field changed since typescript-v0.1.0, so npm keeps 0.1.0 (one tag per
language, docs/publishing.md 5.1).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant