Skip to content

[Bug] MongoDB restore can remain CREATING after client disconnect during base copy #334

Description

@diegotoledano95

Describe the bug

If the client disconnects while RestoreTableFromBackup is copying the
base data with MongoDB $out, the restore can remain permanently stuck in
CREATING.

The restore request creates the target table and its secondary-index metadata
before copying the base data. The restore_backfill_pending marker is written
only after the copy completes.

To Reproduce

  1. Start ExtendDB with MongoDB and a replica set.
  2. Create a table containing a Global Secondary Index or Local Secondary Index.
  3. Create a backup containing data and index entries.
  4. Start RestoreTableFromBackup for that backup.
  5. Disconnect or cancel the client while the $out base-data copy is running.
  6. Describe the restored table and inspect its index status.

Expected behavior

An interrupted restore should either be safely retryable or be recovered by
the background worker. It should not remain permanently stuck in CREATING.

Actual behavior

The restore may exit before setting restore_backfill_pending. The table
remains in CREATING, and its restore-mode index jobs are not processed because
the worker waits for that marker. The restore therefore has no completion path.

Environment

  • ExtendDB version: 0.1.10
  • Operating system: macOS 26.6.2, Apple Silicon
  • MongoDB version: 7.0
  • Rust version: 1.97.1
  • Client SDK/driver and version: aws-sdk-dynamodb 1.113.0 for Rust
  • Deployment method: source
  • MongoDB deployment: single-node replica set
  • Number of ExtendDB instances: 1

Logs / stack trace

No error is necessarily reported after the client disconnects.

Additional context

The current restore flow intentionally creates secondary-index metadata in
CREATING before the base $out copy. The worker now correctly avoids
processing those restore indexes until restore_backfill_pending is set.
This follow-up issue covers recovery when the restore request is interrupted
before that marker is written. Possible solutions include making the restore
state resumable, adding recovery for abandoned restore jobs, or cleaning up
the incomplete target so the restore can be retried safely.

Checklist

  • I have searched existing issues for duplicates
  • I have included the ExtendDB version and environment details
  • I can reproduce this with the latest release

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions