Skip to content

Merge a diff of many small runs in whole granules - #1063

Draft
ejc3 wants to merge 1 commit into
mainfrom
snapshot-merge-whole-granules
Draft

ejc3 wants to merge 1 commit into
mainfrom
snapshot-merge-whole-granules

Conversation

@ejc3

@ejc3 ejc3 commented Oct 4, 2026 •

Copy link
Copy Markdown
Owner

A diff of many small runs is merged in whole 1 MiB granules instead of one write per run.

Follows: #1058 (merged), which runs the merge with workers. Base: main. This is a draft until one question is decided: whether the room rule below, which is a check and not a reservation, is enough (see "What the room does not cover").

The Problem

After #1058 the merge of a restored VM's memory diff still copies each data run with one write. Half of the runs of a large guest's diff are single 4 KiB pages. Writes to one file take that file's lock one at a time, so the workers speed up the reads and not the writes. Each write inside a 128 KiB extent of the reflinked base also splits it.

In fcvm snapshot create of a 128 GiB guest whose diff held 38.1 GiB in 2,251,186 runs, the merge took 554 s and left a file of 5,251,673 extents, five times its base's, which every later restore reads through. On a second host the same merge took about 941 s, where one thread had taken 745 s.

The Solution

A first walk writes nothing and counts the diff's runs. The diff is then written in whole granules when three things hold, and run by run as before when any does not:

  • Many runs. At least 262,144 (MERGE_ROUND_FROM_RUNS).
  • Small runs. Whole granules add at most 64 KiB for each run (MERGE_ROUND_BELOW).
  • Room. The free space beside the file holds the granules and the memory file once more. A filesystem that does not say how much is free gets no granules. A merge is never refused for space.

In whole granules every run is rounded out to 1 MiB, neighbours are joined, and each range is written back whole: the base's bytes with the diff's runs laid over them. The base is read only where the runs do not cover the range.

The merge takes no lock and waits for nothing.

Why These Rules

  • The bar of 262,144 runs. Whole granules save time in proportion to the runs they spare. At the measured 7,200 runs a second that many take 36 s one at a time, which is the most the bar can cost. A small guest's diff is far below it: 58.9 MiB in 4,775 runs on a 1 GiB guest, which whole granules would write as 350 MiB. Its snapshot costs the disk it did and keeps sharing its base's extents.
  • 64 KiB a run. In the first table below, one write per run costs 112 microseconds beyond its bytes, the time 73 KiB take. That diff's granules add 33 KiB a run. Single pages scattered one to a granule add 1 MiB a run: whole granules would rewrite the file for them, take longer than their writes, and leave nothing shared with the base.
  • Room for the file once more. A full snapshot of a guest this size dumps that many bytes while its VM is paused, and running out of space there damages the VM. The merge leaves that room. It also makes two merges that read the same free space both fit: neither writes more than its file's size, and the smaller file's merge fits in the room the larger one leaves. That is why no lock is needed. The amount compared is bytes written, and a compressing filesystem stores them in less.
  • What the room does not cover. The reading of the free space is a check and not a reservation. More large merges or full snapshots at once than one merge beside another, or one merge beside one full snapshot, can still run a short disk out of space. Full snapshots at once can do that today without any merge.
  • One choice for the diff. An earlier state of this branch chose for each 8 MiB window. On the 128 GiB guest that merge took 300 s and wrote 87 percent of what whole granules came to. Whole granules in every window took 263 s on a diff of the same size.

Measurements

The two ways of writing, each applied to the whole file, on the two files of a 128 GiB guest's snapshot (diff of 38.1 GiB in 2,328,992 runs; btrfs, compress-force=zstd:3, page cache dropped before each run, 16 workers):

written as merge writes written extents afterwards
one write per run (#1058) 323 s 2,329,373 38.1 GiB 5,374,185
1 MiB granules 178 s 18,610 111.4 GiB 1,041,076
8 MiB granules 199 s 16,173 126.4 GiB 1,041,213

The cost of granules is disk when a VM touched this much: 111 GiB of the 128 were rewritten, which took 37 GiB after compression where one write per run took 13 GiB.

On the real path, fcvm snapshot create of the 128 GiB guest on one host, each from a VM restored and warmed the same way:

build diff merge written extents afterwards whole command
one write per run with 16 workers (#1058, on main) 38.1 GiB in 2,251,186 runs 554 s 38.1 GiB 5,251,673 14m39s
this pull request 37.9 GiB in 2,355,784 runs 263 s 110.8 GiB 1,045,539 10m50s
an earlier state that chose for each 8 MiB window 37.2 GiB in 2,384,760 runs 300 s 94.3 GiB 1,330,999 12m11s

The second row's line: bytes_merged=40722845696 runs=2355784 writes=20357 bytes_written=118945218560 rounded_bytes=118945218560 base_reads=20132 workers=16 merge_ms=263206. The free space when it ran held the granules and the file once more, so the merge wrote whole granules.

During the merge the host's memory pressure (some avg10 in /proc/pressure/memory, sampled every 5 s) peaked at 12.7 percent in the second row's run and at 13.2 percent in the first's. The source VM was resumed and idle in each run.

Test Results

$ make _test-unit FILTER=-E 'test(/^commands::common::/)'
     Summary [   1.047s] 62 tests run: 62 passed, 1373 skipped

$ make fmt      # leaves the tree unchanged
$ make clippy   # exit 0

$ make _test-root FILTER=-E 'test(/test_user_snapshot_from_clone_uses_parent/)'
  diff merge complete, building atomic update snapshot=user-parent-snap-1680854-0-user bytes_merged=54276096 runs=4405 writes=4406 bytes_written=54276096 rounded_bytes=356515840 base_reads=0 workers=2 merge_ms=408
  ✓ A clone of the merged snapshot holds the first clone's 8 MiB
     Summary [  24.504s] 1 test run: 1 passed, 1785 skipped

Red first, one mutation at a time:

# whole granules without room for the memory file once more
  TRY 1 FAIL  a_merge_writes_run_by_run_when_whole_granules_would_not_leave_room_for_the_file_again
      left: (131072, 131072)
     right: (8192, 131072)

# no bound on the bytes whole granules add for each run
  TRY 1 FAIL  a_diff_whose_runs_are_scattered_is_written_run_by_run_however_many_they_are
      left: MergeStats { diff_bytes: 12288, runs: 3, writes: 2, bytes_written: 3145728, base_reads: 2, workers: 1, rounded_bytes: 3145728 }
     right: MergeStats { diff_bytes: 12288, runs: 3, writes: 3, bytes_written: 12288, base_reads: 0, workers: 1, rounded_bytes: 3145728 }

# whole granules for a diff of any number of runs
  TRY 1 FAIL  a_diff_of_few_runs_is_written_run_by_run_whatever_its_windows_hold
  TRY 1 FAIL  a_merge_holds_the_same_bytes_however_it_is_cut_and_however_many_workers_make_it

# a merge that cannot read the free space beside its base
  TRY 1 FAIL  a_merge_reads_the_free_space_of_the_filesystem_that_holds_its_base
      left: (8192, 0)
     right: (131072, 2)

# every whole-granule write reading the base
  TRY 1 FAIL  a_run_longer_than_a_window_goes_out_in_pieces_that_read_no_base
      left: MergeStats { diff_bytes: 307200, runs: 1, writes: 6, bytes_written: 311296, base_reads: 6, workers: 1, rounded_bytes: 311296 }
     right: MergeStats { diff_bytes: 307200, runs: 1, writes: 6, bytes_written: 311296, base_reads: 1, workers: 1, rounded_bytes: 311296 }

Not run: arm64, a filesystem other than btrfs, and a merge while the source VM is busy.

@coderabbitai

coderabbitai Bot commented Oct 4, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Base automatically changed from snapshot-merge-whole-extents to main October 4, 2026 14:12
@ejc3
ejc3 force-pushed the snapshot-merge-whole-granules branch from 4618ed7 to 05cc1db Compare October 4, 2026 15:26
@ejc3 ejc3 changed the title Write whole granules where a merged diff's runs are many and small Merge a diff of many runs in whole granules Oct 4, 2026
@ejc3
ejc3 force-pushed the snapshot-merge-whole-granules branch from 05cc1db to 2564f73 Compare October 4, 2026 16:13
@ejc3 ejc3 changed the title Merge a diff of many runs in whole granules Merge a diff of many small runs in whole granules Oct 4, 2026
@ejc3
ejc3 force-pushed the snapshot-merge-whole-granules branch from 2564f73 to 7d916a0 Compare October 4, 2026 17:10
The merge of a restored VM's memory diff copies each data run to the base with
one write. Half of the runs of a large guest's diff are single pages. Writes
to one file take the file's lock one at a time, so the 16 workers the merge
runs speed up its reads and not its writes, and each write inside a 128 KiB
extent of the reflinked base splits it. In `fcvm snapshot create` of a 128 GiB
guest whose diff held 38.1 GiB in 2,251,186 runs the merge took 554 s and left
a file of 5,251,673 extents, five times its base's, which every later restore
reads through. On a second host the same merge took about 941 s, where one
thread had taken 745 s.

A first walk writes nothing and counts the diff's runs. The diff is then
written in whole 1 MiB granules (MERGE_GRANULE) when three things hold, and
run by run as before when any does not:

- It has at least 262,144 runs (MERGE_ROUND_FROM_RUNS). Whole granules save
  time in proportion to the runs they spare, and at 7,200 runs a second that
  many take 36 s one at a time. A small guest's diff is far below it (58.9 MiB
  in 4,775 runs on a 1 GiB guest, which whole granules would write as
  350 MiB), so its snapshot costs the disk it did and keeps sharing its base's
  extents.

- Whole granules add at most 64 KiB for each run (MERGE_ROUND_BELOW). In the
  table below one write per run costs 112 microseconds beyond its bytes, the
  time 73 KiB take, and that diff's granules add 33 KiB a run. Single pages
  scattered one to a granule add 1 MiB a run: whole granules would rewrite the
  file for them.

- The free space beside the file holds the granules and the memory file once
  more. That leaves room for a full snapshot of a guest this size, whose dump
  must not run out of space while its VM is paused. Two merges that read the
  same free space both fit: neither writes more than its file's size, and the
  smaller file's merge fits in the room the larger one leaves. So the merge
  takes no lock. The reading is a check and not a reservation: more large
  merges or full snapshots at once than that can still run a short disk out
  of space, as full snapshots at once can without any merge. The amount
  compared is bytes written, and a compressing filesystem stores them in
  less. A filesystem that does not say how much is free gets no granules. A
  merge is never refused for space.

In whole granules every run is rounded out, neighbours are joined, and each
range is written back whole: the base's bytes with the diff's runs laid over
them. The base is read only where the runs do not cover the range. The choice
is made once for the diff: made for each 8 MiB window, the merge of the guest
below took 300 s where whole granules everywhere took 263 s.

Measured on the two files of such a snapshot (btrfs, compress-force=zstd, base
of 1,032,975 extents, page cache dropped before each run, 16 workers, the
whole file merged one way):

  written as         merge   writes      written    extents afterwards
  one write per run  323 s   2,329,373   38.1 GiB   5,374,185
  1 MiB granules     178 s   18,610      111.4 GiB  1,041,076
  8 MiB granules     199 s   16,173      126.4 GiB  1,041,213

In `fcvm snapshot create` of the 128 GiB guest on one host, each from a VM
restored and warmed the same way:

  build                   diff                         merge   written     extents     command
  one write per run       38.1 GiB in 2,251,186 runs   554 s   38.1 GiB    5,251,673   14m39s
  (main, 16 workers)
  this change             37.9 GiB in 2,355,784 runs   263 s   110.8 GiB   1,045,539   10m50s

The second run's line: bytes_merged=40722845696 runs=2355784 writes=20357
bytes_written=118945218560 rounded_bytes=118945218560 base_reads=20132
workers=16 merge_ms=263206. The host's memory pressure (some avg10, sampled
every 5 s) peaked at 12.7 percent during that merge and at 13.2 percent during
the first.

The cost of granules is disk when a VM touched this much: 111 GiB of the 128
were rewritten, 37 GiB after compression, where one write per run took 13.

The "diff merge complete" line gains bytes_written, rounded_bytes (what whole
granules come to) and base_reads. The two free space checks before a snapshot
read the filesystem through the function the merge uses.

Tested (x86_64):

  make _test-unit FILTER=-E 'test(/^commands::common::/)'
                  Summary [   1.047s] 62 tests run: 62 passed, 1373 skipped
  make fmt        leaves the tree unchanged
  make clippy     exit 0
  make _test-root FILTER=-E 'test(/test_user_snapshot_from_clone_uses_parent/)'
                  Summary [  24.504s] 1 test run: 1 passed, 1785 skipped

Red first, one mutation at a time:

  whole granules without room for the memory file once more:
    a_merge_writes_run_by_run_when_whole_granules_would_not_leave_room_for_the_file_again
      left: (131072, 131072)  right: (8192, 131072)
  no bound on the bytes whole granules add for each run:
    a_diff_whose_runs_are_scattered_is_written_run_by_run_however_many_they_are
      left:  writes: 2, bytes_written: 3145728, base_reads: 2
      right: writes: 3, bytes_written: 12288, base_reads: 0
  whole granules for a diff of any number of runs:
    a_diff_of_few_runs_is_written_run_by_run_whatever_its_windows_hold
    a_merge_holds_the_same_bytes_however_it_is_cut_and_however_many_workers_make_it
  a merge that cannot read the free space beside its base:
    a_merge_reads_the_free_space_of_the_filesystem_that_holds_its_base
      left: (8192, 0)  right: (131072, 2)
  every whole-granule write reading the base:
    a_run_longer_than_a_window_goes_out_in_pieces_that_read_no_base
      left: base_reads: 6  right: base_reads: 1

The first three ran before the last two tests were added.

Not run: arm64, a filesystem other than btrfs, and a merge while the source VM
is busy (the source was idle in every run above).
@ejc3
ejc3 force-pushed the snapshot-merge-whole-granules branch from 7d916a0 to f140302 Compare October 4, 2026 19:17

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant