Conversation
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
ejc3
force-pushed
the
snapshot-merge-whole-granules
branch
from
October 4, 2026 15:26
4618ed7 to
05cc1db
Compare
ejc3
force-pushed
the
snapshot-merge-whole-granules
branch
from
October 4, 2026 16:13
05cc1db to
2564f73
Compare
ejc3
force-pushed
the
snapshot-merge-whole-granules
branch
from
October 4, 2026 17:10
2564f73 to
7d916a0
Compare
The merge of a restored VM's memory diff copies each data run to the base with
one write. Half of the runs of a large guest's diff are single pages. Writes
to one file take the file's lock one at a time, so the 16 workers the merge
runs speed up its reads and not its writes, and each write inside a 128 KiB
extent of the reflinked base splits it. In `fcvm snapshot create` of a 128 GiB
guest whose diff held 38.1 GiB in 2,251,186 runs the merge took 554 s and left
a file of 5,251,673 extents, five times its base's, which every later restore
reads through. On a second host the same merge took about 941 s, where one
thread had taken 745 s.
A first walk writes nothing and counts the diff's runs. The diff is then
written in whole 1 MiB granules (MERGE_GRANULE) when three things hold, and
run by run as before when any does not:
- It has at least 262,144 runs (MERGE_ROUND_FROM_RUNS). Whole granules save
time in proportion to the runs they spare, and at 7,200 runs a second that
many take 36 s one at a time. A small guest's diff is far below it (58.9 MiB
in 4,775 runs on a 1 GiB guest, which whole granules would write as
350 MiB), so its snapshot costs the disk it did and keeps sharing its base's
extents.
- Whole granules add at most 64 KiB for each run (MERGE_ROUND_BELOW). In the
table below one write per run costs 112 microseconds beyond its bytes, the
time 73 KiB take, and that diff's granules add 33 KiB a run. Single pages
scattered one to a granule add 1 MiB a run: whole granules would rewrite the
file for them.
- The free space beside the file holds the granules and the memory file once
more. That leaves room for a full snapshot of a guest this size, whose dump
must not run out of space while its VM is paused. Two merges that read the
same free space both fit: neither writes more than its file's size, and the
smaller file's merge fits in the room the larger one leaves. So the merge
takes no lock. The reading is a check and not a reservation: more large
merges or full snapshots at once than that can still run a short disk out
of space, as full snapshots at once can without any merge. The amount
compared is bytes written, and a compressing filesystem stores them in
less. A filesystem that does not say how much is free gets no granules. A
merge is never refused for space.
In whole granules every run is rounded out, neighbours are joined, and each
range is written back whole: the base's bytes with the diff's runs laid over
them. The base is read only where the runs do not cover the range. The choice
is made once for the diff: made for each 8 MiB window, the merge of the guest
below took 300 s where whole granules everywhere took 263 s.
Measured on the two files of such a snapshot (btrfs, compress-force=zstd, base
of 1,032,975 extents, page cache dropped before each run, 16 workers, the
whole file merged one way):
written as merge writes written extents afterwards
one write per run 323 s 2,329,373 38.1 GiB 5,374,185
1 MiB granules 178 s 18,610 111.4 GiB 1,041,076
8 MiB granules 199 s 16,173 126.4 GiB 1,041,213
In `fcvm snapshot create` of the 128 GiB guest on one host, each from a VM
restored and warmed the same way:
build diff merge written extents command
one write per run 38.1 GiB in 2,251,186 runs 554 s 38.1 GiB 5,251,673 14m39s
(main, 16 workers)
this change 37.9 GiB in 2,355,784 runs 263 s 110.8 GiB 1,045,539 10m50s
The second run's line: bytes_merged=40722845696 runs=2355784 writes=20357
bytes_written=118945218560 rounded_bytes=118945218560 base_reads=20132
workers=16 merge_ms=263206. The host's memory pressure (some avg10, sampled
every 5 s) peaked at 12.7 percent during that merge and at 13.2 percent during
the first.
The cost of granules is disk when a VM touched this much: 111 GiB of the 128
were rewritten, 37 GiB after compression, where one write per run took 13.
The "diff merge complete" line gains bytes_written, rounded_bytes (what whole
granules come to) and base_reads. The two free space checks before a snapshot
read the filesystem through the function the merge uses.
Tested (x86_64):
make _test-unit FILTER=-E 'test(/^commands::common::/)'
Summary [ 1.047s] 62 tests run: 62 passed, 1373 skipped
make fmt leaves the tree unchanged
make clippy exit 0
make _test-root FILTER=-E 'test(/test_user_snapshot_from_clone_uses_parent/)'
Summary [ 24.504s] 1 test run: 1 passed, 1785 skipped
Red first, one mutation at a time:
whole granules without room for the memory file once more:
a_merge_writes_run_by_run_when_whole_granules_would_not_leave_room_for_the_file_again
left: (131072, 131072) right: (8192, 131072)
no bound on the bytes whole granules add for each run:
a_diff_whose_runs_are_scattered_is_written_run_by_run_however_many_they_are
left: writes: 2, bytes_written: 3145728, base_reads: 2
right: writes: 3, bytes_written: 12288, base_reads: 0
whole granules for a diff of any number of runs:
a_diff_of_few_runs_is_written_run_by_run_whatever_its_windows_hold
a_merge_holds_the_same_bytes_however_it_is_cut_and_however_many_workers_make_it
a merge that cannot read the free space beside its base:
a_merge_reads_the_free_space_of_the_filesystem_that_holds_its_base
left: (8192, 0) right: (131072, 2)
every whole-granule write reading the base:
a_run_longer_than_a_window_goes_out_in_pieces_that_read_no_base
left: base_reads: 6 right: base_reads: 1
The first three ran before the last two tests were added.
Not run: arm64, a filesystem other than btrfs, and a merge while the source VM
is busy (the source was idle in every run above).
ejc3
force-pushed
the
snapshot-merge-whole-granules
branch
from
October 4, 2026 19:17
7d916a0 to
f140302
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A diff of many small runs is merged in whole 1 MiB granules instead of one write per run.
Follows: #1058 (merged), which runs the merge with workers. Base: main. This is a draft until one question is decided: whether the room rule below, which is a check and not a reservation, is enough (see "What the room does not cover").
The Problem
After #1058 the merge of a restored VM's memory diff still copies each data run with one write. Half of the runs of a large guest's diff are single 4 KiB pages. Writes to one file take that file's lock one at a time, so the workers speed up the reads and not the writes. Each write inside a 128 KiB extent of the reflinked base also splits it.
In
fcvm snapshot createof a 128 GiB guest whose diff held 38.1 GiB in 2,251,186 runs, the merge took 554 s and left a file of 5,251,673 extents, five times its base's, which every later restore reads through. On a second host the same merge took about 941 s, where one thread had taken 745 s.The Solution
A first walk writes nothing and counts the diff's runs. The diff is then written in whole granules when three things hold, and run by run as before when any does not:
MERGE_ROUND_FROM_RUNS).MERGE_ROUND_BELOW).In whole granules every run is rounded out to 1 MiB, neighbours are joined, and each range is written back whole: the base's bytes with the diff's runs laid over them. The base is read only where the runs do not cover the range.
The merge takes no lock and waits for nothing.
Why These Rules
Measurements
The two ways of writing, each applied to the whole file, on the two files of a 128 GiB guest's snapshot (diff of 38.1 GiB in 2,328,992 runs; btrfs,
compress-force=zstd:3, page cache dropped before each run, 16 workers):The cost of granules is disk when a VM touched this much: 111 GiB of the 128 were rewritten, which took 37 GiB after compression where one write per run took 13 GiB.
On the real path,
fcvm snapshot createof the 128 GiB guest on one host, each from a VM restored and warmed the same way:The second row's line:
bytes_merged=40722845696 runs=2355784 writes=20357 bytes_written=118945218560 rounded_bytes=118945218560 base_reads=20132 workers=16 merge_ms=263206. The free space when it ran held the granules and the file once more, so the merge wrote whole granules.During the merge the host's memory pressure (
some avg10in/proc/pressure/memory, sampled every 5 s) peaked at 12.7 percent in the second row's run and at 13.2 percent in the first's. The source VM was resumed and idle in each run.Test Results
Red first, one mutation at a time:
Not run: arm64, a filesystem other than btrfs, and a merge while the source VM is busy.