Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
32 changes: 32 additions & 0 deletions .changeset/shallow-export-forward-root-state.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
---
"loro-crdt": patch
"loro-crdt-map": patch
---

Speed up `export({ mode: "shallow-snapshot" })` by up to ~20x on container-heavy documents with a large retained history.

Building the state at the shallow root used to check the live document out
backwards (latest -> root). That reverse diff makes the richtext/list diff
calculators rebuild a full CRDT tracker from empty for every container touched
in the range, which dominated the export cost: on a doc with ~66k containers
and ~720k streaming-edit ops, shallow export at a mid-history root took ~3.2s
versus ~1.5ms for a full snapshot. When at least 64k ops are retained since
the root, the root state is now reconstructed by replaying the pre-root
history forward into a temporary doc, and the latest state is read from the
live store directly without moving the document. The pre-root prefix is bounded
both relative to the tail (at most 16x the retained op count) and absolutely
(at most 1M ops) and decoded payload size (at most 32 MiB, estimated by
walking op payloads — recursing into nested values, style values, and commit
messages — before any value is copied; op counts miss value sizes, since a Map
write is one atom regardless of payload size), so a document whose
pre-root history is huge or byte-heavy and unrelated to the tail keeps the
previous checkout path. The same export drops to
~370ms and the produced blob is slightly smaller (~17% on the same fixture).
With a small retained range — including a root at the latest version — the
previous checkout path is kept: it is then equally fast and peaks at ~4x less
memory, so exporting a lazily imported document no longer materializes its
whole state. Exported blobs remain logically equivalent; detached or
already-shallow source docs also keep the previous code path.

Also fixes shallow snapshot export resurrecting, as an empty entry, a root
container deleted with `deleteRootContainer` before the shallow root.
2 changes: 1 addition & 1 deletion .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@ moon/_build/
moon_*_fuzz_artifacts*/
dhat-heap.json
.DS_Store
node_modules/
node_modules
.idea/
coverage/
trace-*.json
Expand Down
44 changes: 44 additions & 0 deletions context/internal-encoding.md
Original file line number Diff line number Diff line change
Expand Up @@ -212,6 +212,50 @@ path reuses the root bytes without that check. Containers introduced after the
root are not checked again and can survive either in retained operations (`E`)
or as raw/lazy overlay state bytes.

When the source doc is not shallow, its state is already at the latest
version, and at least `MIN_RETAINED_OPS_FOR_FORWARD_ROOT_STATE` (65536; 16 in
unit tests) ops are retained since the root, `export_shallow_snapshot_inner`
builds the root state by replaying pre-root history forward into a temporary
doc (`export_fast_updates_in_range` pre-encoded under the oplog lock, then
imported), not by checking the live doc out backwards: a reverse checkout
makes the richtext/list diff calculators rebuild a full CRDT tracker from
empty per touched container (the `should_rebuild` path in
`RichtextDiffCalculator::calculate_diff`), which dominated shallow export cost
(~20x slower than forward replay on container-heavy docs). Below the retained-ops
threshold the checkout path is used instead: it ties in time around ~8k
retained ops and peaks at ~4x less memory, which matters for lazily imported
docs (exporting right after import must not materialize the whole state).
The fast path pays for re-encoding and replaying the ENTIRE pre-root history,
so the prefix is also gated: `pre_root_ops <= 16 * ops_num`
(`MAX_PRE_ROOT_TO_RETAINED_OPS_RATIO`; measured crossover — fast wins at
ratio 9, loses at 19), `pre_root_ops <= 1_000_000`
(`MAX_PRE_ROOT_OPS_FOR_FORWARD_REPLAY`), and a decoded-byte cap on the prefix
(`MAX_PRE_ROOT_BYTES_FOR_FORWARD_REPLAY`, 32 MiB) because op counts miss value
sizes — a Map write is one atom regardless of how large its Binary/String
value is. The byte leg runs BEFORE encoding: `estimate_ops_content_bytes`
walks op payloads by reference (arena slices are never copied), recursing into
nested `LoroValue::List`/`Map` and counting everything the block encoder
copies: map and style keys, style values, fractional indexes, root container
names, unknown-op OwnedValue payloads (including MarkStart keys and
MarkStart/ListSet values), and commit messages, with a budget-aware early
exit past the cap, while
`export_fast_updates_in_range` slice-copies values into a fresh store — so the
cap must be checked before any prefix bytes are copied.
A huge unrelated prefix with a large tail must stay on the checkout path —
see the `shallow_export_scalar_prefix` and `shallow_export_byte_prefix`
benches.
The replay doc mirrors the live store's root container entries via
`DocState::existing_retention_roots` (a root-only key scan — never
`iter_all_container_ids`, which calls `load_all`) so accessed-but-op-less root
containers still ship, and it receives a copy of the live doc's
`deleted_root_containers` config so roots deleted before the root are dropped
at flush instead of being resurrected as empty entries (roots deleted after
the root keep their at-root content because flush only drops entries whose
value is empty). Detached or already-shallow sources keep the old checkout
path (the reuse branch handles cached roots; a shallow source's trimmed
history cannot be forward-replayed). The forward path never moves the live
doc, so no state restore is needed.

Pre-shallow frontier safety lives in `loro.rs`: `checkout`, `diff`, and
`revert_to` must return `SwitchToVersionBeforeShallowRoot` instead of traversing
history before the shallow root.
Expand Down
4 changes: 4 additions & 0 deletions crates/loro-internal/Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -128,3 +128,7 @@ harness = false
[[bench]]
name = "jsonpath"
harness = false

[[bench]]
name = "shallow_export"
harness = false
164 changes: 164 additions & 0 deletions crates/loro-internal/benches/shallow_export.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,164 @@
use criterion::{criterion_group, criterion_main, Criterion};
use loro_internal::{encoding::ExportMode, version::Frontiers, LoroDoc};
use std::hint::black_box;

/// Build a doc shaped like a real workspace export: a root list of
/// `OUTER_DOCS` maps, each with a nested list of `INNER_ITEMS` maps, each
/// holding `TEXTS_PER_ITEM` short text containers. Every text is written with
/// `OPS_PER_TEXT` single-character inserts to simulate streaming edits.
///
/// With the defaults this produces ~66k containers and ~720k ops.
fn build_structured_doc() -> (LoroDoc, Frontiers) {
const OUTER_DOCS: usize = 200;
const INNER_ITEMS: usize = 30;
const TEXTS_PER_ITEM: usize = 10;
const OPS_PER_TEXT: usize = 12;

let doc = LoroDoc::new_auto_commit();
doc.set_peer_id(1).unwrap();
let root = doc.get_list("docs");
let mut mid = None;
for i in 0..OUTER_DOCS {
let map = root
.insert_container(i, loro_internal::handler::MapHandler::new_detached())
.unwrap();
let items = map
.insert_container("items", loro_internal::handler::ListHandler::new_detached())
.unwrap();
for j in 0..INNER_ITEMS {
let item = items
.insert_container(j, loro_internal::handler::MapHandler::new_detached())
.unwrap();
for k in 0..TEXTS_PER_ITEM {
let text = item
.insert_container(
&format!("t{k}"),
loro_internal::handler::TextHandler::new_detached(),
)
.unwrap();
for n in 0..OPS_PER_TEXT {
text.insert(n, "x", loro_internal::cursor::PosType::Unicode)
.unwrap();
}
}
}
if i == OUTER_DOCS / 2 {
mid = Some(doc.oplog_frontiers());
}
}
(doc, mid.unwrap())
}

fn shallow_export(c: &mut Criterion) {
let (doc, mid_frontiers) = build_structured_doc();
let mut g = c.benchmark_group("shallow_export");
g.sample_size(10);
g.bench_function("full_snapshot", |b| {
b.iter(|| black_box(doc.export(ExportMode::Snapshot).unwrap()))
});
g.bench_function("shallow_snapshot", |b| {
b.iter(|| {
black_box(
doc.export(ExportMode::shallow_snapshot(&mid_frontiers))
.unwrap(),
)
})
});
g.finish();
}

/// A document imported from a snapshot and never read stays lazy: exporting a
/// shallow snapshot at the latest version must not materialize the whole
/// state. Each iteration exports from a freshly imported doc, so this measures
/// the cold path (setup time is excluded).
fn shallow_export_lazy(c: &mut Criterion) {
let (doc, _) = build_structured_doc();
let full = doc.export(ExportMode::Snapshot).unwrap();
let latest = doc.oplog_frontiers();
let mut g = c.benchmark_group("shallow_export_lazy");
g.sample_size(10);
g.bench_function("at_latest", |b| {
b.iter_batched(
|| {
let lazy = LoroDoc::new();
lazy.import(&full).unwrap();
lazy
},
|lazy| black_box(lazy.export(ExportMode::shallow_snapshot(&latest)).unwrap()),
criterion::BatchSize::LargeInput,
)
});
g.finish();
}

/// Regression guard for the forward-replay gate's prefix bound: a huge
/// pre-root prefix of unrelated scalar overwrites must not be re-encoded and
/// replayed just because the retained tail clears the 65536-op threshold.
/// With the prefix/tail ratio and absolute caps this export stays on the
/// checkout path, whose cost is bounded by the tail.
fn shallow_export_scalar_prefix_heavy(c: &mut Criterion) {
const PREFIX_OPS: usize = 2_000_000;

let doc = LoroDoc::new_auto_commit();
doc.set_peer_id(1).unwrap();
let map = doc.get_map("m");
for i in 0..PREFIX_OPS {
map.insert("k", i as i64).unwrap();
}
let f = doc.oplog_frontiers();
// One big-atom insert past the 65536-op retained threshold.
doc.get_text("t")
.insert(
0,
&"x".repeat(70_000),
loro_internal::cursor::PosType::Unicode,
)
.unwrap();

let mut g = c.benchmark_group("shallow_export_scalar_prefix");
g.sample_size(10);
g.bench_function("export", |b| {
b.iter(|| black_box(doc.export(ExportMode::shallow_snapshot(&f)).unwrap()))
});
g.finish();
}

/// Regression guard for the gate's byte bound: a byte-heavy but low-op prefix
/// (few large Map values) must not be replayed just because the retained tail
/// clears the op threshold — a Map write is one atom regardless of value
/// size. With the encoded-byte cap this export stays on the checkout path.
fn shallow_export_byte_prefix_heavy(c: &mut Criterion) {
let doc = LoroDoc::new_auto_commit();
doc.set_peer_id(1).unwrap();
let map = doc.get_map("m");
// 64 distinct 1 MiB values to the same key: 64 atoms, ~64 MiB of prefix
// history. Distinct values matter — identical strings dedup in the arena.
for i in 0..64 {
let big = format!("{i:08}{}", "v".repeat(1 << 20));
map.insert("k", big.as_str()).unwrap();
}
let f = doc.oplog_frontiers();
doc.get_text("t")
.insert(
0,
&"x".repeat(70_000),
loro_internal::cursor::PosType::Unicode,
)
.unwrap();

let mut g = c.benchmark_group("shallow_export_byte_prefix");
g.sample_size(10);
g.bench_function("export", |b| {
b.iter(|| black_box(doc.export(ExportMode::shallow_snapshot(&f)).unwrap()))
});
g.finish();
}

criterion_group!(
benches,
shallow_export,
shallow_export_lazy,
shallow_export_scalar_prefix_heavy,
shallow_export_byte_prefix_heavy
);
criterion_main!(benches);
7 changes: 7 additions & 0 deletions crates/loro-internal/src/arena.rs
Original file line number Diff line number Diff line change
Expand Up @@ -519,6 +519,13 @@ impl SharedArena {
(self.inner.values.lock()[range]).to_vec()
}

/// Borrow the values in `range` without cloning them (unlike
/// [`Self::get_values`], which clones into a fresh `Vec`).
#[inline]
pub fn with_values<R>(&self, range: Range<usize>, f: impl FnOnce(&[LoroValue]) -> R) -> R {
f(&self.inner.values.lock()[range])
}

pub fn convert_single_op(
&self,
container: &ContainerID,
Expand Down
Loading
Loading