Skip to content

Correct stale and incorrect benchmark documentation - #62

Merged
estebanzimanyi merged 1 commit into
MobilityDB:masterfrom
estebanzimanyi:fix/berlinmod-doc-corrections
Aug 20, 2026
Merged

estebanzimanyi merged 1 commit into
MobilityDB:masterfrom
estebanzimanyi:fix/berlinmod-doc-corrections

Conversation

@estebanzimanyi

Copy link
Copy Markdown
Member

The dataset tables state the sizes of the published archives: brussels_sf0.1.zip is 539 MB and brussels_sf0.2.zip is 937 MB.

The portability-export section lists all seven CSV files that berlinmod_portability_export() writes, including query_periods.csv and query_regions.csv, and gives each one's encoding: trips.csv carries the trip as hex-EWKB, query_points.csv and query_regions.csv carry geometries as EWKT. The example call passes the output SRID explicitly, so it produces the cross-platform dataset rather than depending on the default, and matches the canonical example in berlinmod_export.sql.

The pointer to the cross-platform runners names their location in this repository, BerlinMOD/benchmarks/batch/bench/, holding bench_mbdb.sh, bench_mduck.sh and bench_mspark.sh.

benchmarks/batch/README.md carries the two sections the runner headers cite. "Three-tier index framework" documents the tier semantics each runner implements: PostgreSQL exposes tier 0 through 3, DuckDB tiers 1 through 3, and Spark runs at tier 1 with no --tier flag, since native spatial indexes are a PostgreSQL and DuckDB capability. "NxN mitigations on Spark" documents the four Trips × Trips queries and the rule by which a Spark-optimised variant is preferred over the portable form where one is present.

bench_mduck.sh states the constraint its tiers 2 and 3 carry: a MobilityDuck build whose TRTREE index serializes to disk. Where CREATE INDEX ... USING TRTREE fails on a file-backed database, the fallbacks keep the run going at tier-1 acceleration and the messages name the rebuild that enables the tier.

Documentation and comments only, across two Markdown files and the comment and message text of one runner script.

The published brussels_sf0.1 and sf0.2 archive sizes were understated by
roughly 100x (5.5 MB / 9.6 MB against actual 539 MB / 937 MB); the other six
rows in the two dataset tables are correct.

berlinmod_portability_export() writes seven CSV files, not five: the README
omitted query_periods.csv and query_regions.csv, described trips.csv as WKT
when it is hex-EWKB, and described query_points.csv as WKT when it is EWKT.
The example call now passes the output SRID explicitly rather than relying on
the 4326 default, matching the canonical example in berlinmod_export.sql:169.

The README pointed at MobilitySpark's berlinmod/run_mbdb.sh and run_mduck.sh;
neither exists in that repository. The cross-platform runners live here, in
BerlinMOD/benchmarks/batch/bench/.

Three runner comments referenced README sections that were never written --
"Three-tier index framework" (bench_mbdb.sh, bench_mduck.sh, bench_mspark.sh)
and "NxN mitigations on Spark" (bench_mspark.sh). Both sections are now
written in benchmarks/batch/README.md, documenting the tier semantics as the
three runners actually implement them.

bench_mduck.sh gated its Tiers 2 and 3 on MobilityDuck PRs #143 and #144,
both of which were closed without being merged. The actual constraint is
MobilityDuck #285: before it, CREATE INDEX ... USING TRTREE on a file-backed
database failed at commit with "The implementation of this index WAL
serialization does not exist."
@estebanzimanyi
estebanzimanyi merged commit d633cd5 into MobilityDB:master Aug 20, 2026
2 checks passed
@estebanzimanyi
estebanzimanyi deleted the fix/berlinmod-doc-corrections branch August 21, 2026 09:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant