Repository navigation
Correct stale and incorrect benchmark documentation - #62
Merged
estebanzimanyi merged 1 commit intoAug 20, 2026
Merged
estebanzimanyi merged 1 commit into
estebanzimanyi merged 1 commit into
Conversation
The published brussels_sf0.1 and sf0.2 archive sizes were understated by roughly 100x (5.5 MB / 9.6 MB against actual 539 MB / 937 MB); the other six rows in the two dataset tables are correct. berlinmod_portability_export() writes seven CSV files, not five: the README omitted query_periods.csv and query_regions.csv, described trips.csv as WKT when it is hex-EWKB, and described query_points.csv as WKT when it is EWKT. The example call now passes the output SRID explicitly rather than relying on the 4326 default, matching the canonical example in berlinmod_export.sql:169. The README pointed at MobilitySpark's berlinmod/run_mbdb.sh and run_mduck.sh; neither exists in that repository. The cross-platform runners live here, in BerlinMOD/benchmarks/batch/bench/. Three runner comments referenced README sections that were never written -- "Three-tier index framework" (bench_mbdb.sh, bench_mduck.sh, bench_mspark.sh) and "NxN mitigations on Spark" (bench_mspark.sh). Both sections are now written in benchmarks/batch/README.md, documenting the tier semantics as the three runners actually implement them. bench_mduck.sh gated its Tiers 2 and 3 on MobilityDuck PRs #143 and #144, both of which were closed without being merged. The actual constraint is MobilityDuck #285: before it, CREATE INDEX ... USING TRTREE on a file-backed database failed at commit with "The implementation of this index WAL serialization does not exist."
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The dataset tables state the sizes of the published archives:
brussels_sf0.1.zipis 539 MB andbrussels_sf0.2.zipis 937 MB.The portability-export section lists all seven CSV files that
berlinmod_portability_export()writes, includingquery_periods.csvandquery_regions.csv, and gives each one's encoding:trips.csvcarries the trip as hex-EWKB,query_points.csvandquery_regions.csvcarry geometries as EWKT. The example call passes the output SRID explicitly, so it produces the cross-platform dataset rather than depending on the default, and matches the canonical example inberlinmod_export.sql.The pointer to the cross-platform runners names their location in this repository,
BerlinMOD/benchmarks/batch/bench/, holdingbench_mbdb.sh,bench_mduck.shandbench_mspark.sh.benchmarks/batch/README.mdcarries the two sections the runner headers cite. "Three-tier index framework" documents the tier semantics each runner implements: PostgreSQL exposes tier 0 through 3, DuckDB tiers 1 through 3, and Spark runs at tier 1 with no--tierflag, since native spatial indexes are a PostgreSQL and DuckDB capability. "NxN mitigations on Spark" documents the fourTrips×Tripsqueries and the rule by which a Spark-optimised variant is preferred over the portable form where one is present.bench_mduck.shstates the constraint its tiers 2 and 3 carry: a MobilityDuck build whose TRTREE index serializes to disk. WhereCREATE INDEX ... USING TRTREEfails on a file-backed database, the fallbacks keep the run going at tier-1 acceleration and the messages name the rebuild that enables the tier.Documentation and comments only, across two Markdown files and the comment and message text of one runner script.