Skip to content

converter: fail fast on unreadable svo2 streams; record per-camera svo2 frame counts in metadata.json - #8

Open
anthony-liang-tri wants to merge 5 commits into
mainfrom
converter-svo2-guards
Open

anthony-liang-tri wants to merge 5 commits into
mainfrom
converter-svo2-guards

Conversation

@anthony-liang-tri

Copy link
Copy Markdown

Why

5 recordings in the YAM backfill (3 CleanUpSpill, 1 PlaceTeaBagInTeapot, 1 BreadInTheBag) crashed in _build_lowdim with numpy ValueError: negative dimensions are not allowed. Root cause (probed with pyzed inside the ingest image): one camera's svo2 opens and its index reports 1,273 frames, but every grab() fails from frame 0 — a truncated/corrupt payload. _extract_svo2_synchronized then printed 0 synchronized frames extracted, the start-time alignment trimmed 1–3 leading frames from 0, and the negative frame count reached np.tile.

What

  • _check_initial_grabs: a stream whose first grab fails raises a RuntimeError naming the camera and its index total (instead of silently extracting nothing). A second guard raises if the sync loop extracts 0 frames.
  • _check_aligned_frames: convert_recording refuses a non-positive common frame count after alignment.
  • metadata.json gains svo2_frame_counts (per-camera index totals from the pre-scan) and aligned_frame_count, so the usual off-by-one between cameras is visible downstream (yam-data copies this metadata verbatim into each episode's episode.json).
  • tests/test_converter_guards.py (pure-python, no ZED): 6 tests.

Stacks on converter-no-depth (4 commits, not yet PR'd); only the last commit is this change. Recordings with an unreadable stream are unusable for the 4-camera training presets and are now reported as such by the backfill runner instead of crashing.

convert_recording/convert_task decode svo2 with a ZED NEURAL (or learned-stereo) depth pass that
dominates runtime (~6 min/episode). For RGB+proprio training materializations depth is unused, so add
compute_depth=True (default, unchanged) threaded through _extract_svo2_synchronized and
_build_sequence_metadata. With compute_depth=False: ZED opens DEPTH_MODE.NONE, no predictor, no depth/
written, metadata drops the depth label — RGB + lowdim only, ~6x faster and ~half the storage.
…SDK calibration dir)

Lets a non-root ZED SDK install point at a writable settings dir for per-serial factory
calibration (default /usr/local/zed/settings needs root). Env-gated; no behavior change when unset.
Note: the path must end with a trailing slash - the SDK concatenates '<path>SN<serial>.conf'.
…ding (default True; lets callers skip the NEURAL depth pass)
…; record svo2_frame_counts + aligned_frame_count in metadata.json

An svo2 whose index reports frames but whose first grab fails (truncated/corrupt payload, seen on 5
CleanUpSpill/PlaceTeaBag/BreadInTheBag recordings) used to yield '0 synchronized frames extracted', then
the start-time alignment trimmed leading frames below zero and _build_lowdim crashed with numpy
'negative dimensions are not allowed'. Now the stream is named in a RuntimeError at the first grab (and
again if the sync loop extracts nothing), and convert_recording refuses a non-positive aligned count.
metadata.json gains the per-camera svo2 index totals and the common aligned count so per-camera
mismatches (the usual off-by-one) are visible downstream (yam-data copies this metadata into episode.json).

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant