Repository navigation
converter: fail fast on unreadable svo2 streams; record per-camera svo2 frame counts in metadata.json - #8
Open
anthony-liang-tri wants to merge 5 commits into
Open
anthony-liang-tri wants to merge 5 commits into
anthony-liang-tri wants to merge 5 commits into
Conversation
convert_recording/convert_task decode svo2 with a ZED NEURAL (or learned-stereo) depth pass that dominates runtime (~6 min/episode). For RGB+proprio training materializations depth is unused, so add compute_depth=True (default, unchanged) threaded through _extract_svo2_synchronized and _build_sequence_metadata. With compute_depth=False: ZED opens DEPTH_MODE.NONE, no predictor, no depth/ written, metadata drops the depth label — RGB + lowdim only, ~6x faster and ~half the storage.
…SDK calibration dir) Lets a non-root ZED SDK install point at a writable settings dir for per-serial factory calibration (default /usr/local/zed/settings needs root). Env-gated; no behavior change when unset. Note: the path must end with a trailing slash - the SDK concatenates '<path>SN<serial>.conf'.
…ding (default True; lets callers skip the NEURAL depth pass)
…of aborting the whole task
…; record svo2_frame_counts + aligned_frame_count in metadata.json An svo2 whose index reports frames but whose first grab fails (truncated/corrupt payload, seen on 5 CleanUpSpill/PlaceTeaBag/BreadInTheBag recordings) used to yield '0 synchronized frames extracted', then the start-time alignment trimmed leading frames below zero and _build_lowdim crashed with numpy 'negative dimensions are not allowed'. Now the stream is named in a RuntimeError at the first grab (and again if the sync loop extracts nothing), and convert_recording refuses a non-positive aligned count. metadata.json gains the per-camera svo2 index totals and the common aligned count so per-camera mismatches (the usual off-by-one) are visible downstream (yam-data copies this metadata into episode.json).
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
5 recordings in the YAM backfill (3 CleanUpSpill, 1 PlaceTeaBagInTeapot, 1 BreadInTheBag) crashed in
_build_lowdimwith numpyValueError: negative dimensions are not allowed. Root cause (probed with pyzed inside the ingest image): one camera's svo2 opens and its index reports 1,273 frames, but everygrab()fails from frame 0 — a truncated/corrupt payload._extract_svo2_synchronizedthen printed0 synchronized frames extracted, the start-time alignment trimmed 1–3 leading frames from 0, and the negative frame count reachednp.tile.What
_check_initial_grabs: a stream whose first grab fails raises aRuntimeErrornaming the camera and its index total (instead of silently extracting nothing). A second guard raises if the sync loop extracts 0 frames._check_aligned_frames:convert_recordingrefuses a non-positive common frame count after alignment.metadata.jsongainssvo2_frame_counts(per-camera index totals from the pre-scan) andaligned_frame_count, so the usual off-by-one between cameras is visible downstream (yam-data copies this metadata verbatim into each episode'sepisode.json).tests/test_converter_guards.py(pure-python, no ZED): 6 tests.Stacks on
converter-no-depth(4 commits, not yet PR'd); only the last commit is this change. Recordings with an unreadable stream are unusable for the 4-camera training presets and are now reported as such by the backfill runner instead of crashing.