Website: https://dsta022.github.io/Loop-Engineering-for-VLA/.
An end-to-end toolkit for collecting, merging, auditing, enriching, and iteratively improving multimodal LeRobot datasets for vision-language-action models. It supports standard RGB-only LeRobot datasets and RGB-D datasets with lossless depth sidecars, and it treats RGB video, robot actions, states, metadata, task semantics, language annotations, and human feedback as one auditable engineering loop: Capture · Audit · Enrich · Improve.
The offline pipeline never imports LeRobot or camera SDKs. The recorder in
record/ is a self-contained, framework-free data collection layer that talks
to SO-100/SO-101 arms and cameras directly and writes LeRobot v3.0 datasets:
record/rgb_record/: RGB-only recording.record/rgbd_record/: Orbbec / Intel RealSense RGB-D recording, plus human-guided policy feedback.record/visuo-tactile_record/: visuo-tactile recording, in progress.
Recorded datasets flow straight into the merge, audit, clean, enrich, feedback, and upload steps below. Three guiding rules run through every stage:
- Conservative decisions. Only objective structural, numerical, or decoding
failures become
drop; semantic and threshold-based findings arereview; a check with missing preconditions recordsnot_checkedinstead of passing. - Sidecars, never in-place edits. Depth, feedback provenance, language, training weights, and reports live outside the dataset; source data is read-only.
- Pluggable, model-free by default. Evaluators, annotators, triggers, and
the train/evaluate/deploy trio load by
module:attribute; the base install downloads no weights.
The merged RGB-D VLA dataset is published on Hugging Face as DerekLX/lerobot_derek_depth. Its root contains the LeRobot subdirectories directly:
data/
depth_sidecar/
meta/
videos/
Samples from one pick_up_cups_dataset episode with two cups in the scene.
Left to right: global robot collection view, wrist camera view, and front depth visualization.
If the video does not render in your Markdown viewer, open task_demo.mp4 directly.
| Front camera, 9 sampled episodes | Wrist camera, 9 sampled episodes | Front depth, 9 sampled episodes |
|---|---|---|
![]() |
![]() |
![]() |
pipeline/ # entry points: merge, audit, clean, feedback build,
# language enrichment, policy iteration, HF upload
record/ # framework-free recorder: recorder/ library, camera_profiles,
# rgb_record/, rgbd_record/ (+ feedback), visuo-tactile_record/
tools/ # studios, depth visualization, completeness check,
# threshold calibration
utils/ # shared audit, metadata, media, and semantic helpers
semantic_backends/ # built-in --semantic-evaluator plugins (VLA checkpoint scoring)
language_backends/ # built-in --annotator plugins (kinematic segment annotator)
policy_backends/ # reference trainer / evaluator / deployer plugins
policy_improvement/ # feedback capture, intervention, and policy iteration contracts
tests/ # unit tests with synthetic datasets
configs/ # example configuration files
docs/ # project website (GitHub Pages source), plugin contracts,
# audit strategy, and the operating playbook
.github/workflows/ # tests on Linux and Windows; GitHub Pages deployment
assets/ # README media
data/ # local datasets (not tracked by git)
Python 3.12 is recommended. Create an environment, install the project in
editable mode, and add ffmpeg/ffprobe for video decoding and trimming:
conda create -n vla_data_check python=3.12 -y
conda activate vla_data_check
conda install -c conda-forge ffmpeg -y
python -m pip install -e ".[record]"The editable install exposes console commands and lets every script run by path from any directory:
| Command | Script | Purpose |
|---|---|---|
vla-merge |
pipeline/dataset_merge.py |
merge datasets into one real dataset |
vla-audit |
pipeline/data_quality_audit.py |
conservative quality audit |
vla-clean |
pipeline/dataset_clean_build.py |
build a clean dataset from a drop list |
vla-enrich |
pipeline/trajectory_language_enrichment.py |
language sidecars for approved episodes |
vla-feedback-build |
pipeline/feedback_dataset_build.py |
weighted feedback training dataset |
vla-iterate |
pipeline/policy_iteration.py |
train, evaluate, promote, deploy |
vla-push |
pipeline/hf_push.py |
upload to Hugging Face |
vla-completeness |
tools/dataset_completeness_check.py |
structure-only check |
vla-calibrate |
tools/threshold_calibration.py |
review-reason precision and threshold recommendation |
vla-depth-vis |
tools/depth_visualize.py |
depth PNG statistics and previews |
vla-task-studio, vla-language-studio |
tools/*_studio.py |
local windows for profiles and enrichment |
Optional extras: .[realsense], .[orbbec] for camera SDKs and .[policy]
for the lerobot runtime used by the VLA checkpoint evaluator, the feedback
recorder's in-loop policy, and the offline policy evaluator. Faster Hugging
Face uploads: $env:HF_XET_HIGH_PERFORMANCE="1".
Camera hardware facts live once in record/camera_profiles.yaml; select a set
at record time with --camera_profile=<name>, override any field inline
(--dataset.num_episodes=5), or add --mock for an offline dry run with
synthetic cameras and joints. Calibration uses lerobot's cache location, so
arms calibrated with lerobot work unchanged; add --calibrate to a teleoperate
command to calibrate a new arm. Recording keys: Right = end episode, Left =
re-record, Esc = stop. Set display_data: true to open a live preview window
per camera (depth is colorized next to RGB; needs a GUI build of OpenCV).
Recording into an existing dataset root is refused unless you pass
--resume=true, which appends episodes.
python .\record\rgbd_record\find_cameras.py opencv
python .\record\rgb_record\rgb_teleoperate.py --config_path=.\record\rgb_record\configs\head_wrist_rgb_teleoperate.yaml
python .\record\rgb_record\rgb_record.py --config_path=.\record\rgb_record\configs\head_wrist_rgb_record.yamlDepth-enabled cameras are rejected on this path. The output is a standard
LeRobot RGB dataset with data/, meta/, and videos/.
pip install pyorbbecsdk2 # and/or: pip install pyrealsense2
python .\record\rgbd_record\find_cameras.py orbbec
python .\record\rgbd_record\find_cameras.py realsense
python .\record\rgbd_record\record.py --config_path=.\record\rgbd_record\configs\record_rgbd.yamlSupported backends are orbbec (Femto Bolt), intelrealsense (D405 / D435 /
D435i), and opencv. record_rgbd_realsense.yaml is a ready all-RealSense
rig. Every camera with use_depth: true writes lossless uint16 PNG depth
frames under depth_sidecar/ with a manifest at
meta/rgbd_vla_depth_recording.json. See record/README.md.
python ./pipeline/dataset_merge.py --src-root . --out-dir .\lerobot_derek_depth --dataset-glob "*_dataset" --copy-mode copyThe merge rewrites episode, frame, and task indices, metadata, parquet
tables, video references, and depth sidecars into one physical dataset with
real files. Episode metadata split across several
meta/episodes/chunk-*/file-*.parquet shards is read as a whole. Add
--overwrite to replace an existing output.
python ./pipeline/data_quality_audit.py .\lerobot_derek_depth `
--depth-check header --video-check decode --video-content-check sample `
--sample-frames 8 --depth-sample-frames 8 --workers 4Deterministic checks cover dataset metadata, episode ranges, index and
timestamp integrity, action/state numerics, depth PNG structure and content,
RGB video existence, ffprobe dimensions and codec, full decode, and sampled
black/white/near-constant/frozen content confined to each episode's own time
range. Distribution outliers, exact trajectory duplicates, and low motion are
review-only. --workers runs ffprobe, decode, and content sampling for unique
video files in parallel before the episode loop; verdicts are unchanged.
Outputs in quality_audit_reports/<dataset_name>/ (or --out-dir):
quality_report.csv # one row per episode: status, reasons, not_checked, metrics_json
quality_summary.json # global status, thresholds, audit_script_version, fingerprint
keep_episodes.txt / review_episodes.txt / drop_episodes.txt
Semantic evaluation is expensive, so run the audit in two stages during
iteration: a deterministic pass first, then --semantic-only on top of the
saved report:
python ./pipeline/data_quality_audit.py .\lerobot_derek_depth --semantic-only `
--semantic-evaluator my_quality_backend:create_evaluator `
--semantic-failure-threshold 0.5 --semantic-pass-threshold 0.9The rerun preserves every structural finding, replaces the previous
semantic_* results, and records audit_mode: semantic_reuse. Reuse is refused when the dataset fingerprint, episode set, or
lengths changed, and --feedback-sidecar is rejected in that mode. Run one
final single-pass audit before publishing. The full six-stage procedure is in
docs/audit_operations_playbook.md; the
decision policy is in docs/data_quality_audit_strategy.md.
python ./tools/task_semantic_studio.pyThe studio turns a detailed task description into a validated, editable
profile: task key, name, instruction, aliases, objects, ordered stages with
observable conditions and weights, success and failure criteria, preferred
video keys, and optional threshold overrides. Profiles are saved as JSON and
passed to the audit and the enrichment step with --task-profile-dir. With a
profile directory configured, an episode whose task text matches no profile,
or that contains several distinct tasks, is sent to review. The built-in
parser is a conservative heuristic; install a stronger parser as a
module:attribute plugin.
An optional evaluator receives the task text, episode-scoped video segments, the Task Profile, and the action/state streams, and returns a progress curve and/or a success probability:
class MySemanticEvaluator:
name = "my_video_language_evaluator"
def evaluate(self, episode):
return {"progress_curve": [0.05, 0.30, 0.62], "success_probability": 0.41,
"failure_stage": "grasp", "details": {"counterfactual_margin": 0.35}}python ./pipeline/data_quality_audit.py .\lerobot_derek_depth `
--semantic-evaluator my_quality_backend:create_evaluator --semantic-config .\semantic_evaluator.json `
--semantic-failure-threshold 0.5 --semantic-pass-threshold 0.9 --task-profile-dir .\task_profilesDecision bands are conservative: below the failure threshold is
semantic_low_score_review, below the pass threshold is
semantic_uncertain_score_review, and only scores at or above the pass
threshold can enter keep. Evaluator errors go to review, structural drops
skip inference, and three optional reliability contracts demote unstable
scores: sampling dispersion (score_samples / score_std), the
counterfactual margin over deliberately corrupted instructions, and per-view
agreement with --semantic-per-view. Every field, rule, and threshold is
documented in docs/plugins.md.
Built-in backend. semantic_backends.vla_checkpoint_evaluator:create_evaluator
turns a trained lerobot policy checkpoint into a data-quality scorer. At
sampled anchors the policy predicts an action chunk from the camera frames,
joint state, and task text; the score measures how much better than a
hold-position baseline it reproduces the demonstrated actions (1.0 reproduces
the demonstration). The same anchors are re-scored under corrupted
instructions to fill the counterfactual margin.
python ./pipeline/data_quality_audit.py .\lerobot_derek_depth --semantic-only `
--semantic-evaluator semantic_backends.vla_checkpoint_evaluator:create_evaluator `
--semantic-config .\configs\vla_checkpoint_evaluator.example.jsonIt needs the policy runtime (pip install -e ".[policy]"), so run this stage
in the training environment. A low score can also mean the episode is out of
distribution for the current policy; calibrate the thresholds before trusting
automatic passes.
A second group of semantic checks needs no model. Each is review-only and
records not_checked when its inputs are missing:
- Progress-curve shape (needs an evaluator):
semantic_progress_regression_review,semantic_progress_idle_tail_review,semantic_progress_idle_head_review. - Stage ordering (needs an evaluator and a Task Profile):
semantic_stage_order_violation_review,semantic_stage_skipped_review,semantic_stage_unknown_review,semantic_stage_dwell_outlier_review. - Action/language consistency (needs nothing):
semantic_motion_gripper_pattern_reviewcompares the verb in the task text with the observed gripper actuation, andsemantic_motion_vertical_direction_reviewchecks directional verbs against--motion-vertical-index. The gripper channel is auto-detected or set with--motion-gripper-index;--motion-check noneturns the group off.
Counts per check are written to the semantic_consistency block of
quality_summary.json so each check's precision can be measured on its own.
Review reasons and semantic bands are conservative defaults, not calibrated classifiers. After a human has labelled a set of episodes, measure them:
python ./tools/threshold_calibration.py `
--quality-report .\quality_audit_reports\lerobot_derek_depth\quality_report.csv `
--labels .\lerobot_derek_depth_labels.csvThe labels file is a CSV with episode_index and a label column (reject,
fail, bad, drop, 0, false, no mean rejected). The tool prints the
precision and recall of every review and drop reason, the keep/review/drop
confusion against the labels, and recommends --semantic-failure-threshold
and --semantic-pass-threshold values that meet --precision-target and
--pass-false-rate. Results are written to calibration_report.json.
The third layer describes what happens inside an approved trajectory. It runs after quality review so annotation compute is spent only on usable data:
python ./pipeline/trajectory_language_enrichment.py .\lerobot_derek_depth `
--task-profile-dir .\task_profiles `
--annotator language_backends.kinematic_annotator:create_annotator `
--approved-episodes .\approved_episodes.txtBy default only keep episodes are processed; --approved-episodes adds
human-approved review episodes, and drop is always rejected. The quality
report defaults to quality_audit_reports/<dataset_name>/quality_report.csv,
and its quality_summary.json must belong to the selected dataset. Outputs are
external sidecars under language_enrichment_outputs/<dataset_name>/:
per-episode annotation JSON files, language_annotations.jsonl,
subtasks.jsonl, events.jsonl, state_descriptions.jsonl,
subtask_boundary_alignment.jsonl, and enrichment_manifest.json.
An annotator returns time-aligned subtasks, events, state descriptions, and a
summary; every annotation is validated for frame bounds, ordering,
non-overlap, required text, confidence ranges, and Task Profile stage names,
and subtask boundaries are checked against kinematic change points
(language_subtask_boundary_unaligned_review). The built-in kinematic
annotator segments episodes at gripper transitions and arm near-stops and
labels them with Task Profile stages; treat it as a baseline for a
language-model annotator. python ./tools/trajectory_language_studio.py opens
a window for the same settings. Contract details: docs/plugins.md.
The fourth layer collects targeted human feedback from states a deployed policy visits, keeps control provenance outside the dataset, and gates the next policy on evidence plus human approval.
python ./record/rgbd_record/feedback_record.py `
--config_path=./record/rgbd_record/configs/record_rgbd_feedback.yaml `
--policy.path=./outputs/train/run/checkpoints/last/pretrained_modelKeys: i pauses the policy, aligns the teleoperator, and begins Recovery;
c marks the first Correction frame; i again completes the intervention,
saves the episode, and returns control to the policy; esc requests a clean
stop that never cuts an intervention in half. Capture modes in the config:
corrections_only (recommended; only Recovery/Correction frames are stored and
training-eligible), event_buffer (keeps a bounded pre-event window; h
saves a manual event; an optional automatic_trigger plugin can flag windows
in a background thread without transferring control), and continuous.
Autonomous frames are never expert targets.
Handover to the human is verified before control transfers. handover: capability drives an actuated leader to the follower pose; handover: position_match supports a passive leader: the follower holds still while the
operator moves the leader by hand, and control transfers only once every joint
is within handover_tolerance. Policy inference sits behind
recorder.policy_runner.PolicyRunner; the default adapter runs any lerobot
checkpoint, and --mock uses a stub policy.
Each session writes feedback_outputs/<dataset>/<session>/ with
session_manifest.json, frame_feedback.parquet, episode_feedback.jsonl, and
intervention_segments.jsonl.
python ./pipeline/data_quality_audit.py ./data/so100_rgbd_feedback `
--feedback-sidecar ./feedback_outputs/so100_rgbd_feedback/session-001
python ./pipeline/feedback_dataset_build.py `
--base-dataset ./data/original_clean_dataset `
--feedback-dataset ./data/so100_rgbd_feedback `
--feedback-sidecar ./feedback_outputs/so100_rgbd_feedback/session-001 `
--quality-report ./quality_audit_reports/so100_rgbd_feedback/quality_report.csv `
--approved-episodes ./approved_feedback_episodes.txt `
--language-enrichment ./language_enrichment_outputs/so100_rgbd_feedback `
--out-dir ./iteration_datasets/iteration-001The audit validates sidecar alignment (episode lengths, contiguous frames,
intervention ranges, phase labels, training eligibility) and sends new
pending feedback to review until a human approval list exists. The builder
accepts corrections-only episodes, refuses episodes with autonomous frames
until they are segmented, and writes training_sidecars/training_weights.parquet
(base demonstrations 1.0, Recovery 1.0, Correction 2.0 by default) plus a
manifest confirming that no autonomous failure action became an expert target.
python ./pipeline/policy_iteration.py ./configs/policy_iteration.example.jsonThree plugins own the model-specific work. The reference backends in
policy_backends/ run a configurable training command, score the checkpoint
offline on a held-out dataset with the VLA checkpoint evaluator, and publish
it into a target directory with an atomic current.json pointer. Promotion
requires the configured success rate, intervention rate, and critical-failure
limits plus a separate human approval file; deployment is skipped when any
gate fails, and iteration_manifest.json records every stage. The offline
evaluator's success_rate is a proxy computed from demonstrations; supply
measured intervention_rate and critical_failures from real trials.
python ./tools/dataset_completeness_check.py .\lerobot_derek_depthRuns the structural checks only. Useful right after a merge or clean build.
python ./pipeline/dataset_clean_build.py .\lerobot_derek_depth `
--drop-list .\quality_audit_reports\lerobot_derek_depth\drop_episodes.txt `
--out-dir .\lerobot_derek_depth_clean --video-mode trim-reencodeThe source dataset is preserved. trim-reencode physically removes dropped
video segments; copy-referenced is faster but may keep unreferenced segments
inside copied files. The output's cleaning_manifest.json records the source,
the kept, reviewed, and dropped episodes with their reasons, the index
mappings, and the audit version, time, and thresholds that produced the drop
list.
$env:HF_TOKEN="hf_your_write_token" # or: huggingface-cli login
$env:HF_XET_HIGH_PERFORMANCE="1"
python ./pipeline/hf_push.py --dry-run
python ./pipeline/hf_push.py --dataset-dir .\lerobot_derek_depth_clean --dataset-name lerobot_derek_depth_cleanThe token comes from HF_TOKEN or the cached login; it is never stored in the
repository. Uploads validate the dataset layout, refuse junction, symlink, or
reparse paths, and use resumable large-folder uploads when available.
python ./tools/depth_visualize.py .\lerobot_derek_depth\depth_sidecar --stats
python ./tools/depth_visualize.py .\lerobot_derek_depth\depth_sidecar --out .\depth_vispython -m unittest discover -s tests
python -m unittest discover -s record/recorder/tests -t .Both suites use synthetic datasets only; tests that need ffmpeg skip
themselves when it is missing. The same suites run on Linux and Windows in
GitHub Actions for every push and pull request.
See CONTRIBUTING.md for the development setup, the conventions (conservative decisions, sidecars, framework-free core), and the pull-request checklist. Changes are tracked in CHANGELOG.md.
@dataset{dereklx_lerobot_derek_depth_2026,
author = {DerekLX},
title = {lerobot_derek_depth},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/datasets/DerekLX/lerobot_derek_depth}
}This project is released under the MIT License.




