Repository navigation
Cosmos3-Edge: is the shipped LIBERO 10D representation supported by the public forward-dynamics checkpoint? #249
Description
Activity
I can only speak to what the repo pins, not to what the released weights saw in training, but three of these are answerable from the tree. All four files cited below are identical between the revision you pinned,
fa18818, and currentmain, so the line numbers hold on your pin.The 10D layout and
global_raware consistent with each other and both come from the same place.libero_lerobot_dataset.py:12-15definesframe_wise_relativeas the stored 7D delta[dpos(3), drot_axisangle(3), gripper(1)], with only the rotation re-encoded, giving[dpos(3), rot6d(6), gripper(1)]atrotation_space="6d".global_rawis the statistics bucket thataction_normalization="quantile_rot"selects rather than a mode you set yourself, atlibero_lerobot_dataset.py:131-132.For question 1 the recipe is
action_policy_libero_nano.py:196-218, which pins every item on your list in one block:fps=20,image_size=256soconcat_viewis 256x512,action_space="frame_wise_relative",rotation_space="6d",pose_coordinate_frame="native",action_normalization="quantile_rot",format_prompt_as_json=True. Its own comment names the matching data,nvidia/LIBERO_LeRobot_v3at 20 FPS, chosen to match the bundled stats and the 20 Hz eval.That recipe is
mode="wam"on line 205, though, not forward dynamics. The only forward-dynamics post-training recipe in the tree isaction_fd_droid_posttrain.py, which is DROID withmode="forward_dynamics"andaction_space="ee_pose"on lines 184-185. The representation you are asking about is the one the policy recipes use, and the sole FD recipe uses absolute end-effector pose. There is no LIBERO forward-dynamics config inposttrain_config/at either revision.On question 3, every config in that directory builds on
NANO_MODEL_CONFIG, the FD one included.edge_model_config.pyis referenced only from the vision SFT experiments, so I could not find an action-conditioned Edge recipe in-tree.Nothing in the LIBERO loader refuses forward dynamics:
modeis passed to the base dataset and_MODE_CHOICESincludes it (base_dataset.py:25). I have not run it in that mode, since I do not have the dataset locally, so that is the absence of a blocker and not evidence it produces sensible targets.Which leaves the part only NVIDIA can settle, and it is the part you actually asked: whether the released Cosmos3-Edge weights were trained or validated on this LIBERO representation for forward dynamics, and whether 2 and 3 have a planned answer.
Thank you for checking the pinned code and clearly distinguishing the available recipes from the training coverage of the released weights.
This confirms that our LIBERO encoding is consistent with the policy recipe, but does not establish support for forward dynamics with the public Edge checkpoint.
One small clarification: the DROID FD
ee_posepath appears to map tomidtrain, which derives relative poses from the recorded absolute pose sequence. So the source states are absolute, but the model input appears to be relative.For the NVIDIA maintainers, the remaining question is: was this Edge checkpoint trained or validated for LIBERO forward dynamics? If not, is there a supported checkpoint or adaptation path you recommend for this use case?
We will avoid treating loader compatibility as evidence of prediction capability. Thanks again.
I haven't personally trained edged for forward dynamics. @ychao-nvidia may have more information on this.
Hello, I would like to clarify the checkpoint-specific input contract before doing further forward-dynamics experiments.
For
nvidia/Cosmos3-Edgerevisiona9d944e2c6a1bf9f48b92ad16348e70c5f1836ba, is theliberodomain trained/validated for forward dynamics with the shippedframe_wise_relative10D control representation andglobal_rawquantile statistics, or is this loader only infrastructure for post-trained checkpoints?The code being compared is pinned to cosmos-framework revision
fa1881878ab9b234583d89319d56c7903be5161d. I understand that a supported loader and a LIBERO policy post-training recipe do not by themselves establish that the generic Edge checkpoint is trained for this forward-dynamics use.This question is specifically about external candidate actions conditioning future observations, not policy success rates or generic image-to-video quality. Thank you.