Skip to content

Cosmos3-Edge: is the shipped LIBERO 10D representation supported by the public forward-dynamics checkpoint? #249

Description

@pome223

Hello, I would like to clarify the checkpoint-specific input contract before doing further forward-dynamics experiments.

For nvidia/Cosmos3-Edge revision a9d944e2c6a1bf9f48b92ad16348e70c5f1836ba, is the libero domain trained/validated for forward dynamics with the shipped frame_wise_relative 10D control representation and global_raw quantile statistics, or is this loader only infrastructure for post-trained checkpoints?

The code being compared is pinned to cosmos-framework revision fa1881878ab9b234583d89319d56c7903be5161d. I understand that a supported loader and a LIBERO policy post-training recipe do not by themselves establish that the generic Edge checkpoint is trained for this forward-dynamics use.

  1. If supported, which checkpoint-specific recipe defines normalization, frame timing, rotation convention, gripper convention, concat-view orientation/resizing, and prompt formatting? Is there a matched LIBERO forward-dynamics input/output example?
  2. If not, is an NVIDIA-released LIBERO forward-dynamics checkpoint available or planned?
  3. Which action-conditioned adaptation recipe is validated for Edge rather than Nano?

This question is specifically about external candidate actions conditioning future observations, not policy success rates or generic image-to-video quality. Thank you.

Activity

  1. harshitwandhare commented on Sep 11, 2026

    @harshitwandhare
    Contributor

    I can only speak to what the repo pins, not to what the released weights saw in training, but three of these are answerable from the tree. All four files cited below are identical between the revision you pinned, fa18818, and current main, so the line numbers hold on your pin.

    The 10D layout and global_raw are consistent with each other and both come from the same place. libero_lerobot_dataset.py:12-15 defines frame_wise_relative as the stored 7D delta [dpos(3), drot_axisangle(3), gripper(1)], with only the rotation re-encoded, giving [dpos(3), rot6d(6), gripper(1)] at rotation_space="6d". global_raw is the statistics bucket that action_normalization="quantile_rot" selects rather than a mode you set yourself, at libero_lerobot_dataset.py:131-132.

    For question 1 the recipe is action_policy_libero_nano.py:196-218, which pins every item on your list in one block: fps=20, image_size=256 so concat_view is 256x512, action_space="frame_wise_relative", rotation_space="6d", pose_coordinate_frame="native", action_normalization="quantile_rot", format_prompt_as_json=True. Its own comment names the matching data, nvidia/LIBERO_LeRobot_v3 at 20 FPS, chosen to match the bundled stats and the 20 Hz eval.

    That recipe is mode="wam" on line 205, though, not forward dynamics. The only forward-dynamics post-training recipe in the tree is action_fd_droid_posttrain.py, which is DROID with mode="forward_dynamics" and action_space="ee_pose" on lines 184-185. The representation you are asking about is the one the policy recipes use, and the sole FD recipe uses absolute end-effector pose. There is no LIBERO forward-dynamics config in posttrain_config/ at either revision.

    On question 3, every config in that directory builds on NANO_MODEL_CONFIG, the FD one included. edge_model_config.py is referenced only from the vision SFT experiments, so I could not find an action-conditioned Edge recipe in-tree.

    Nothing in the LIBERO loader refuses forward dynamics: mode is passed to the base dataset and _MODE_CHOICES includes it (base_dataset.py:25). I have not run it in that mode, since I do not have the dataset locally, so that is the absence of a blocker and not evidence it produces sensible targets.

    Which leaves the part only NVIDIA can settle, and it is the part you actually asked: whether the released Cosmos3-Edge weights were trained or validated on this LIBERO representation for forward dynamics, and whether 2 and 3 have a planned answer.

  2. pome223 commented on Sep 11, 2026

    @pome223
    Author

    Thank you for checking the pinned code and clearly distinguishing the available recipes from the training coverage of the released weights.

    This confirms that our LIBERO encoding is consistent with the policy recipe, but does not establish support for forward dynamics with the public Edge checkpoint.

    One small clarification: the DROID FD ee_pose path appears to map to midtrain, which derives relative poses from the recorded absolute pose sequence. So the source states are absolute, but the model input appears to be relative.

    For the NVIDIA maintainers, the remaining question is: was this Edge checkpoint trained or validated for LIBERO forward dynamics? If not, is there a supported checkpoint or adaptation path you recommend for this use case?

    We will avoid treating loader compatibility as evidence of prediction capability. Thanks again.

  3. fwd4 commented on Sep 11, 2026

    @fwd4
    Collaborator

    I haven't personally trained edged for forward dynamics. @ychao-nvidia may have more information on this.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions