Skip to content

Recommended framework for RL post-training given cosmos-rl is in maintenance mode? #701

Description

@yicwang

We've been exploring RL training of flow-matching based VLA models using cosmos-rl, and it has been working well for us. With this repo now in maintenance mode and future development focused on Cosmos 3, we noticed that cosmos-framework currently covers SFT / action post-training only, and the Cosmos 3 technical report does not mention RL anywhere in the training pipeline. Question: what is the recommended path going forward for sim-in-the-loop RL workloads (GRPO on VLAs / world-action models)? Will RL support land in cosmos-framework, NeMo RL, or elsewhere?

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Fields

    No fields configured for issues without a type.

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions