We've been exploring RL training of flow-matching based VLA models using cosmos-rl, and it has been working well for us. With this repo now in maintenance mode and future development focused on Cosmos 3, we noticed that cosmos-framework currently covers SFT / action post-training only, and the Cosmos 3 technical report does not mention RL anywhere in the training pipeline. Question: what is the recommended path going forward for sim-in-the-loop RL workloads (GRPO on VLAs / world-action models)? Will RL support land in cosmos-framework, NeMo RL, or elsewhere?
We've been exploring RL training of flow-matching based VLA models using cosmos-rl, and it has been working well for us. With this repo now in maintenance mode and future development focused on Cosmos 3, we noticed that cosmos-framework currently covers SFT / action post-training only, and the Cosmos 3 technical report does not mention RL anywhere in the training pipeline. Question: what is the recommended path going forward for sim-in-the-loop RL workloads (GRPO on VLAs / world-action models)? Will RL support land in cosmos-framework, NeMo RL, or elsewhere?