It is mentioned in the model card on the huggingface website that the Cosmos Reason 2
Supports object detection with 2D/3D point localization and bounding box coordinates with reasoning explanations and labels.
Is there an example for that ? Is it possible to input both the RGB and depth images and generate a Scene Graph the objects 3D bounding boxes directly ?
It is mentioned in the model card on the huggingface website that the Cosmos Reason 2
Supports object detection with 2D/3D point localization and bounding box coordinates with reasoning explanations and labels.
Is there an example for that ? Is it possible to input both the RGB and depth images and generate a Scene Graph the objects 3D bounding boxes directly ?