Haoxuan Li1,* · Sixu Yan1,* · Lianghui Zhu1,* · Xuanlai Tang2 · Shikang Wang1 · Xinggang Wang1,†
1Huazhong University of Science and Technology · 2KEENON Robotics
*Equal contribution · †Corresponding author
Bridge3D equips pretrained 2D Vision–Language–Action models with 3D geometry guidance for precise robotic manipulation. Implicit Fusion augments image features with geometric priors from a 3D foundation model, while Explicit Conditioning grounds action denoising in an explicit 3D semantic field. Layer-wise linear probing identifies where geometric conditioning is most useful, improving learning efficiency. Experiments on RoboTwin 2.0 and six real-world tasks demonstrate improved spatial precision and data efficiency.
@article{li2026bridge3d,
title={Bridge3D: Enabling Vision-Language-Action Models to See and Act in 3D},
author={Li, Haoxuan and Yan, Sixu and Zhu, Lianghui and Tang, Xuanlai and Wang, Shikang and Wang, Xinggang},
journal={arXiv preprint arXiv:2609.24525},
year={2026}
}