TY - JOUR
T1 - Enhance 3D Visual Grounding through LiDAR and Radar Point Clouds Fusion for Autonomous Driving
AU - Guan, Runwei
AU - Liu, Jianan
AU - Ouyang, Ningwei
AU - Liang, Shaofeng
AU - Liu, Daizong
AU - Sun, Xiaolou
AU - Zheng, Lianqing
AU - Xu, Ming
AU - Huang, Tao
AU - Yue, Yutao
AU - Mao, Guoqiang
AU - Xiong, Hui
PY - 2026/8/17
Y1 - 2026/8/17
N2 - Embodied outdoor scene understanding forms the foundation for autonomous agents to perceive, analyze, and react to dynamic driving environments. In contrast to vision-only frameworks, point cloud sensors such as LiDAR provide rich depth and fine-grained 3D representations, while the emerging 4D millimeter-wave radar detects object motions and velocity. By bypassing visual textures, this active dual-sensor combination inherently enables privacy-preserving 3D perception while maintaining direct geometric and kinematic awareness under adverse conditions. The integration of these two modalities provides more flexible querying conditions for natural language, thereby supporting more accurate 3D visual grounding. To this end, we propose TPCNet, the first outdoor 3D visual grounding model upon the paradigm of prompt-guided point cloud sensor combination. TPCNet employs Bidirectional Agent Cross-Attention (BACA) for dynamic, text-aligned feature fusion. Moreover, a Dynamic Gated Graph Fusion (DGGF) module is developed to filter background noise and locate the regions of interest identified by the queries, followed by the C3D-RECHead, which anchors bounding box regression based on the nearest object edge to the ego-vehicle. Experimental results demonstrate that TPCNet achieves state-of-the-art performance on both the Talk2Radar and Talk2Car datasets. We release the code at https://github.com/GuanRunwei/TPCNet.
AB - Embodied outdoor scene understanding forms the foundation for autonomous agents to perceive, analyze, and react to dynamic driving environments. In contrast to vision-only frameworks, point cloud sensors such as LiDAR provide rich depth and fine-grained 3D representations, while the emerging 4D millimeter-wave radar detects object motions and velocity. By bypassing visual textures, this active dual-sensor combination inherently enables privacy-preserving 3D perception while maintaining direct geometric and kinematic awareness under adverse conditions. The integration of these two modalities provides more flexible querying conditions for natural language, thereby supporting more accurate 3D visual grounding. To this end, we propose TPCNet, the first outdoor 3D visual grounding model upon the paradigm of prompt-guided point cloud sensor combination. TPCNet employs Bidirectional Agent Cross-Attention (BACA) for dynamic, text-aligned feature fusion. Moreover, a Dynamic Gated Graph Fusion (DGGF) module is developed to filter background noise and locate the regions of interest identified by the queries, followed by the C3D-RECHead, which anchors bounding box regression based on the nearest object edge to the ego-vehicle. Experimental results demonstrate that TPCNet achieves state-of-the-art performance on both the Talk2Radar and Talk2Car datasets. We release the code at https://github.com/GuanRunwei/TPCNet.
UR - https://arxiv.org/abs/2503.08336
U2 - 10.1109/TITS.2026.3717183
DO - 10.1109/TITS.2026.3717183
M3 - Article
SN - 1524-9050
SP - 1
EP - 14
JO - IEEE Transactions on Intelligent Transportation Systems
JF - IEEE Transactions on Intelligent Transportation Systems
ER -