Skip to main navigation Skip to search Skip to main content

Enhance 3D Visual Grounding through LiDAR and Radar Point Clouds Fusion for Autonomous Driving

  • Runwei Guan
  • , Jianan Liu
  • , Ningwei Ouyang
  • , Shaofeng Liang
  • , Daizong Liu
  • , Xiaolou Sun
  • , Lianqing Zheng
  • , Ming Xu
  • , Tao Huang
  • , Yutao Yue*
  • , Guoqiang Mao
  • , Hui Xiong*
  • *Corresponding author for this work
    • The Hong Kong University and Technology (Guangzhou)
    • Mononai AI
    • Institute of Deep Perception Technology
    • Institute of Deep Perception Technology
    • The Hong Kong University of Science and Technology (Guangzhou)
    • The HongKong University of Science and Technology (Guangzhou)

    Research output: Contribution to journalArticlepeer-review

    Abstract

    Embodied outdoor scene understanding forms the foundation for autonomous agents to perceive, analyze, and react to dynamic driving environments. In contrast to vision-only frameworks, point cloud sensors such as LiDAR provide rich depth and fine-grained 3D representations, while the emerging 4D millimeter-wave radar detects object motions and velocity. By bypassing visual textures, this active dual-sensor combination inherently enables privacy-preserving 3D perception while maintaining direct geometric and kinematic awareness under adverse conditions. The integration of these two modalities provides more flexible querying conditions for natural language, thereby supporting more accurate 3D visual grounding. To this end, we propose TPCNet, the first outdoor 3D visual grounding model upon the paradigm of prompt-guided point cloud sensor combination. TPCNet employs Bidirectional Agent Cross-Attention (BACA) for dynamic, text-aligned feature fusion. Moreover, a Dynamic Gated Graph Fusion (DGGF) module is developed to filter background noise and locate the regions of interest identified by the queries, followed by the C3D-RECHead, which anchors bounding box regression based on the nearest object edge to the ego-vehicle. Experimental results demonstrate that TPCNet achieves state-of-the-art performance on both the Talk2Radar and Talk2Car datasets. We release the code at https://github.com/GuanRunwei/TPCNet.
    Original languageEnglish
    Pages (from-to)1-14
    JournalIEEE Transactions on Intelligent Transportation Systems
    DOIs
    Publication statusPublished - 17 Aug 2026

    Cite this