Skip to main navigation Skip to search Skip to main content

Multimodal fusion for complete motion trajectory reasoning in articulated object manipulation

  • Jie Song
  • , Xing Wu*
  • , Quan Qian
  • , Jianbiao Dai
  • , Jianjia Wang
  • , Jun Song
  • , Qingzhe Cui
  • *Corresponding author for this work
    • Shanghai University
    • Ltd.
    • University of Saint Joseph
    • Hong Kong Baptist University

    Research output: Contribution to journalArticlepeer-review

    Abstract

    The manipulation of articulated objects for part-level motion is crucial due to their prevalence in real-world applications. Although current manipulation methods have improved interaction quality, they share a common issue: neglecting the completeness of motion trajectories. For example, when we use a front-loading drum washing machine, the expected action is to manipulate the door from fully closed to fully open. However, these methods might only result in it being half-open. To tackle this limitation, we introduce a novel framework for optimizing motion trajectories based on multimodal fusion. Specifically, we explicitly model trajectory completeness and propose a motion trajectory construction paradigm (MTCP). This paradigm is applied to a large-scale dataset containing a wide range of articulated objects, generating high-quality motion trajectories for multimodal fusion. Furthermore, to handle trajectory homogeneity, we propose a trajectory enhancement policy (TEP) that enriches the trajectory set by capturing the multimodal distribution of feasible trajectories. Subsequently, to enhance the learning efficiency and task adaptability of Multimodal Large Language Models (MLLMs), we propose a learning strategy for 3D perception inspired by 2D perception (3PI2P), complemented by a progressive reasoning approach. This strategy integrates a dual-branch input design using RGB images and depth maps, combined with six forms of visual question answering tasks, to achieve collaborative reasoning and deep fusion of 2D semantics and 3D geometric information. The robustness and generalizability of the framework are demonstrated through evaluations in both simulation and real-world environments.

    Original languageEnglish
    Article number104659
    JournalInformation Fusion
    Volume137
    DOIs
    Publication statusPublished - Jan 2027

    Keywords

    • Articulated objects
    • Multimodal large language models
    • Part-level motion
    • Trajectory completeness

    Cite this