Skip to main navigation Skip to search Skip to main content

A Text-Guided Cross-Hierarchical Fusion and Multi-Task Learning Framework for Multimodal Sentiment Analysis

  • University of Liverpool
  • Ruijin-XJTLU Intelligent Medicine Institute, Shanghai Jiao Tong University
  • Ruijin-XJTLU Intelligent Medicine Institute, Shanghai Jiao Tong University
  • Zhengzhou University of Light Industry
  • Norwegian University of Science and Technology
  • University College London (UCL)
  • Chengdu University of Information Technology

Research output: Contribution to journalArticlepeer-review

3 Citations (Scopus)

Abstract

Existing multimodal sentiment analysis (MSA) methods have achieved strong performance, but they still face several challenges: (1) insufficient utilization of textual modality information; (2) limited effectiveness in jointly modeling hierarchical multimodal features (including deep and shallow features); and (3) inadequate exploration of the independent characteristics of each modality. These issues may cause models to overlook important emotional patterns in text and hinder effective learning of cross-hierarchical features with salient emotional cues as well as modality-specific characteristics. To address these challenges, we propose a multi-task learning framework that jointly models three unimodal prediction tasks and one multimodal sentiment prediction task. The multimodal branch produces the final prediction output, while the unimodal branches, under the supervision of loss functions, help optimize the network parameters and promote the learning of modality-specific characteristics. The framework contains three innovative modules: (1) the Raw Cross-Modal Information Fusion Module (RCMIF), built upon graph convolution, for learning shallow multimodal representations; (2) the Cleaned Cross-Modal Information Fusion Module (CCMIF), which captures deeper multimodal information via dynamic graph convolution and attention mechanisms; and (3) the Bilinear Attention Deep Independent Characteristics Mining Module (BADIC), which explores unimodal independent characteristics by leveraging bilinear pooling and related techniques. Notably, both BADIC and CCMIF exploit textual guidance, while RCMIF and CCMIF collaboratively learn cross-hierarchical features. Extensive experiments on multiple datasets (MOSI, MOSEI, CH-SIMS, and CH-SIMS-V2) show that the proposed framework achieves improvements of approximately 0.5%-1% in binary classification accuracy and F1 score compared with existing methods.

Original languageEnglish
Article number108995
JournalNeural Networks
Volume202
Issue number108995
DOIs
Publication statusPublished - Oct 2026

Keywords

  • Hierarchical feature fusion
  • Multimodal sentiment analysis
  • Textual information enhancement
  • Unimodal characteristic extraction

Cite this