Projects per year
Abstract
Existing multimodal sentiment analysis (MSA) methods have achieved strong performance, but they still face several challenges: (1) insufficient utilization of textual modality information; (2) limited effectiveness in jointly modeling hierarchical multimodal features (including deep and shallow features); and (3) inadequate exploration of the independent characteristics of each modality. These issues may cause models to overlook important emotional patterns in text and hinder effective learning of cross-hierarchical features with salient emotional cues as well as modality-specific characteristics. To address these challenges, we propose a multi-task learning framework that jointly models three unimodal prediction tasks and one multimodal sentiment prediction task. The multimodal branch produces the final prediction output, while the unimodal branches, under the supervision of loss functions, help optimize the network parameters and promote the learning of modality-specific characteristics. The framework contains three innovative modules: (1) the Raw Cross-Modal Information Fusion Module (RCMIF), built upon graph convolution, for learning shallow multimodal representations; (2) the Cleaned Cross-Modal Information Fusion Module (CCMIF), which captures deeper multimodal information via dynamic graph convolution and attention mechanisms; and (3) the Bilinear Attention Deep Independent Characteristics Mining Module (BADIC), which explores unimodal independent characteristics by leveraging bilinear pooling and related techniques. Notably, both BADIC and CCMIF exploit textual guidance, while RCMIF and CCMIF collaboratively learn cross-hierarchical features. Extensive experiments on multiple datasets (MOSI, MOSEI, CH-SIMS, and CH-SIMS-V2) show that the proposed framework achieves improvements of approximately 0.5%-1% in binary classification accuracy and F1 score compared with existing methods.
| Original language | English |
|---|---|
| Article number | 108995 |
| Journal | Neural Networks |
| Volume | 202 |
| Issue number | 108995 |
| DOIs | |
| Publication status | Published - Oct 2026 |
Keywords
- Hierarchical feature fusion
- Multimodal sentiment analysis
- Textual information enhancement
- Unimodal characteristic extraction
Projects
- 4 Active
-
Suzhou Industrial Park Affective Computing & Interactive Health Interdisciplinary Innovation (Research) Platform
Xu, Z. (PI), lim, E. (Team member), Leach, M. (Team member), Yue, Y. (Team member), Selig, T. (Team member), Wang, W. (Team member), Zhang, X. (Team member), Zhu, X. (Team member), Pan, Y. (Team member), Xiang, N. (Team member), Zhang, H. (Team member), Wang, Y. (Team member), Chen, Y. (Team member), Zhang, C. (Team member), Liu, P. (Team member), Wang, J. (Team member), Sun, Q. (Team member), Li, Y. (Team member), 黎秋宇 (Team member), 刘禹含 (Team member), 高一凡 (Team member), 朱铭徽 (Team member), 白万里 (Team member), 张吉儿 (Team member), 徐豪谡 (Team member) & 胡晓松 (Team member)
1/10/25 → 30/09/27
Project: Governmental Research Project
-
Advancing AI-Assisted Elderly Care: Emotional Monitoring, Edge Intelligence, and Privacy-Focused Interaction Systems
Xu, Z. (PI)
1/06/25 → 31/05/28
Project: Internal Research Project
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver