Skip to main navigation Skip to search Skip to main content

Multi-scale local-global fusion network with temporal attention for speech emotion recognition

Research output: Contribution to journalArticlepeer-review

Abstract

Speech Emotion Recognition (SER) is essential for affective computing, intelligent dialogue, and mental health assessment. However, the dynamic nature of speech signals poses challenges for accurately capturing both local acoustic cues and long-range emotional context. To address this, we propose a novel architecture named Multi-scale Local-Global Fusion Network (MLGFNet). MLGFNet features three core modules: (1) a Local Feature Extraction Module (LFEM), which proposes Inception-style multi-branch depthwise convolutions to extract emotional patterns at different temporal resolutions; and (2) a Global Key Context Focusing Module (GKCF), which presents hierarchical strip convolutions to generate frame-wise attention maps, allowing the network to highlight emotionally salient frames across the utterance. Then (3) a new Attention-based Fusion Mechanism is designed to adaptively integrate the outputs from (1) and (2). We evaluate the proposed MLGFNet on four public SER datasets: IEMOCAP, RAVDESS, SAVEE, and EMOVO. The experimental results show that MLGFNet consistently outperforms competitive baselines in terms of unweighted and weighted average recall. Ablation studies verify the effectiveness of LFEM and GKCF individually and jointly. Furthermore, t-Distributed Stochastic Neighbor Embedding(t-SNE) visualizations demonstrate that MLGFNet learns more separable and emotion-aware feature representations. These findings highlight MLGFNet’s robustness and interpretability, making it a promising solution for speech emotion recognition.

Original languageEnglish
Article number478
JournalNeural Computing and Applications
Volume38
Issue number12
DOIs
Publication statusPublished - Jun 2026

Cite this