TY - JOUR
T1 - Multi-scale local-global fusion network with temporal attention for speech emotion recognition
AU - Ma, Wenning
AU - Jin, Nanlin
N1 - Publisher Copyright:
© The Author(s), under exclusive licence to Springer-Verlag London Ltd., part of Springer Nature 2026.
PY - 2026/6
Y1 - 2026/6
N2 - Speech Emotion Recognition (SER) is essential for affective computing, intelligent dialogue, and mental health assessment. However, the dynamic nature of speech signals poses challenges for accurately capturing both local acoustic cues and long-range emotional context. To address this, we propose a novel architecture named Multi-scale Local-Global Fusion Network (MLGFNet). MLGFNet features three core modules: (1) a Local Feature Extraction Module (LFEM), which proposes Inception-style multi-branch depthwise convolutions to extract emotional patterns at different temporal resolutions; and (2) a Global Key Context Focusing Module (GKCF), which presents hierarchical strip convolutions to generate frame-wise attention maps, allowing the network to highlight emotionally salient frames across the utterance. Then (3) a new Attention-based Fusion Mechanism is designed to adaptively integrate the outputs from (1) and (2). We evaluate the proposed MLGFNet on four public SER datasets: IEMOCAP, RAVDESS, SAVEE, and EMOVO. The experimental results show that MLGFNet consistently outperforms competitive baselines in terms of unweighted and weighted average recall. Ablation studies verify the effectiveness of LFEM and GKCF individually and jointly. Furthermore, t-Distributed Stochastic Neighbor Embedding(t-SNE) visualizations demonstrate that MLGFNet learns more separable and emotion-aware feature representations. These findings highlight MLGFNet’s robustness and interpretability, making it a promising solution for speech emotion recognition.
AB - Speech Emotion Recognition (SER) is essential for affective computing, intelligent dialogue, and mental health assessment. However, the dynamic nature of speech signals poses challenges for accurately capturing both local acoustic cues and long-range emotional context. To address this, we propose a novel architecture named Multi-scale Local-Global Fusion Network (MLGFNet). MLGFNet features three core modules: (1) a Local Feature Extraction Module (LFEM), which proposes Inception-style multi-branch depthwise convolutions to extract emotional patterns at different temporal resolutions; and (2) a Global Key Context Focusing Module (GKCF), which presents hierarchical strip convolutions to generate frame-wise attention maps, allowing the network to highlight emotionally salient frames across the utterance. Then (3) a new Attention-based Fusion Mechanism is designed to adaptively integrate the outputs from (1) and (2). We evaluate the proposed MLGFNet on four public SER datasets: IEMOCAP, RAVDESS, SAVEE, and EMOVO. The experimental results show that MLGFNet consistently outperforms competitive baselines in terms of unweighted and weighted average recall. Ablation studies verify the effectiveness of LFEM and GKCF individually and jointly. Furthermore, t-Distributed Stochastic Neighbor Embedding(t-SNE) visualizations demonstrate that MLGFNet learns more separable and emotion-aware feature representations. These findings highlight MLGFNet’s robustness and interpretability, making it a promising solution for speech emotion recognition.
UR - https://link.springer.com/article/10.1007/s00521-026-12157-1
UR - https://www.scopus.com/pages/publications/105041605236
U2 - 10.1007/s00521-026-12157-1
DO - 10.1007/s00521-026-12157-1
M3 - Article
AN - SCOPUS:105041605236
SN - 0941-0643
VL - 38
JO - Neural Computing and Applications
JF - Neural Computing and Applications
IS - 12
M1 - 478
ER -