Abstract
Speech Enhancement aims to effectively separate and restore clean speech signals from noisy inputs, where the performance is critical for downstream applications such as communications, hearing aids, and speech recognition. While current deep learning-based SE methods have achieved remarkable progress in improving overall speech quality, most mainstream architectures prioritize global optimization of time–frequency features, often neglecting the exploitation of intrinsic relationships among channel-dimension features after encoding. This limitation restricts the capability to model and reconstruct fine-grained spectral structures, particularly in frequency bands contributing significantly to human perception. To address this, we propose a Kolmogorov–Arnold Network (KAN)-based Channel Attention Module for Speech Enhancement Networks (KCAM-SENet). The model employs an encoder–decoder backbone integrated with Transformer blocks to capture long-range dependencies. Its core innovation lies in the proposed KAN-based Channel Attention Module (KCAM). This module utilizes a dual-layer channel-spatial attention structure to deeply fuse inter-channel interactive features with non-linear spatial information. Specifically, within the spatial attention branch, we introduce learnable KANs to replace conventional fixed activation functions, adaptively generating more precise attention maps to highlight critical regions in the time–frequency representation. Experimental results on VoiceBank+DEMAND dataset and DNS Challenge 2020 dataset demonstrate that KCAM-SENet outperforms existing state-of-the-art models in terms of overall metrics such as PESQ and CSIG. Furthermore, detailed frequency-band analysis validates the significant advantage of the proposed method in enhancing high-frequency components. Ablation studies also confirm the effectiveness of the joint attention mechanism and the non-linear activation design within KCAM.
| Original language | English |
|---|---|
| Article number | 103425 |
| Journal | Speech Communication |
| Volume | 182 |
| DOIs | |
| Publication status | Published - Jul 2026 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 3 Good Health and Well-being
Keywords
- Channel attention
- Dual branch
- Kolmogorov–Arnold network
- Speech enhancement
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver