Skip to main navigation Skip to search Skip to main content

KCAM-SENet: Speech enhancement network with KAN-based channel attention module

  • Linhui Sun*
  • , Zhaowei Ding
  • , Yuhang Qin
  • , Tingting Wang
  • , Shengchen Li
  • , Xi Shao
  • , Eng Siong Chng
  • *Corresponding author for this work
  • Nanjing University of Posts and Telecommunications
  • Nanyang Technological University

Research output: Contribution to journalArticlepeer-review

Abstract

Speech Enhancement aims to effectively separate and restore clean speech signals from noisy inputs, where the performance is critical for downstream applications such as communications, hearing aids, and speech recognition. While current deep learning-based SE methods have achieved remarkable progress in improving overall speech quality, most mainstream architectures prioritize global optimization of time–frequency features, often neglecting the exploitation of intrinsic relationships among channel-dimension features after encoding. This limitation restricts the capability to model and reconstruct fine-grained spectral structures, particularly in frequency bands contributing significantly to human perception. To address this, we propose a Kolmogorov–Arnold Network (KAN)-based Channel Attention Module for Speech Enhancement Networks (KCAM-SENet). The model employs an encoder–decoder backbone integrated with Transformer blocks to capture long-range dependencies. Its core innovation lies in the proposed KAN-based Channel Attention Module (KCAM). This module utilizes a dual-layer channel-spatial attention structure to deeply fuse inter-channel interactive features with non-linear spatial information. Specifically, within the spatial attention branch, we introduce learnable KANs to replace conventional fixed activation functions, adaptively generating more precise attention maps to highlight critical regions in the time–frequency representation. Experimental results on VoiceBank+DEMAND dataset and DNS Challenge 2020 dataset demonstrate that KCAM-SENet outperforms existing state-of-the-art models in terms of overall metrics such as PESQ and CSIG. Furthermore, detailed frequency-band analysis validates the significant advantage of the proposed method in enhancing high-frequency components. Ablation studies also confirm the effectiveness of the joint attention mechanism and the non-linear activation design within KCAM.

Original languageEnglish
Article number103425
JournalSpeech Communication
Volume182
DOIs
Publication statusPublished - Jul 2026

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 3 - Good Health and Well-being
    SDG 3 Good Health and Well-being

Keywords

  • Channel attention
  • Dual branch
  • Kolmogorov–Arnold network
  • Speech enhancement

Cite this