Abstract
Depression is a prevalent mental disorder, and early detection and diagnosis are crucial for its prevention and treatment. Speech-based depression detection represents an efficient and convenient approach within the current landscape of computer-aided detection methods. However, challenges remain in effectively and reliably extracting features and classifying speech patterns to distinguish individuals with depression from those without. This paper introduces an audio feature set for depression analysis, referred to as SJTU-LWDLab DACD. Based on this feature set, we propose a novel method for identifying patients with depression using summed graph convolutional networks to mitigate inaccuracies that arise from the loss of spatial features, such as height and depth, during the structured fusion of multiple depression audio features. Experimental results demonstrate that the accuracy of depression recognition in speech can reach 92.4%. The method proposed in this paper provides objective indicators and a foundation for the auxiliary identification of depression.
| Original language | English |
|---|---|
| Pages (from-to) | 309-320 |
| Number of pages | 12 |
| Journal | Annals of the New York Academy of Sciences |
| Volume | 1550 |
| Issue number | 1 |
| DOIs | |
| Publication status | Published - Aug 2025 |
| Externally published | Yes |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 3 Good Health and Well-being
Keywords
- audio feature
- DACD data set
- depression detection
- SGCNs
- structured fusion
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver