Skip to main navigation Skip to search Skip to main content

A comparative Re-assessment of feature extractors for deep speaker embeddings

    • University of Eastern Finland
    • Université de Lorraine

    Research output: Chapter in Book or Report/Conference proceedingConference Proceedingpeer-review

    12 Citations (Scopus)

    Abstract

    Modern automatic speaker verification relies largely on deep neural networks (DNNs) trained on mel-frequency cepstral coefficient (MFCC) features. While there are alternative feature extraction methods based on phase, prosody and long-term temporal operations, they have not been extensively studied with DNN-based methods. We aim to fill this gap by providing extensive re-assessment of 14 feature extractors on VoxCeleb and SITW datasets. Our findings reveal that features equipped with techniques such as spectral centroids, group delay function, and integrated noise suppression provide promising alternatives to MFCCs for deep speaker embeddings extraction. Experimental results demonstrate up to 16.3% (VoxCeleb) and 25.1% (SITW) relative decrease in equal error rate (EER) to the baseline.

    Original languageEnglish
    Title of host publicationInterspeech 2020
    PublisherInternational Speech Communication Association
    Pages3221-3225
    Number of pages5
    ISBN (Print)9781713820697
    DOIs
    Publication statusPublished - 2020
    Event21st Annual Conference of the International Speech Communication Association, INTERSPEECH 2020 - Shanghai, China
    Duration: 25 Oct 202029 Oct 2020

    Publication series

    NameProceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
    Volume2020-October
    ISSN (Print)2308-457X
    ISSN (Electronic)1990-9772

    Conference

    Conference21st Annual Conference of the International Speech Communication Association, INTERSPEECH 2020
    Country/TerritoryChina
    CityShanghai
    Period25/10/2029/10/20

    Keywords

    • Deep speaker embeddings
    • Feature extraction
    • Speaker verification

    Cite this