Skip to main navigation Skip to search Skip to main content

Open-Source Conversational AI with SpeechBrain 1.0

  • Mirco Ravanelli
  • , Titouan Parcollet
  • , Adel Moumen
  • , Sylvain de Langen
  • , Cem Subakan
  • , Peter Plantinga
  • , Yingzhi Wang
  • , Pooneh Mousavi
  • , Luca Della Libera
  • , Artem Ploujnikov
  • , Francesco Paissan
  • , Davide Borra
  • , Salah Zaiem
  • , Zeyu Zhao
  • , Shucong Zhang
  • , Georgios Karakasidis
  • , Sung Lin Yeh
  • , Pierre Champion
  • , Aku Rouhe
  • , Rudolf Braun
  • Florian Mai, Juan Zuluaga-Gomez, Seyed Mahed Mousavi, Andreas Nautsch, Ha Nguyen, Xuechen Liu, Sangeet Sagar, Jarod Duret, Salima Mdhaffar, Gaëlle Laperrière, Mickael Rouvier, Renato De Mori, Yannick Estève
  • Concordia University
  • Mila - Quebec Artificial Intelligence Institute
  • University of Montreal
  • Samsung
  • University of Cambridge
  • Avignon Université
  • Université Laval
  • Zaion
  • Fondazione Bruno Kessler
  • Aalto University
  • University of Bologna
  • Ecole Nationale Superieure des Telecommunications
  • University of Edinburgh
  • Institut national de recherche en informatique et en automatique
  • Silo AI
  • Idiap
  • KU Leuven
  • EPFL
  • University of Trento
  • Research Organization of Information and Systems, National Institute of Informatics
  • Saarland University
  • McGill University

Research output: Contribution to journalArticlepeer-review

50 Citations (Scopus)

Abstract

SpeechBrain1 is an open-source Conversational AI toolkit based on PyTorch, focused particularly on speech processing tasks such as speech recognition, speech enhancement, speaker recognition, text-to-speech, and much more. It promotes transparency and replicability by releasing both the pre-trained models and the complete “recipes” of code and algorithms required for training them. This paper presents SpeechBrain 1.0, a significant milestone in the evolution of the toolkit, which now has over 200 recipes for speech, audio, and language processing tasks, and more than 100 models available on Hugging Face. SpeechBrain 1.0 introduces new technologies to support diverse learning modalities, Large Language Model (LLM) integration, and advanced decoding strategies, along with novel models, tasks, and modalities. It also includes a new benchmark repository, offering researchers a unified platform for evaluating models across diverse tasks.

Original languageEnglish
JournalJournal of Machine Learning Research
Volume25
Publication statusPublished - 2024
Externally publishedYes

Keywords

  • Conversational AI
  • deep learning
  • open-source
  • speech processing

Cite this