TY - JOUR
T1 - Open-Source Conversational AI with SpeechBrain 1.0
AU - Ravanelli, Mirco
AU - Parcollet, Titouan
AU - Moumen, Adel
AU - de Langen, Sylvain
AU - Subakan, Cem
AU - Plantinga, Peter
AU - Wang, Yingzhi
AU - Mousavi, Pooneh
AU - Libera, Luca Della
AU - Ploujnikov, Artem
AU - Paissan, Francesco
AU - Borra, Davide
AU - Zaiem, Salah
AU - Zhao, Zeyu
AU - Zhang, Shucong
AU - Karakasidis, Georgios
AU - Yeh, Sung Lin
AU - Champion, Pierre
AU - Rouhe, Aku
AU - Braun, Rudolf
AU - Mai, Florian
AU - Zuluaga-Gomez, Juan
AU - Mousavi, Seyed Mahed
AU - Nautsch, Andreas
AU - Nguyen, Ha
AU - Liu, Xuechen
AU - Sagar, Sangeet
AU - Duret, Jarod
AU - Mdhaffar, Salima
AU - Laperrière, Gaëlle
AU - Rouvier, Mickael
AU - De Mori, Renato
AU - Estève, Yannick
N1 - Publisher Copyright:
© 2024 Mirco Ravanelli, Titouan Parcollet, et al.
PY - 2024
Y1 - 2024
N2 - SpeechBrain1 is an open-source Conversational AI toolkit based on PyTorch, focused particularly on speech processing tasks such as speech recognition, speech enhancement, speaker recognition, text-to-speech, and much more. It promotes transparency and replicability by releasing both the pre-trained models and the complete “recipes” of code and algorithms required for training them. This paper presents SpeechBrain 1.0, a significant milestone in the evolution of the toolkit, which now has over 200 recipes for speech, audio, and language processing tasks, and more than 100 models available on Hugging Face. SpeechBrain 1.0 introduces new technologies to support diverse learning modalities, Large Language Model (LLM) integration, and advanced decoding strategies, along with novel models, tasks, and modalities. It also includes a new benchmark repository, offering researchers a unified platform for evaluating models across diverse tasks.
AB - SpeechBrain1 is an open-source Conversational AI toolkit based on PyTorch, focused particularly on speech processing tasks such as speech recognition, speech enhancement, speaker recognition, text-to-speech, and much more. It promotes transparency and replicability by releasing both the pre-trained models and the complete “recipes” of code and algorithms required for training them. This paper presents SpeechBrain 1.0, a significant milestone in the evolution of the toolkit, which now has over 200 recipes for speech, audio, and language processing tasks, and more than 100 models available on Hugging Face. SpeechBrain 1.0 introduces new technologies to support diverse learning modalities, Large Language Model (LLM) integration, and advanced decoding strategies, along with novel models, tasks, and modalities. It also includes a new benchmark repository, offering researchers a unified platform for evaluating models across diverse tasks.
KW - Conversational AI
KW - deep learning
KW - open-source
KW - speech processing
UR - https://www.scopus.com/pages/publications/105018580751
M3 - Article
AN - SCOPUS:105018580751
SN - 1532-4435
VL - 25
JO - Journal of Machine Learning Research
JF - Journal of Machine Learning Research
ER -