TY - GEN
T1 - Augmentation Invariant Discrete Representation for Generative Spoken Language Modeling
AU - Gat, Itai
AU - Kreuk, Felix
AU - Nguyen, Tu Anh
AU - Lee, Ann
AU - Copet, Jade
AU - Synnaeve, Gabriel
AU - Dupoux, Emmanuel
AU - Adi, Yossi
N1 - Publisher Copyright: © IWSLT 2023.All rights reserved.
PY - 2023
Y1 - 2023
N2 - Generative Spoken Language Modeling research focuses on optimizing speech Language Models (LMs) using raw audio recordings without accessing any textual supervision. Such speech LMs usually operate over discrete units obtained from quantizing internal representations of self-supervised models. Although such units show impressive modeling results, their robustness capabilities have not been extensively investigated. This work focuses on improving the invariance of discrete input representations to non-spoken augmentations for generative spoken language modeling. First, we formally define how to measure the robustness of such representations to various signal variations that do not alter the spoken information (e.g., time-stretch). Next, we empirically demonstrate how current state-of-the-art representation models lack robustness to such variations. To overcome this, we propose an effective and efficient method to learn invariant discrete speech representation for generative spoken language modeling. The proposed approach is based on applying a set of signal transformations to the speech signal and optimizing the model using an iterative pseudo-labeling scheme. Our method significantly improves over the evaluated baselines when considering encoding and modeling metrics. We additionally evaluate our method on the speech-to-speech translation task, considering Spanish-English and French-English translations, and show the proposed approach outperforms the evaluated baselines.
AB - Generative Spoken Language Modeling research focuses on optimizing speech Language Models (LMs) using raw audio recordings without accessing any textual supervision. Such speech LMs usually operate over discrete units obtained from quantizing internal representations of self-supervised models. Although such units show impressive modeling results, their robustness capabilities have not been extensively investigated. This work focuses on improving the invariance of discrete input representations to non-spoken augmentations for generative spoken language modeling. First, we formally define how to measure the robustness of such representations to various signal variations that do not alter the spoken information (e.g., time-stretch). Next, we empirically demonstrate how current state-of-the-art representation models lack robustness to such variations. To overcome this, we propose an effective and efficient method to learn invariant discrete speech representation for generative spoken language modeling. The proposed approach is based on applying a set of signal transformations to the speech signal and optimizing the model using an iterative pseudo-labeling scheme. Our method significantly improves over the evaluated baselines when considering encoding and modeling metrics. We additionally evaluate our method on the speech-to-speech translation task, considering Spanish-English and French-English translations, and show the proposed approach outperforms the evaluated baselines.
UR - https://www.scopus.com/pages/publications/85174969361
M3 - Conference contribution
T3 - 20th International Conference on Spoken Language Translation, IWSLT 2023 - Proceedings of the Conference
SP - 465
EP - 477
BT - 20th International Conference on Spoken Language Translation, IWSLT 2023 - Proceedings of the Conference
A2 - Salesky, Elizabeth
A2 - Federico, Marcello
A2 - Carpuat, Marine
T2 - 20th International Conference on Spoken Language Translation, IWSLT 2023
Y2 - 13 July 2023 through 14 July 2023
ER -