Speechbert: an Audio-And-Text Jointly Learned Language Model for End-To-End Spoken Question Answering

SpeechBERT: An Audio-and-text Jointly Learned Language Model for End-to-end Spoken Question Answering Yung-Sung Chuang Chi-Liang Liu Hung-yi Lee Lin-shan Lee College of Electrical Engineering and Computer Science, National Taiwan University fchuangyungsung,liangtaiwan1230,[email protected], [email protected] Abstract swers. Fine-grained information is also important to predict the exact position of the answer span from a very long context. This While various end-to-end models for spoken language under- paper is probably the first known attempt to try to perform such standing tasks have been explored recently, this paper is prob- a difficult SQA task with an end-to-end approach. ably the first known attempt to challenge the very difficult Substantial improvements in text question answering task of end-to-end spoken question answering (SQA). Learning (TQA) tasks trained and tested on text data have been observed from the very successful BERT model for various text process- after the large-scale self-supervised pre-trained language mod- ing tasks, here we proposed an audio-and-text jointly learned els appeared, such as BERT [10] and GPT [11]. Instead of SpeechBERT model. This model outperformed the conven- learning the TQA tasks from scratch, these models were first tional approach of cascading ASR with the following text ques- pre-trained on a large unannotated text corpus to learn self- tion answering (TQA) model on datasets including ASR errors supervised representations for the general purpose and then in answer spans, because the end-to-end model was shown to fine-tuned on the downstream TQA dataset. Results compa- be able to extract information out of audio data before ASR rable to human performance on SQuAD datasets [12, 13] were produced errors. When ensembling the proposed end-to-end obtained in this way. However, previous work [14] indicated model with the cascade architecture, even better performance that such an approach may not be easily extendable to the SQA was achieved. In addition to the potential of end-to-end SQA, task on audio data by directly cascading an ASR module in the the SpeechBERT can also be considered for many other spo- front. Not only the ASR caused serious problems, but very of- ken language understanding tasks just as BERT for many text ten the true answer spans included name entities or OOV words processing tasks. which cannot be correctly recognized at all, thus cannot be iden- 1. Introduction tified by the following TQA module trained on text. That is why Various spoken language processing tasks, such as transla- all questions with recognition errors in answer spans were dis- tion [1], retrieval [2], summarization [3] and understanding [4] carded in the previous work [14]. This led to the end-to-end have been very successful with a standard cascade architecture: approach to SQA proposed here. an ASR front-end module transforming the speech signals into The BERT model [10] useful in TQA tasks was able to text form, followed by the downstream task module (such as transform word tokens into contextualized embeddings carrying translation) trained on text taking the ASR output as normal text plenty of information. For SQA task, it is certainly desirable to input. However, the end-to-end approach trying to consider the transform audio words (audio signals for word tokens) also into two modules as a whole is always attractive for the following such embeddings, but much more challenging. The same word reasons. The two modules in the cascade architecture locally token can have millions of different audio signal realizations in optimize the two tasks with different criteria, while the end-to- different utterances. The boundaries for audio words in utter- end approach may obtain globally optimized performance for ances are not available. The next problem is even much more the overall task. The ASR module minimizes the WER, which challenging. BERT can learn semantic information of a word is not necessarily directly proportional to the performance mea- token based on its context in text form. Audio words in au- sure of the overall task. Much information is inevitably lost dio data have context only in audio form, which is much more when the speech signals are transformed into text with errors, noisy, confusing, unpredictable, and difficult to handle. Learn- and the errors can’t be recovered in the following module. The ing semantics from audio context is really hard. end-to-end approach allows the possibility of capturing infor- Audio Word2Vec [15] was the first effort to transform audio mation directly from the speech signals not shown in the ASR words with known boundaries into embeddings carrying pho- output and offering overall performance less limited by the ASR netic information only, no semantics at all. Speech2Vec [16] accuracy. then tried to imitate the training process of skip-gram or CBOW Some spoken language tasks such as translation, retrieval, in Word2Vec [17] to extract some semantic features. It was and understanding (intent classification and slot filling) have proposed [18] to learn to jointly segment the audio words out been achieved with end-to-end approaches [5, 6, 7, 8, 9], al- of utterance and extract the embeddings. Other efforts then though it remains difficult to obtain significantly better results tried to align the audio word embeddings with text word embed- than the standard cascade architecture. These tasks are primar- dings [19, 20]. Some more recent works [21, 22, 23, 24, 25, 26] ily sentence-level, in which extracting some local information even tried to obtain embeddings for audio signals using ap- in a short utterance may be adequate. Spoken question answer- proaches very similar to BERT, but primarily extracting acous- ing (SQA) considered here is known as a much more difficult tic information with very limited semantics. These approaches problem. The inputs to the SQA task are much longer spoken may be able to extract some semantics with audio word embed- paragraphs. In addition to understanding the literal meaning, dings, but the level of semantics obtained was still very far from the global information in the paragraphs needs to be organized, that required for the SQA tasks. and sophisticated reasoning is usually required to find the an- In this paper, we propose an audio-and-text jointly learned Audio data sets Initial Phonetic-semantic Text BERT Text data sets Speech Joint Embedding for Audio Pre-training features (audio words) Words (2.2) (2.1) Reconstruction Loss in (1) SpeechBERT SpeechBERT Pre-training Fine-tuning on SQA (2.4) (2.3) Audio Encoder Audio Decoder L1 Loss in (2) Figure 1: Overall training process for SpeechBERT. Speech features ... SpeechBERT model for the end-to-end SQA task. Speech- BERT Word Embedding Layer BERT is a pre-trained model learned from audio and text audio word datasets, so as to be able to produce embeddings for audio words, and these embeddings can be properly aligned with the Figure 2: Training procedure for the Initial Phonetic-Semantic embeddings for the corresponding text words offered by a typ- Joint Embedding. After training, the the encoded vector (z in ical text BERT model trained with text datasets. Standard pre- red) obtained here is used to train the SpeechBERT. training and fine-tuning as text BERT are also performed. When where k is the index for training audio words. This process en- used in the SQA task, performance comparable to or better than ables the vector z to capture the phonetic structure information the cascade architecture was obtained. of the audio words but not semantics at all, which is not ade- 2. SpeechBERT for End-to-end SQA quate for the goal here. So we make the vector z be constrained by a L1-distance loss (lower right) further: Here we assume the training audio datasets include the ground X truth transcripts, and an off-the-shelf ASR engine is available. LL1 = kz − Emb(t)k1 ; (2) So we can segment the audio training data into audio words (au- k dio signals for the underlying word tokens) by forced alignment. where t is the token label for the audio word x and Emb is the For audio testing data we can use the ASR engine to obtain au- embedding layer of the Text BERT trained in Sec. 2.1. In this dio words including their boundaries too, although with errors. way, we use the token label of the audio word to access the se- The overall training process for the SpeechBERT for end-to- mantic information about the considered audio word extracted end SQA proposed here is shown in Fig. 1. We first use the text by the Text BERT as long as it is within the vocabulary of the dataset to pre-train a Text BERT (upper right block, Sec. 2.1), Text BERT. So, the autoencoder model learns to keep the pho- based on which we train the initial phonetic-semantic joint em- netic structure of the audio word so as to reconstruct the origi- bedding with the audio words in the training audio dataset (up- nal speech features x as much as possible, but at the same time per left block, Sec. 2.2). The training of SpeechBERT (or a it learns to fit to the embedding distribution of the Text BERT shared BERT model for both text and audio) is then in the which carries plenty of semantic information for the word to- bottom of the figure, including pre-training (Sec 2.3) and fine- kens. This makes it possible for the model to learn a joint em- tuning on SQA (Sec. 2.4). The details are given below.

Speechbert: an Audio-And-Text Jointly Learned Language Model for End-To-End Spoken Question Answering

Malware Classification with BERT

Bros:Apre-Trained Language Model for Understanding Textsin Document

Information Extraction Based on Named Entity for Tourism Corpus

NLP with BERT: Sentiment Analysis Using SAS® Deep Learning and Dlpy Doug Cairns and Xiangxiang Meng, SAS Institute Inc

Unified Language Model Pre-Training for Natural

Question Answering by Bert

Distilling BERT for Natural Language Understanding

Arxiv:2104.05274V2 [Cs.CL] 27 Aug 2021 NLP ﬁeld Due to Its Excellent Performance in Various Tasks

Efficient Training of BERT by Progressively Stacking

Word Sense Disambiguation with Transformer Models

Improving the Prosody of RNN-Based English Text-To-Speech Synthesis by Incorporating a BERT Model

Recombining Frames and Roles in Frame-Semantic Parsing