End-to-end contextual speech recognition using class language models and a token passing decoder
End-to-end modeling (E2E) of automatic speech recognition (ASR) blends all the components of a traditional speech recognition system into a unified model. Although it simplifies training and decoding pipelines, the unified model is hard to adapt when mismatch exists between training and test data. In this work, we focus on contextual speech recognition, which is particularly challenging for E2E models because it introduces significant mismatch between training and test data. To improve the performance in the presence of complex contextual information, we propose to use class-based language models(CLM) that can populate the classes with contextdependent information in real-time. To enable this approach to scale to a large number of class members and minimize search errors, we propose a token passing decoder with efficient token recombination for E2E systems for the first time. We evaluate the proposed system on general and contextual ASR, and achieve relative 62% Word Error Rate(WER) reduction for contextual ASR without hurting performance for general ASR. We show that the proposed method performs well without modification of the decoding hyper-parameters across tasks, making it a general solution for E2E ASR.
Authors

Are you an author of this paper? Check the Twitter handle we have for you is correct.

Zhehuai Chen (add twitter)
Mahaveer Jain (add twitter)
Yongqiang Wang (add twitter)
Michael L. Seltzer (add twitter)
Christian Fuegen (add twitter)
Ask The Authors

Ask the authors of this paper a question or leave a comment.

Read it. Rate it.
#1. Which part of the paper did you read?

#2. The paper contains new data or analyses that is openly accessible?
#3. The conclusion is supported by the data and analyses?
#4. The conclusion is of scientific interest?
#5. The result is likely to lead to future research?

Github
User:
None (add)
Repo:
None (add)
Stargazers:
0
Forks:
0
Open Issues:
0
Network:
0
Subscribers:
0
Language:
None
Youtube
Link:
None (add)
Views:
0
Likes:
0
Dislikes:
0
Favorites:
0
Comments:
0
Other
Sample Sizes (N=):
Inserted:
Words Total:
Words Unique:
Source:
Abstract:
None
12/05/18 06:01PM
4,568
1,641
Tweets
eiichiroi: 積み論文が一本増えた > End-to-end contextual speech recognition using class language models and a token passing decoder https://t.co/BAakiJvLL9
MuhammadThalhah: RT @Memoirs: End-to-end contextual speech recognition using class language models and a token passing decoder. https://t.co/j09MiIyd7a
desantis: RT @Memoirs: End-to-end contextual speech recognition using class language models and a token passing decoder. https://t.co/j09MiIyd7a
Memoirs: End-to-end contextual speech recognition using class language models and a token passing decoder. https://t.co/j09MiIyd7a
arxivml: "End-to-end contextual speech recognition using class language models and a token passing decoder", Zhehuai Chen, M… https://t.co/NVn9BJt4Ow
BrundageBot: End-to-end contextual speech recognition using class language models and a token passing decoder. Zhehuai Chen, Mahaveer Jain, Yongqiang Wang, Michael L. Seltzer, and Christian Fuegen https://t.co/Jp7L1tlfIe
Images
Related