Audio-attention discriminative language model for ASR rescoring

Gandhe, Ankur; Rastrow, Ariya

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:1912.03363 (eess)

[Submitted on 6 Dec 2019 (v1), last revised 18 Feb 2020 (this version, v2)]

Title:Audio-attention discriminative language model for ASR rescoring

Authors:Ankur Gandhe, Ariya Rastrow

View PDF

Abstract:End-to-end approaches for automatic speech recognition (ASR) benefit from directly modeling the probability of the word sequence given the input audio stream in a single neural network. However, compared to conventional ASR systems, these models typically require more data to achieve comparable results. Well-known model adaptation techniques, to account for domain and style adaptation, are not easily applicable to end-to-end systems. Conventional HMM-based systems, on the other hand, have been optimized for various production environments and use cases. In this work, we propose to combine the benefits of end-to-end approaches with a conventional system using an attention-based discriminative language model that learns to rescore the output of a first-pass ASR system. We show that learning to rescore a list of potential ASR outputs is much simpler than learning to generate the hypothesis. The proposed model results in 8% improvement in word error rate even when the amount of training data is a fraction of data used for training the first-pass system.

Comments:	4 pages, 1 figure, Accepted at ICASSP 2020
Subjects:	Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD)
Cite as:	arXiv:1912.03363 [eess.AS]
	(or arXiv:1912.03363v2 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.1912.03363

Submission history

From: Ankur Gandhe [view email]
[v1] Fri, 6 Dec 2019 22:09:07 UTC (121 KB)
[v2] Tue, 18 Feb 2020 18:03:03 UTC (234 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Audio-attention discriminative language model for ASR rescoring

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Audio-attention discriminative language model for ASR rescoring

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators