ISCA Archive Interspeech 2004
ISCA Archive Interspeech 2004

Vocabulary and language model adaptation using information retrieval

Brigitte Bigi, Yan Huang, Renato De Mori

The goal of vocabulary optimization is to construct a vocabulary with exactly those words that are the most likely to appear in the test data. We will present a new approach to reduce the out-of-vocabulary (OOV) rate by adapting the vocabulary model during the ASR process. This method can also be used for the statistial language model (SLM) adaptation. An information retrieval system is used after the first pass of the ASR system to obtain a set of relevant documents. These documents are then used to generate the new vocabulary and/or corpus. In this paper, we propose a new retrieving method well-adapted for this purpose. Experiments were carried out on French with a 28% OOV rate reduction. Experiments were also carried out on English for the SLM adaptation, with 7.9% perplexity reduction, and minor WER improvement.


doi: 10.21437/Interspeech.2004-489

Cite as: Bigi, B., Huang, Y., Mori, R.D. (2004) Vocabulary and language model adaptation using information retrieval. Proc. Interspeech 2004, 1361-1364, doi: 10.21437/Interspeech.2004-489

@inproceedings{bigi04_interspeech,
  author={Brigitte Bigi and Yan Huang and Renato De Mori},
  title={{Vocabulary and language model adaptation using information retrieval}},
  year=2004,
  booktitle={Proc. Interspeech 2004},
  pages={1361--1364},
  doi={10.21437/Interspeech.2004-489}
}