FROM SENONES TO CHENONES: TIED CONTEXT-DEPENDENT GRAPHEMES FOR HYBRID SPEECH RECOGNITION
Abstract
There is an implicit assumption that traditional hybrid approaches for automatic speech recognition (ASR) cannot directly model graphemes and need to rely on phonetic lexicons to get competitive performance, especially on English which has poor grapheme-phoneme correspondence. In this work, we show for the first time that, on English, hybrid ASR systems can in fact model graphemes effectively by leveraging tied context-dependent graphemes, i.e., chenones. Our chenone-based systems significantly outperform equivalent senone baselines by 4.5% to 11.1% relative on three different English datasets. Our results on Librispeech are state-of-the-art compared to other hybrid approaches and competitive with previously published end-to-end numbers. Further analysis shows that chenones can better utilize powerful acoustic models and large training data, and require context- and position-dependent modeling to work well. Chenone-based systems also outperform senone baselines on proper noun and rare word recognition, an area where the latter is traditionally thought to have an advantage. Our work provides an alternative for end-to-end ASR and establishes that hybrid systems can be improved by dro
中文速览
传统混合语音识别系统(hybrid ASR)长期被认为必须依赖专家标注的音素词典才能在英语上取得有竞争力的表现,因为英语字母与发音的对应关系极为不规则。这项研究提出了一种叫做"chenone"的单元——即捆绑的上下文相关字素(tied context-dependent grapheme)——让混合系统可以直接用字母建模而完全抛弃音素知识。在三个英语数据集(包括公开的Librispeech及两个内部大规模数据集)上,基于chenone的系统比等价的音素基线(senone)相对降低词错误率4.5%到11.1%,在Librispeech上达到当时混合方法的最优水平,与端到端方法也具有可比性。进一步分析表明,chenone能更充分地利用大模型容量和海量训练数据,在专有名词和罕见词识别上同样超过音素基线——而后者过去被认为在这方面具有天然优势。这项工作打破了"混合系统必须依赖语言学知识"的固有认知,为构建更简单、更高效的语音识别系统提供了新路径。
原文 arXiv:1910.01493;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1910.01493v2