An Analysis of Neural Language Modeling at Multiple Scales
Stephen Merity Affiliation: Salesforce Research, Palo Alto, CA – 94301 Correspondence to: Nitish Shirish Keskar Affiliation: Salesforce Research, Palo Alto, CA – 94301 Richard Socher Affiliation: Salesforce Research, Palo Alto, CA – 94301
Abstract
Many of the leading approaches in language modeling introduce novel, complex and specialized architectures. We take existing state-of-the-art word level language models based on LSTMs and QRNNs and extend them to both larger vocabularies as well as character-level granularity. When properly tuned, LSTMs and QRNNs achieve state-of-the-art results on character-level (Penn Treebank, enwik8) and word-level (WikiText-103) datasets, respectively. Results are obtained in only 12 hours (WikiText-103) to 2 days (enwik8) using a single modern GPU.
原文 arXiv:1803.08240;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1803.08240v1