Nonparametric Masked Language Modeling
Sewon MinWeijia ShiMike LewisXilun ChenWen-tau YihHannaneh HajishirziLuke Zettlemoyer Affiliation: University of Washington Affiliation: Meta AI Affiliation: Meta AI Affiliation: Meta AI Affiliation: Allen Institute for
Abstract
Existing language models (LMs) predict tokens with a softmax over a finite vocabulary, which can make it difficult to predict rare tokens or phrases. We introduce NpM, the first nonparametric masked language model that replaces this softmax with a nonparametric distribution over every phrase in a reference corpus. NpM fills in the [MASK] solely from retrieving a token from a text corpus. We show that NpM can be efficiently trained with a contrastive objective and an in-batch approximation to full corpus retrieval. Zero-shot evaluation on 16 tasks including classification, fact probing and question answering demonstrates that NpM outperforms significantly larger parametric models, either with or without a retrieve-and-generate approach. It is particularly better at dealing with rare patterns (word senses or facts) and predicting rare or nearly unseen words (e.g., non-Latin script). We release the model and code at github.com/facebookresearch/NPM.
原文 arXiv:2212.01349;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2212.01349v2