From Word to Sense Embeddings: A Survey on Vector Representations of Meaning
\nameJose Camacho-Collados \addrSchool of Computer Science and Informatics Cardiff University United Kingdom \AND\nameMohammad Taher Pilehvar \addrSchool of Computer Engineering Iran University of Science and Technology Tehran, Iran
Abstract
Over the past years, distributed semantic representations have proved to be effective and flexible keepers of prior knowledge to be integrated into downstream applications. This survey focuses on the representation of meaning. We start from the theoretical background behind word vector space models and highlight one of their major limitations: the meaning conflation deficiency, which arises from representing a word with all its possible meanings as a single vector. Then, we explain how this deficiency can be addressed through a transition from the word level to the more fine-grained level of word senses (in its broader acceptation) as a method for modelling unambiguous lexical meaning. We present a comprehensive overview of the wide range of techniques in the two main branches of sense representation, i.e., unsupervised and knowledge-based. Finally, this survey covers the main evaluation procedures and applications for this type of representation, and provides an analysis of four of its important aspects: interpretability, sense granularity, adaptability to different domains and compositionality.
中文速览
词义表示(word sense representation)研究长期面临一个核心难题:传统词向量把一个词的所有含义压缩成单一向量,导致"语义混淆"——比如"mouse"的"鼠标"和"老鼠"两个义项被混为一谈,连带使语义上毫不相关的"rat"和"screen"在向量空间中被错误地拉近。为此,研究者们提出了从词级别向更细粒度的词义级别(word sense)过渡的思路,并发展出两大类方法:一类是无监督方法,直接从文本语料中自动归纳词义;另一类是基于知识的方法,借助WordNet等词汇知识库中预定义的义项来构建词义表示。这篇综述系统梳理了这两类方法的技术路线、评测基准与下游应用,并从可解释性、义项粒度、跨领域适应性和组合性四个维度对它们进行了深入比较分析。这项工作为自然语言处理中更精准的语义建模提供了全面的理论参照和方法导图,对机器翻译、问答系统等众多NLP任务具有直接的指导价值。
原文 arXiv:1805.04032;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1805.04032v3