Massively Multilingual Word Embeddings
Waleed Ammar George Mulcaire Yulia Tsvetkov Affiliation: School of Computer Science, Carnegie Mellon University, Pittsburgh, PA, USA Affiliation: School of Computer Science, Carnegie Mellon University, Pittsburgh, PA, USA Affiliation: Computer Science、Engineering, University of Washington, Seattle, WA, Guillaume Lample Chris Dyer Noah A. Smith Affiliation: School of Computer Science, Carnegie Mellon University, Pittsburgh, PA, USA Affiliation: School of Computer Science, Carnegie Mellon University, Pittsburgh, PA, USA Affiliation: Computer Science、Engineering, University of Washington, Seattle, WA,
Abstract
We introduce new methods for estimating and evaluating embeddings of words in more than fifty languages in a single shared embedding space. Our estimation methods, multiCluster and multiCCA, use dictionaries and monolingual data; they do not require parallel data. Our new evaluation method, multiqvec-cca, is shown to correlate better than previous ones with two downstream tasks (text categorization and parsing). We also describe a web portal for evaluation that will facilitate further research in this area, along with open-source releases of all our methods.
原文 arXiv:1602.01925;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1602.01925v2