New word analogy corpus for exploring embeddings of Czech words
Lukáš Svoboda 1122 Tomáš Brychcín 1122
Abstract
The word embedding methods have been proven to be very useful in many tasks of NLP (Natural Language Processing). Much has been investigated about word embeddings of English words and phrases, but only little attention has been dedicated to other languages.
中文速览
词向量(word embedding)方法在英语自然语言处理中表现出色,但对于捷克语这样形态极为丰富的语言,其效果几乎无人研究。作者为此专门构建了一个包含两万余道类比题的捷克语评测语料库,覆盖语义(如首都-国家、反义词)和句法(如名词复数、职业阴阳性)等多个维度,并在捷克语维基百科语料上分别训练了CBOW、Skip-gram和GloVe三种词向量模型加以测试。结果显示,CBOW整体表现最好,GloVe最差,所有模型在句法类任务(动词时态、名词复数)上的准确率普遍高于语义类任务(国家总统等),且增加训练轮数和向量维度通常能带来明显提升。这项工作不仅揭示了现有词向量方法在高度屈折变化语言上的局限,所发布的评测语料库也为后续斯拉夫语系乃至其他形态复杂语言的词向量研究提供了重要基准。
原文 arXiv:1608.00789;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1608.00789v1