Unsupervised Learning of Sentence Embeddings using Compositional n-Gram Features
Matteo Pagliardini* Affiliation: Iprova SA, Switzerland Email: Prakhar Gupta* Affiliation: EPFL, Switzerland Email: Martin Jaggi Affiliation: EPFL, Switzerland Email:
Abstract
The recent tremendous success of unsupervised word embeddings in a multitude of applications raises the obvious question if similar methods could be derived to improve embeddings (i.e. semantic representations) of word sequences as well. We present a simple but efficient unsupervised objective to train distributed representations of sentences. Our method outperforms the state-of-the-art unsupervised models on most benchmark tasks, highlighting the robustness of the produced general-purpose sentence embeddings.
原文 arXiv:1703.02507;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1703.02507v3