Unnatural Language Processing: Bridging the Gap Between Synthetic and Natural Language Data
Alana Marzoev Affiliation: Massachusetts Institute of Technology Correspondence to: Samuel Madden Affiliation: Massachusetts Institute of Technology M. Frans Kaashoek Affiliation: Massachusetts Institute of Technology Michael Cafarella Affiliation: University of Michigan Jacob Andreas Affiliation: Massachusetts Institute of Technology
Abstract
Large, human-annotated datasets are central to the development of natural language processing models. Collecting these datasets can be the most challenging part of the development process. We address this problem by introducing a general-purpose technique for “simulation-to-real" transfer in language understanding problems with a delimited set of target behaviors, making it possible to develop models that can interpret natural utterances without natural training data.
原文 arXiv:2004.13645;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2004.13645v1