Linguistic Analysis of Pretrained Sentence Encoders with Acceptability Judgments
Alex Affiliation: Dept. of LinguisticsNew York University10 Washington PlaceNew York, NY 10003 Samuel R. Affiliation: Dept. of LinguisticsNew York University10 Washington PlaceNew York, NY 10003 Affiliation: Dept. of Computer ScienceNew York University60 Fifth AvenueNew York, NY 10011 Affiliation: Center for Data ScienceNew York University60 Fifth AvenueNew York, NY 10011
Abstract
Recent work on evaluating grammatical knowledge in pretrained sentence encoders gives a fine-grained view of a small number of phenomena. We introduce a new analysis dataset that also has broad coverage of linguistic phenomena. We annotate the development set of the Corpus of Linguistic Acceptability (Warstadt et al. 2018, CoLA;) for the presence of 13 classes of syntactic phenomena including various forms of argument alternations, movement, and modification. We use this analysis set to investigate the grammatical knowledge of three pretrained encoders: BERT Devlin et al. 2018, GPT Radford et al. 2018, and the BiLSTM baseline from Warstadt et al. 2018 We find that these models have a strong command of complex or non-canonical argument structures like ditransitives (Sue gave Dan a book) and passives (The book was read). Sentences with long-distance dependencies like questions (What do you think I ate?) challenge all models, but for these, BERT and GPT have a distinct advantage over the baseline. We conclude that recent sentence encoders, despite showing near-human performance on acceptability classification overall, still fail to make fine-grained grammaticality distinctions for man
原文 arXiv:1901.03438;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1901.03438v4