Pre-training Is (Almost) All You Need: An Application to Commonsense Reasoning
Alexandre Tamborrino Thanks: Equal contribution. Nicola Pellicanò Baptiste Pannier Affiliation: Pascal Voitot, Louise Naudin Affiliation: Samsung Strategy and Innovation Center Email:
Abstract
Fine-tuning of pre-trained transformer models has become the standard approach for solving common NLP tasks Devlin et al. 2019. Most of the existing approaches rely on a randomly initialized classifier on top of such networks. We argue that this fine-tuning procedure is sub-optimal as the pre-trained model has no prior on the specific classifier labels, while it might have already learned an intrinsic textual representation of the task. In this paper, we introduce a new scoring method that casts a plausibility ranking task in a full-text format and leverages the masked language modeling head tuned during the pre-training phase. We study commonsense reasoning tasks where the model must rank a set of hypotheses given a premise, focusing on the COPA Gordon et al. 2012, Swag Zellers et al. 2018, HellaSwag Zellers et al. 2019 and CommonsenseQA Talmor et al. 2019 datasets. By exploiting our scoring method without fine-tuning, we are able to produce strong baselines (e.g. 80% test accuracy on COPA) that are comparable to supervised approaches. Moreover, when fine-tuning directly on the proposed scoring function, we show that our method provides a much more stable training phase across ran
原文 arXiv:2004.14074;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2004.14074v1