Simplified TinyBERT: Knowledge Distillation for Document Retrieval
Xuanang Chen(🖂) Affiliation: University of Chinese Academy of Sciences, Beijing, China Affiliation: Institute of Software, Chinese Academy of Sciences, Beijing, China E-mail {benhe, Ben He(🖂) Affiliation: University of Chinese Academy of Sciences, Beijing, China Affiliation: Institute of Software, Chinese Academy of Sciences, Beijing, China E-mail {benhe, Kai Hui Thanks: This work has been done before joining Amazon. Affiliation: Amazon Alexa, Berlin, Germany E-mail Le Sun Affiliation: Institute of Software, Chinese Academy of Sciences, Beijing, China E-mail {benhe, Yingfei Sun(🖂) Affiliation: University of Chinese Academy of Sciences, Beijing, China
Abstract
Despite the effectiveness of utilizing the BERT model for document ranking, the high computational cost of such approaches limits their uses. To this end, this paper first empirically investigates the effectiveness of two knowledge distillation models on the document ranking task. In addition, on top of the recently proposed TinyBERT model, two simplifications are proposed. Evaluations on two different and widely-used benchmarks demonstrate that Simplified TinyBERT with the proposed simplifications not only boosts TinyBERT, but also significantly outperforms BERT-Base when providing 15 $\times$ speedup.
原文 arXiv:2009.07531;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2009.07531v2