Differentially Private Fine-tuning of Language ModelsThanks: Aside from the first and second authors, all other authors are listed in alphabetical order.
Da Yu Thanks: Sun Yat-sen University. Work was done while an intern at Microsoft Research Asia. Saurabh Naik Thanks: Microsoft. {snaik, Arturs Backurs Thanks: Microsoft Research. {arturs.backurs, sigopi, huseyin.inan, jakul, amonteiroman, Sivakanth Gopi Huseyin A. Inan Gautam Kamath Thanks: Cheriton School of Computer Science, University of Waterloo. Supported by an NSERC Discovery Grant. Janardhan Kulkarni Yin Tat Lee Thanks: University of Washington and Microsoft Research. Andre Manoel Lukas Wutschitz Sergey Yekhanin Huishuai Zhang Thanks: Microsoft Research Asia.
Abstract
We give simpler, sparser, and faster algorithms for differentially private fine-tuning of large-scale pre-trained language models, which achieve the state-of-the-art privacy versus utility tradeoffs on many standard NLP tasks. We propose a meta-framework for this problem, inspired by the recent success of highly parameter-efficient methods for fine-tuning. Our experiments show that differentially private adaptations of these approaches outperform previous private algorithms in three important dimensions: utility, privacy, and the computational and memory cost of private training. On many commonly studied datasets, the utility of private models approaches that of non-private models. For example, on the MNLI dataset we achieve an accuracy of $87.8\%$ using RoBERTa-Large and $83.5\%$ using RoBERTa-Base with a privacy budget of $\varepsilon=6.7$ . In comparison, absent privacy constraints, RoBERTa-Large achieves an accuracy of $90.2\%$ . Our findings are similar for natural language generation tasks. Privately fine-tuning with DART, GPT-2-Small, GPT-2-Medium, GPT-2-Large, and GPT-2-XL achieve BLEU scores of 38.5, 42.0, 43.1, and 43.8 respectively (privacy budget of $\varepsilon=6.8,\de
原文 arXiv:2110.06500;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2110.06500v2