Privacy Amplification by Subsampling: Tight Analyses via Couplings and Divergences
Borja Balle111Corresponding e-mail: Amazon Research, Cambridge, UK Gilles Barthe IMDEA Software Institute, Madrid, Spain Marco Gaboardi University at Buffalo (SUNY), Buffalo, USA
Abstract
Differential privacy comes equipped with multiple analytical tools for the design of private data analyses. One important tool is the so-called “privacy amplification by subsampling” principle, which ensures that a differentially private mechanism run on a random subsample of a population provides higher privacy guarantees than when run on the entire population. Several instances of this principle have been studied for different random subsampling methods, each with an ad-hoc analysis. In this paper we present a general method that recovers and improves prior analyses, yields lower bounds and derives new instances of privacy amplification by subsampling. Our method leverages a characterization of differential privacy as a divergence which emerged in the program verification community. Furthermore, it introduces new tools, including advanced joint convexity and privacy profiles, which might be of independent interest.
中文速览
随机子采样能让差分隐私(differential privacy)机制的隐私保证得到"放大",但此前各种子采样方式的分析都是各自为战、证明繁琐,且难以判断结果是否最优。本文提出了一套统一的分析框架,核心是借助来自程序验证领域的α-散度(α-divergence)来刻画差分隐私,并引入了"高级联合凸性"(advanced joint convexity)和"隐私轮廓"(privacy profiles)两个新工具,将子采样后的输出混合分布的隐私分析转化为对散度的精确估计。基于这套方法,论文不仅统一复现并改进了文献中已有的所有子采样放大结论,还推导出若干新实例,并通过通用下界证明了多数结果的紧致性。这项工作填补了隐私放大理论中系统性工具缺失的空白,对机器学习中随机优化、贝叶斯推断等大量依赖采样操作的私有算法设计具有直接指导意义。
原文 arXiv:1807.01647;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1807.01647v2