Privacy and Statistical Risk: Formalisms and Minimax Bounds
Rina Foygel Barber Department of Statistics University of Chicago John C. Duchi Departments of Statistics and Electrical Engineering Stanford University
Abstract
We explore and compare a variety of definitions for privacy and disclosure limitation in statistical estimation and data analysis, including (approximate) differential privacy, testing-based definitions of privacy, and posterior guarantees on disclosure risk. We give equivalence results between the definitions, shedding light on the relationships between different formalisms for privacy. We also take an inferential perspective, where—building off of these definitions—we provide minimax risk bounds for several estimation problems, including mean estimation, estimation of the support of a distribution, and nonparametric density estimation. These bounds highlight the statistical consequences of different definitions of privacy and provide a second lens for evaluating the advantages and disadvantages of different techniques for disclosure limitation.
中文速览
在统计数据发布中,如何在保护个人隐私的同时仍能准确估计总体参数,是一个核心难题。这篇论文系统梳理并比较了多种隐私定义——包括差分隐私(differential privacy)及其近似版本、基于假设检验的隐私定义以及后验披露风险保证——并给出了它们之间的等价关系,厘清了不同隐私框架的内在联系。在此基础上,作者从统计推断角度出发,推导了均值估计、分布支撑估计和非参数密度估计等问题在各类隐私约束下的极小化极大(minimax)风险下界,并构造了达到这些下界的具体估计方案。结果表明,不同隐私定义下的最优估计误差在矩条件的依赖方式上高度相似,但在数据维度的依赖上存在显著差异——某些较弱的隐私定义能在高维场景中获得更好的统计性能,代价是安全保障有所降低。这项工作为量化隐私保护强度与统计估计精度之间的权衡提供了严格的理论依据,有助于实践者在设计隐私保护统计程序时做出更有据可依的选择。
原文 arXiv:1412.4451;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1412.4451v1