Unsolved Problems in ML Safety
Dan Hendrycks UC Berkeley Nicholas Carlini Google John Schulman OpenAI Jacob Steinhardt UC Berkeley
Abstract
Machine learning (ML) systems are rapidly increasing in size, are acquiring new capabilities, and are increasingly deployed in high-stakes settings. As with other powerful technologies, safety for ML should be a leading research priority. In response to emerging safety challenges in ML, such as those introduced by recent large-scale models, we provide a new roadmap for ML Safety and refine the technical problems that the field needs to address. We present four problems ready for research, namely withstanding hazards (“Robustness”), identifying hazards (“Monitoring”), steering ML systems (“Alignment”), and reducing deployment hazards (“Systemic Safety”). Throughout, we clarify each problem’s motivation and provide concrete research directions.
中文速览
机器学习系统正在被大规模部署到医疗、自动驾驶、军事指挥等高风险场景,而现有系统存在脆弱性、不透明性和难以对齐人类意图等深层隐患,这些问题随着模型能力的增强只会愈发严峻。研究者提出了一套机器学习安全(ML Safety)技术路线图,将核心挑战归纳为四大方向:让系统能抵御极端事件与对抗攻击的"鲁棒性"(Robustness)、能识别异常与恶意行为的"监控"(Monitoring)、能与人类意图保持一致的"对齐"(Alignment),以及从系统层面降低整体部署风险的"系统性安全"(Systemic Safety)。论文为每个方向梳理了清晰的研究动机,并给出了可在近期落地推进的具体研究议题。这项工作的重要性在于,安全工程必须在系统设计早期介入——就像航空法规"写在血液里"的教训所揭示的那样——若等到灾难发生后再亡羊补牢,代价将无法承受,而一个成熟的安全研究社区能从根本上降低ML技术大规模落地的风险。
原文 arXiv:2109.13916;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2109.13916v5