Can machines learn morality? The experiment
Liwei Jiang♣♡♣♡\clubsuit\heartsuit Jena D. Hwang♡♡\heartsuit Chandra Bhagavatula♡♡\heartsuit Ronan Le Bras♡♡\heartsuit Jenny Liang♡♡\heartsuit Jesse Dodge♡♡\heartsuit Keisuke Sakaguchi♡♡\heartsuit Maxwell Forbes♣♣\clubsuit Jon Borchardt♡♡\heartsuit Saadia Gabriel♣♣\clubsuit Yulia Tsvetkov♣♣\clubsuit Oren Etzioni♡♡\heartsuit Maarten Sap♡♡\heartsuit Regina Rini††\dagger Yejin Choi♣♡♣♡\clubsuit\heartsuit ♣♣\clubsuitPaul G. Allen School of Computer Science、Engineering, University of Washington ♡♡\heartsuitAllen Institute for Artificial Intelligence ††\daggerPhilosophy Department, York University
Abstract
As AI systems become increasingly powerful and pervasive, there are growing concerns about machines’ morality or a lack thereof. Yet, teaching morality to machines is a formidable task, as morality remains among the most intensely debated questions in humanity, let alone for AI. Existing AI systems deployed to millions of users, however, are already making decisions loaded with moral implications, which poses a seemingly impossible challenge: teaching machines moral sense, while humanity continues to grapple with it.
中文速览
机器究竟能不能学会道德判断?这篇论文提出了一个名为 Delphi 的人工智能系统,专门用来对自然语言描述的日常情境(如"帮朋友散布假新闻")做出道德评估。研究团队构建了一个包含 170 万条众包道德判断的数据集"常识规范库"(Commonsense Norm Bank),以此作为"道德教科书"对深度神经网络进行训练,让模型从大量具体例子中自下而上地学习人类共有的道德直觉,而非依赖预设的抽象规则。实验结果显示,Delphi 在未见过的道德情境上判断准确率达 92.8%,远超 GPT-3 的 60.2%,且对微小的情境变化表现出较强的鲁棒性;但同时,Delphi 也暴露出对数据偏见的敏感性和前后不一致等局限。这项工作的重要性在于:当前 AI 系统已在简历筛选、自动驾驶等高风险场景中大量应用,却普遍缺乏对人类价值观的感知,Delphi 为探索如何让机器具备基本道德常识提供了一个具有实际参考价值的早期框架。
原文 arXiv:2110.07574;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2110.07574v2