Aligning AI With Shared Human Values
Dan Hendrycks UC Berkeley、Collin Burns* Columbia University、Steven Basart UChicago \ANDAndrew Critch UC Berkeley、Jerry Li Microsoft、Dawn Song UC Berkeley、Jacob Steinhardt UC Berkeley\AND Equal Contribution.
Abstract
We show how to assess a language model’s knowledge of basic concepts of morality. We introduce the ETHICS dataset, a new benchmark that spans concepts in justice, well-being, duties, virtues, and commonsense morality. Models predict widespread moral judgments about diverse text scenarios. This requires connecting physical and social world knowledge to value judgements, a capability that may enable us to steer chatbot outputs or eventually regularize open-ended reinforcement learning agents. With the ETHICS dataset, we find that current language models have a promising but incomplete ability to predict basic human ethical judgements. Our work shows that progress can be made on machine ethics today, and it provides a steppingstone toward AI that is aligned with human values.
中文速览
让机器理解人类道德观念一直缺乏系统性的衡量标准,为此研究者构建了 ETHICS 这一大规模基准数据集,涵盖正义、义务、美德、功利主义和常识道德五大伦理维度,共超过13万条自然语言场景样本。数据集要求模型将现实世界的物理与社会知识和道德判断联系起来,例如判断"从别人钱包里拿钱"与"捡路边的一分硬币"在道德上的不同。实验结果表明,在大规模文本上预训练的语言模型经过微调后表现出初步但尚不完善的道德推理能力,说明当前技术已经能够在机器伦理上取得进展,但模型对道德相关世界知识的掌握仍有很大欠缺。这项工作的意义在于为衡量AI系统的价值观对齐程度提供了第一块垫脚石,未来可用于约束聊天机器人的输出或为开放世界强化学习智能体提供道德正则化信号。
原文 arXiv:2008.02275;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2008.02275v6