Aha.
正在载入中英对照阅读…

arXiv:2304.07297 · 中英对照阅读

Language Instructed Reinforcement Learning for Human-AI Coordination

Hengyuan Hu、Dorsa Sadigh

中文速览

多智能体强化学习在缺少高质量人类行为数据时,常会学到虽然得分高、却不符合人类习惯的协作方式,导致人与AI难以配合。instructRL让人用自然语言说明希望AI遵循的策略,再借助大语言模型根据指令生成行动偏好,并把它作为强化学习的参考,推动训练结果收敛到符合人类期待的均衡。实验表明,它在设计的Say-Select游戏和Hanabi基准上都能学出遵守指令、表现强且风格不同的策略,而且人类知道指令后,与AI的协作效果显著提升。这个方法不依赖大规模人类示范数据,却能用简单的语言约定减少协作中的“猜心”和试错,为训练更容易理解、配合的人机协作智能体提供了实用思路。

摘要

One of the fundamental quests of AI is to produce agents that coordinate well with humans. This problem is challenging, especially in domains that lack high quality human behavioral data, because multi-agent reinforcement learning (RL) often converges to different equilibria from the ones that humans prefer. We propose a novel framework, instructRL , that enables humans to specify what kind of strategies they expect from their AI partners through natural language instructions. We use pretrained large language models to generate a prior policy conditioned on the human instruction and use the prior to regularize the RL objective. This leads to the RL agent converging to equilibria that are aligned with human preferences. We show that instructRL converges to human-like policies that satisfy the given instructions in a proof-of-concept environment as well as the challenging Hanabi benchmark. Finally, we show that knowing the language instruction significantly boosts human-AI coordination performance through human evaluations in Hanabi.

术语表

human-AI coordination
人-AI 协作
multi-agent reinforcement learning (MARL)
多智能体强化学习(MARL)
reinforcement learning (RL)
强化学习(RL)
instructRL
instructRL
natural language instructions
自然语言指令
large language models (LLMs)
大语言模型(LLMs)
prior policy
先验策略
RL objective regularization
强化学习目标正则化
equilibrium policies
均衡策略
joint policies
联合策略
equilibrium selection
均衡选择
human priors
人类先验
human-like policies
类人策略
human-compatible policies
人类兼容策略
Say-Select
Say-Select
best response strategy
最优响应策略
Q-learning
Q 学习
PPO
近端策略优化(PPO)
Hanabi benchmark
Hanabi 基准测试集
inductive bias
归纳偏置
language abstraction
语言抽象
intrinsic rewards
内在奖励
foundation models
基础模型
in-context learning
上下文学习
reward specification
奖励指定
CLIP
CLIP
contrastive video-text model
对比式视频-文本模型
human behavioral data
人类行为数据