A Data-Driven Approach for Learning to Control Computers
Peter Humphreys David Raposo Toby Pohlen Gregory Thornton Rachita Chhaparia Alistair Muldal Josh Abramson Petko Georgiev Adam Santoro Timothy Lillicrap
Abstract
It would be useful for machines to use computers as humans do so that they can aid us in everyday tasks. This is a setting in which there is also the potential to leverage large-scale expert demonstrations and human judgements of interactive behaviour, which are two ingredients that have driven much recent success in AI. Here we investigate the setting of computer control using keyboard and mouse, with goals specified via natural language. Instead of focusing on hand-designed curricula and specialized action spaces, we focus on developing a scalable method centered on reinforcement learning combined with behavioural priors informed by actual human-computer interactions. We achieve state-of-the-art and human-level mean performance across all tasks within the MiniWob++ benchmark, a challenging suite of computer control problems, and find strong evidence of cross-task transfer. These results demonstrate the usefulness of a unified human-agent interface when training machines to use computers. Altogether our results suggest a formula for achieving competency beyond MiniWob++ and towards controlling computers, in general, as a human would.
中文速览
让机器像人类一样使用电脑完成日常任务,这是一个既实用又充满挑战的目标。研究团队聚焦于用键盘和鼠标操控电脑、以自然语言描述任务的设置,采用强化学习(Reinforcement Learning, RL)与行为克隆(Behavioural Cloning, BC)相结合的方法,并收集了比以往研究多达400倍的真实人类操作示范数据来训练智能体。在涵盖点击、表单填写等多种计算机操作的MiniWob++基准测试中,该方法首次达到了人类平均水平,并显著超越了此前所有方案,同时还观察到跨任务的知识迁移现象。这项工作表明,让机器与人类共用同一套键鼠交互界面、并辅以大规模人类示范数据,是训练通用计算机控制智能体的关键所在,为未来实现更广泛的"人类级"电脑操控能力提供了清晰的路线图。
原文 arXiv:2202.08137;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2202.08137v2