Shaking the foundations: delusions in sequence models for interaction and control
Pedro A. Ortega Deepmind Safety Analysis Markus Kunesch Deepmind Safety Analysis Grégoire Delétang Deepmind Safety Analysis Tim Genewein Deepmind Safety Analysis Jordi Grau-Moya Deepmind Safety Analysis Joel Veness DeepMind Jonas Buchli DeepMind Jonas Degrave DeepMind Bilal Piot DeepMind Julien Perolat DeepMind Tom Everitt DeepMind Corentin Tallec DeepMind Emilio Parisotto DeepMind Tom Erez DeepMind Yutian Chen DeepMind Scott Reed DeepMind Marcus Hutter DeepMind Nando de Freitas DeepMind Shane Legg DeepMind
Abstract
The recent phenomenal success of language models has reinvigorated machine learning research, and large sequence models such as transformers are being applied to a variety of domains. One important problem class that has remained relatively elusive however is purposeful adaptive behavior. Currently there is a common perception that sequence models “lack the understanding of the cause and effect of their actions” leading them to draw incorrect inferences due to auto-suggestive delusions. In this report we explain where this mismatch originates, and show that it can be resolved by treating actions as causal interventions. Finally, we show that in supervised learning, one can teach a system to condition or intervene on data by training with factual and counterfactual error signals respectively.
原文 arXiv:2110.10819;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2110.10819v1