Third-Person Imitation Learning
Bradly C. Stadie Affiliation: OpenAI Affiliation: UC Berkeley, Department of Statistics Pieter Abbeel Affiliation: OpenAI Affiliation: UC Berkeley, Departments of EECS and ICSI{bstadie, pieter, Ilya Sutskever Affiliation: OpenAI
Abstract
Reinforcement learning (RL) makes it possible to train agents capable of achieving sophisticated goals in complex and uncertain environments. A key difficulty in reinforcement learning is specifying a reward function for the agent to optimize. Traditionally, imitation learning in RL has been used to overcome this problem. Unfortunately, hitherto imitation learning methods tend to require that demonstrations are supplied in the first-person: the agent is provided with a sequence of states and a specification of the actions that it should have taken. While powerful, this kind of imitation learning is limited by the relatively hard problem of collecting first-person demonstrations. Humans address this problem by learning from third-person demonstrations: they observe other humans perform tasks, infer the task, and accomplish the same task themselves.
原文 arXiv:1703.01703;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1703.01703v2