Agile Autonomous Driving using End-to-End Deep Imitation Learning
Yunpeng Pan1, Ching-An Cheng1, Kamil Saigol1, Keuntaek Lee2, Xinyan Yan1, Evangelos A. Theodorou1, and Byron Boots1 1Institute for Robotics and Intelligent Machines, 2School of Electrical and Computer Engineering Georgia Institute of Technology, Atlanta, Georgia 30332–0250
Abstract
We present an end-to-end imitation learning system for agile, off-road autonomous driving using only low-cost on-board sensors. By imitating a model predictive controller equipped with advanced sensors, we train a deep neural network control policy to map raw, high-dimensional observations to continuous steering and throttle commands. Compared with recent approaches to similar tasks, our method requires neither state estimation nor on-the-fly planning to navigate the vehicle. Our approach relies on, and experimentally validates, recent imitation learning theory. Empirically, we show that policies trained with online imitation learning overcome well-known challenges related to covariate shift and generalize better than policies trained with batch imitation learning. Built on these insights, our autonomous driving system demonstrates successful high-speed off-road driving, matching the state-of-the-art performance.
中文速览
高速越野自动驾驶长期依赖昂贵的GPS、IMU以及大量在线实时规划,成本高且难以普及。这篇文章提出了一套端到端的模仿学习(imitation learning)系统:让廉价的小比例赛车只靠单目摄像头和轮速传感器,通过模仿一个配备高级传感器的模型预测控制器(MPC)专家,直接学出从原始图像到油门和方向盘指令的深度神经网络策略,完全不需要状态估计或在线规划。研究在理论上将DAgger算法推广到连续动作空间,并在实车实验中验证了在线模仿学习(online imitation learning)能有效缓解协变量偏移(covariate shift)问题,表现明显优于离线批量模仿学习(batch imitation learning)。最终,该系统在真实越野赛道上实现了平均6 m/s、最高8 m/s的高速自驾,换算到实车相当于时速144 km/h,与依赖昂贵硬件的现有最优方法持平,证明了低成本传感器配合在线模仿学习完全可以胜任高难度的高速越野驾驶任务。
原文 arXiv:1709.07174;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1709.07174v6