A General Safety Framework for Learning-Based Control in Uncertain Robotic Systems
Jaime F. Fisac∗*1 Anayo K. Akametalu∗*1 Melanie N. Zeilinger2 Shahab Kaynama3 Jeremy Gillula4 and Claire J. Tomlin1 ∗* The first two authors contributed equally. This work was supported by the NSF CPS project ActionWebs under grant 0931843, NSF CPS project FORCES under grant 1239166, and by ONR under the HUNT, SMARTS and Embedded Humans MURIs, and by AFOSR under the CHASE MURI. The research of J. F. Fisac received funding from the “la Caixa” Foundation. The research of A.K. Akametalu received funding from the NSF Bridge to Doctorate program. The research of M.N. Zeilinger received funding from the EU FP7 (FP7/2007-2013) under grant PIOFGA-2011-301436- COGENT . 1 Department of Electrical Engineering and Computer Sciences, University of California, Berkeley. Cory Hall, Berkeley, CA 94720, United States. 2 Department of Mechanical and Process Engineering, ETH Zurich. Rämistrasse 101, 8092 Zürich, Switzerland. 3 Clearpath Robotics. 1425 Strasburg Rd, Suite 2A, Kitchener, ON N2R 1H2, Canada. 4 Electronic Frontier Foundation. 815 Eddy St, San Francisco, CA 94109, United States. Email: {jfisac, kakametalu, tomlin} @eecs.berkeley.edu,
Abstract
The proven efficacy of learning-based control schemes strongly motivates their application to robotic systems operating in the physical world. However, guaranteeing correct operation during the learning process is currently an unresolved issue, which is of vital importance in safety-critical systems. We propose a general safety framework based on Hamilton-Jacobi reachability methods that can work in conjunction with an arbitrary learning algorithm. The method exploits approximate knowledge of the system dynamics to guarantee constraint satisfaction while minimally interfering with the learning process. We further introduce a Bayesian mechanism that refines the safety analysis as the system acquires new evidence, reducing initial conservativeness when appropriate while strengthening guarantees through real-time validation. The result is a least-restrictive, safety-preserving control law that intervenes only when (a) the computed safety guarantees require it, or (b) confidence in the computed guarantees decays in light of new observations. We prove theoretical safety guarantees combining probabilistic and worst-case analysis and demonstrate the proposed framework experimentally on a
中文速览
让机器人在"边学边干"的过程中不出事故,是强化学习落地的核心难题。这篇文章提出了一套把汉密顿-雅可比可达性分析(Hamilton-Jacobi reachability)与贝叶斯高斯过程(Gaussian process)相结合的安全监督框架:系统利用已知的近似动力学模型预先划定一个"安全集",只在机器人快要冲出安全边界时才强制接管控制权,其余时间完全放手让强化学习算法自由探索;与此同时,框架会持续收集飞行数据,用贝叶斯推断实时校验模型假设是否仍然成立,一旦发现真实扰动超出预期就及时收紧保护,反之则放宽限制以减少对学习的干预。作者在四旋翼无人机(quadrotor)上进行了真实飞行实验,无人机在策略梯度强化学习的全程中从未坠机,并在飞行中遭受强烈外部扰动时也能安全规避。这项工作首次将在线模型验证机制引入可达性分析,为学习型机器人系统提供了一套既有最坏情况理论保证、又具备概率自适应能力的通用安全框架。
原文 arXiv:1705.01292;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1705.01292v3