Learning Stabilizable Nonlinear Dynamics with Contraction-Based Regularization
Sumeet Singh Department of Aeronautics and Astronautics, Stanford University ††thanks: Spencer M. Richards Department of Aeronautics and Astronautics, Stanford University ††thanks: Vikas Sindhwani Google Brain Robotics, New York ††thanks: Jean-Jacques E. Slotine Department of Mechanical Engineering, Massachusetts Institute of Technology ††thanks: Marco Pavone Department of Aeronautics and Astronautics, Stanford University ††thanks:
Abstract
We propose a novel framework for learning stabilizable nonlinear dynamical systems for continuous control tasks in robotics. The key contribution is a control-theoretic regularizer for dynamics fitting rooted in the notion of stabilizability, a constraint which guarantees the existence of robust tracking controllers for arbitrary open-loop trajectories generated with the learned system. Leveraging tools from contraction theory and statistical learning in Reproducing Kernel Hilbert Spaces, we formulate stabilizable dynamics learning as a functional optimization with convex objective and bi-convex functional constraints. Under a mild structural assumption and relaxation of the functional constraints to sampling-based constraints, we derive the optimal solution with a modified Representer theorem. Finally, we utilize random matrix feature approximations to reduce the dimensionality of the search parameters and formulate an iterative convex optimization algorithm that jointly fits the dynamics functions and searches for a certificate of stabilizability. We validate the proposed algorithm in simulation for a planar quadrotor, and on a quadrotor hardware testbed emulating planar dynamics
中文速览
机器人控制中,从少量数据学到的动力学模型往往质量很差,导致规划出的轨迹根本无法被稳定跟踪。这篇论文提出了一种把"可镇定性(stabilizability)"作为正则化约束嵌入动力学学习过程的框架——简单说,就是在拟合动力学模型时,同步要求存在一个能指数级收敛地跟踪任意参考轨迹的反馈控制器,从而把控制理论的需求直接写进学习目标里。具体实现上,作者借助收缩理论(contraction theory)中的控制收缩度量(Control Contraction Metric,CCM)将可镇定性条件转化为凸的线性矩阵不等式,再结合再生核希尔伯特空间(RKHS)的表示定理和随机矩阵特征近似,推导出一套可交替迭代求解的有限维凸优化算法。在仿真和真实四旋翼飞行器实验中,该方法用仅150条带噪声样本就能稳定跟踪远超训练分布的高难度轨迹,而传统回归方法学到的模型则频繁失控坠机。这项工作揭示了"控制论约束即正则化"的核心思想——把下游任务的可控性要求直接嵌入模型学习,既大幅提升数据效率,又从根本上保障了模型用于规划和控制时的可靠性。
原文 arXiv:1907.13122;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1907.13122v1