On Passivity, Reinforcement Learning and Higher-Order Learning in Multi-Agent Finite Games
Bolin Gao and Lacra Pavel This work was supported by NSERC Grant (261764). B. Gao and L. Pavel are with the Department of Electrical and Computer Engineering, University of Toronto, Canada.
Abstract
In this paper, we propose a passivity-based methodology for analysis and design of reinforcement learning in multi-agent finite games. Starting from a known exponentially-discounted reinforcement learning scheme, we show that convergence to a Nash distribution can be shown in the class of games characterized by the monotonicity property of their (negative) payoff. We further exploit passivity to propose a class of higher-order schemes that preserve convergence properties, can improve the speed of convergence and can even converge in cases whereby their first-order counterpart fail to converge. We demonstrate these properties through numerical simulations for several representative games.
中文速览
多智能体博弈中的强化学习往往很难证明收敛性,尤其是在超出势博弈范畴的更广泛游戏类型(如剪刀石头布)中。本文把连续时间指数折扣强化学习(EXP-D-RL)重新表述为一个反馈互联系统,并借助控制理论中的"均衡无关无源性"(equilibrium-independent passivity,EIP)框架来分析其收敛行为:核心思路是把智能体的得分更新动态视为一个输出严格无源系统,把博弈本身视为一个单调算子,两者的闭环互联自然保证了系统稳定。基于此,作者首先证明了在满足单调性的更大一类博弈(涵盖势博弈、零和博弈及标准剪刀石头布等)中,EXP-D-RL均收敛到纳什分布(logit均衡);进一步,他们利用软最大映射的余强制性(cocoercivity)将收敛范围扩展到部分"弱不稳定"博弈。在此基础上,论文给出了一套设计高阶学习动态的系统方法:通过在反馈路径中串联一个无源的线性时不变正实系统来构造二阶方案,数值仿真表明新方案在保持收敛保证的同时速度更快,甚至能在一阶方案发散的场景下实现收敛。这项工作为强化学习与无源性/控制理论之间架起了一座桥梁,为在更广泛博弈类别中设计有收敛保证的多智能体学习算法提供了原则性工具。
原文 arXiv:1808.04464;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1808.04464v1