On Passivity, Reinforcement Learning and Higher-Order Learning in Multi-Agent Finite Games
Bolin Gao Lacra Pavel Thanks: This work was supported by NSERC Grant (261764). B. Gao and L. Pavel are with the Department of Electrical and Computer Engineering, University of Toronto, Canada.
Abstract
In this paper, we propose a passivity-based methodology for analysis and design of reinforcement learning in multi-agent finite games. Starting from a known exponentially-discounted reinforcement learning scheme, we show that convergence to a Nash distribution can be shown in the class of games characterized by the monotonicity property of their (negative) payoff. We further exploit passivity to propose a class of higher-order schemes that preserve convergence properties, can improve the speed of convergence and can even converge in cases whereby their first-order counterpart fail to converge. We demonstrate these properties through numerical simulations for several representative games.
原文 arXiv:1808.04464;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1808.04464v1