RLCard: A toolkit for Reinforcement Learning in Card Games
Daochen Zha1, Kwei-Herng Lai1, Yuanpu Cao1, Songyi Huang2, Ruzhe Wei1††, Junyu Guo1††, Xia Hu1 1 Department of Computer Science and Engineering, Texas A、M University, College Station, USA 2 Simon Fraser University, BC, Canada {daochen.zha, {guojunyu, Authors contribute during the visit at Texas A、M University.
Abstract
We present RLCard, an open-source toolkit for reinforcement learning research in card games. It supports various card environments with easy-to-use interfaces, including Blackjack, Leduc Hold’em, Texas Hold’em, UNO, Dou Dizhu and Mahjong. The goal of RLCard is to bridge reinforcement learning and imperfect information games, and push forward the research of reinforcement learning in domains with multiple agents, large state and action space, and sparse reward. In this paper, we provide an overview of the key components in RLCard, a discussion of the design principles, a brief introduction of the interfaces, and comprehensive evaluations of the environments. The codes and documents are available at https://github.com/datamllab/rlcard.
中文速览
扑克、麻将、斗地主这类卡牌游戏天然具备多玩家对抗、状态空间巨大、动作组合爆炸、奖励稀疏等特性,是检验强化学习算法的绝佳战场,但此前缺乏一套统一、易用的研究平台。RLCard 就是为填补这一空白而生的开源工具包,它将 Blackjack、Leduc Hold'em、Texas Hold'em、UNO、斗地主和麻将六款游戏封装成标准化的强化学习接口,既支持基于采样的算法,也支持需要回溯游戏树的方法,研究者无需深入了解游戏细节便能直接专注于算法开发。实验表明,DQN、NFSP、CFR 等主流算法在小规模游戏上已能稳定提升,但在 UNO、麻将、斗地主等大规模环境中表现高度不稳定,说明这些环境仍存在大量尚待突破的研究空间。RLCard 的意义在于:它为不完美信息博弈与强化学习研究搭建了一座标准化的桥梁,使后续算法的开发、对比和复现变得更加便捷。
原文 arXiv:1910.04376;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1910.04376v2