DeepStack: Expert-Level Artificial Intelligence in Heads-Up No-Limit Poker
Matej Moravčík Affiliation: Department of Computing Science, University of Alberta,Edmonton, Alberta, T6G2E8, Canada Affiliation: Department of Applied Mathematics, Charles University,Prague, Czech Republic Affiliation: These authors contributed equally to this work and are listed in alphabetical order. Martin Schmid Affiliation: Department of Computing Science, University of Alberta,Edmonton, Alberta, T6G2E8, Canada Affiliation: Department of Applied Mathematics, Charles University,Prague, Czech Republic Affiliation: These authors contributed equally to this work and are listed in alphabetical order. Neil Burch Affiliation: Department of Computing Science, University of Alberta,Edmonton, Alberta, T6G2E8, Canada Viliam Lisý Affiliation: Department of Computing Science, University of Alberta,Edmonton, Alberta, T6G2E8, Canada Affiliation: Department of Computer Science, FEE, Czech Technical University,Prague, Czech Republic Dustin Morrill Affiliation: Department of Computing Science, University of Alberta,Edmonton, Alberta, T6G2E8, Canada Nolan Bard Affiliation: Department of Computing Science, University of Alberta,Edmonton, Alberta, T6G2E8, Canada Trevor Davis Affiliation: Department of Computing Science, University of Alberta,Edmonton, Alberta, T6G2E8, Canada Kevin Waugh Affiliation: Department of Computing Science, University of Alberta,Edmonton, Alberta, T6G2E8, Canada Michael Johanson Affiliation: Department of Computing Science, University of Alberta,Edmonton, Alberta, T6G2E8, Canada Michael Bowling Affiliation: Department of Computing Science, University of Alberta,Edmonton, Alberta, T6G2E8, Canada Affiliation: To whom correspondence should be addressed; E-mail:
Abstract
Although DeepStack uses ideas from abstraction, it is fundamentally different from abstraction-based approaches. DeepStack restricts the number of actions in its lookahead trees, much like action abstraction [25, 26]. However, each re-solve in DeepStack starts from the actual public state and so it always perfectly understands the current situation. The algorithm also never needs to use the opponent’s actual action to obtain correct ranges or opponent counterfactual values, thereby avoiding translation of opponent bets. We used hand clustering as inputs to our counterfactual value functions, much like explicit card abstraction approaches [27, 28]. However, our clustering is used to estimate counterfactual values at the end of a lookahead tree rather than limiting what information the player has about their cards when acting. We later show that these differences result in a strategy substantially more difficult to exploit.
原文 arXiv:1701.01724;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1701.01724v3