BBQ-Networks: Efficient Exploration in Deep Reinforcement Learning for Task-Oriented Dialogue Systems
Zachary Lipton⋆ Xiujun Li† Jianfeng Gao† Lihong Li‡ Faisal Ahmed† Li Deng§∗ ⋆Carnegie Mellon University, Pittsburgh, PA, USA ⋆Amazon AI, Palo Alto, CA, USA †Microsoft Research, Redmond, WA, USA ‡Google Inc., Kirkland, WA, USA §Citadel, Seattle, WA, USA This work was done while ZL, LL、LD were with Microsoft.
Abstract
We present a new algorithm that significantly improves the efficiency of exploration for deep Q-learning agents in dialogue systems. Our agents explore via Thompson sampling, drawing Monte Carlo samples from a Bayes-by-Backprop neural network. Our algorithm learns much faster than common exploration strategies such as $\epsilon$ -greedy, Boltzmann, bootstrapping, and intrinsic-reward-based ones. Additionally, we show that spiking the replay buffer with experiences from just a few successful episodes can make Q-learning feasible when it might otherwise fail.
原文 arXiv:1608.05081;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1608.05081v4