Aha.
正在载入中英对照阅读…

arXiv:1709.06136 · 中英对照阅读

Iterative Policy Learning in End-to-End Trainable Task-Oriented Neural Dialog Models

Bing Liu, Ian Lane

中文速览

任务型对话系统要在多轮交流中帮用户完成订餐、查信息等目标,但强化学习通常依赖可靠的用户模拟器,而设计这样的模拟器并不比设计对话系统容易。研究先用真实对话语料分别监督训练一个基础对话代理和用户模拟器,再让两者围绕随机用户目标进行对话,并用任务完成度产生的奖励反复联合优化双方策略;两者都采用可端到端训练的神经网络。实验显示,这种方法相比单纯监督学习和只训练对话代理的强化学习基线,在任务成功率和总奖励上都有明显提升。它的重要性在于不再依赖大量手工规则,就能同时提升代理和模拟器,为更灵活、能应对多样场景的任务型对话系统提供了数据驱动的训练路径。

摘要

In this paper, we present a deep reinforcement learning (RL) framework for iterative dialog policy optimization in end-to-end task-oriented dialog systems. Popular approaches in learning dialog policy with RL include letting a dialog agent to learn against a user simulator. Building a reliable user simulator, however, is not trivial, often as difficult as building a good dialog agent. We address this challenge by jointly optimizing the dialog agent and the user simulator with deep RL by simulating dialogs between the two agents. We first bootstrap a basic dialog agent and a basic user simulator by learning directly from dialog corpora with supervised training. We then improve them further by letting the two agents to conduct task-oriented dialogs and iteratively optimizing their policies with deep RL. Both the dialog agent and the user simulator are designed with neural network models that can be trained end-to-end. Our experiment results show that the proposed method leads to promising improvements on task success rate and total task reward comparing to supervised training and single-agent RL training baseline models.

术语表

deep reinforcement learning (RL)
深度强化学习(RL)
dialog policy optimization
对话策略优化
end-to-end task-oriented dialog system
端到端任务型对话系统
user simulator
用户模拟器
dialog agent
对话代理
iterative dialog policy learning
迭代式对话策略学习
supervised training
监督训练
dialog corpus
对话语料库
task success rate
任务成功率
total task reward
任务总奖励
natural language understanding (NLU)
自然语言理解(NLU)
dialog state tracking (DST)
对话状态跟踪(DST)
dialog management (DM)
对话管理(DM)
credit assignment
信用分配
policy shaping
策略塑形
goal fulfilling task
目标达成任务
partially observable Markov decision process (POMDP)
部分可观测马尔可夫决策过程(POMDP)
dialog state
对话状态
system action
系统动作
end-to-end memory network
端到端记忆网络
belief tracking
信念状态跟踪
belief state
信念状态
knowledge base (KB)
知识库(KB)
dialog simulation
对话模拟
multi-agent reinforcement learning
多智能体强化学习
stochastic game
随机博弈
co-adaptation
协同适应
data-driven learning
数据驱动学习
policy robustness
策略鲁棒性