arXiv:1709.06136 · 中英对照阅读
Iterative Policy Learning in End-to-End Trainable Task-Oriented Neural Dialog Models
中文速览
任务型对话系统要在多轮交流中帮用户完成订餐、查信息等目标,但强化学习通常依赖可靠的用户模拟器,而设计这样的模拟器并不比设计对话系统容易。研究先用真实对话语料分别监督训练一个基础对话代理和用户模拟器,再让两者围绕随机用户目标进行对话,并用任务完成度产生的奖励反复联合优化双方策略;两者都采用可端到端训练的神经网络。实验显示,这种方法相比单纯监督学习和只训练对话代理的强化学习基线,在任务成功率和总奖励上都有明显提升。它的重要性在于不再依赖大量手工规则,就能同时提升代理和模拟器,为更灵活、能应对多样场景的任务型对话系统提供了数据驱动的训练路径。
摘要
In this paper, we present a deep reinforcement learning (RL) framework for iterative dialog policy optimization in end-to-end task-oriented dialog systems. Popular approaches in learning dialog policy with RL include letting a dialog agent to learn against a user simulator. Building a reliable user simulator, however, is not trivial, often as difficult as building a good dialog agent. We address this challenge by jointly optimizing the dialog agent and the user simulator with deep RL by simulating dialogs between the two agents. We first bootstrap a basic dialog agent and a basic user simulator by learning directly from dialog corpora with supervised training. We then improve them further by letting the two agents to conduct task-oriented dialogs and iteratively optimizing their policies with deep RL. Both the dialog agent and the user simulator are designed with neural network models that can be trained end-to-end. Our experiment results show that the proposed method leads to promising improvements on task success rate and total task reward comparing to supervised training and single-agent RL training baseline models.
术语表
- deep reinforcement learning (RL)
- 深度强化学习(RL)
- dialog policy optimization
- 对话策略优化
- end-to-end task-oriented dialog system
- 端到端任务型对话系统
- user simulator
- 用户模拟器
- dialog agent
- 对话代理
- iterative dialog policy learning
- 迭代式对话策略学习
- supervised training
- 监督训练
- dialog corpus
- 对话语料库
- task success rate
- 任务成功率
- total task reward
- 任务总奖励
- natural language understanding (NLU)
- 自然语言理解(NLU)
- dialog state tracking (DST)
- 对话状态跟踪(DST)
- dialog management (DM)
- 对话管理(DM)
- credit assignment
- 信用分配
- policy shaping
- 策略塑形
- goal fulfilling task
- 目标达成任务
- partially observable Markov decision process (POMDP)
- 部分可观测马尔可夫决策过程(POMDP)
- dialog state
- 对话状态
- system action
- 系统动作
- end-to-end memory network
- 端到端记忆网络
- belief tracking
- 信念状态跟踪
- belief state
- 信念状态
- knowledge base (KB)
- 知识库(KB)
- dialog simulation
- 对话模拟
- multi-agent reinforcement learning
- 多智能体强化学习
- stochastic game
- 随机博弈
- co-adaptation
- 协同适应
- data-driven learning
- 数据驱动学习
- policy robustness
- 策略鲁棒性