PLATO-K: Internal and External Knowledge Enhanced Dialogue Generation
Siqi Bao Huang He111The DuSinc and Diamante are two companion papers of PLATO-K. Please refer to the original papers for more details. Jun Xu111The DuSinc and Diamante are two companion papers of PLATO-K. Please refer to the original papers for more details. Hua Lu111The DuSinc and Diamante are two companion papers of PLATO-K. Please refer to the original papers for more details. Fan Wang Hua Wu222Compared to PLATO-XL, PLATO-K also makes some modifications on the model design and training process, including the model scale (11B -> 22B parameters), computation amount (150B -> 200B tokens), position encoding (naive relative -> rotary position encoding), etc. According to previous studies (Kaplan et al., 2020; Hoffmann et al., 2022; Su et al., 2021) and our preliminary experiments, these modifications also bring minor improvements. Han Zhou Wenquan Wu Zheng-Yu Niu Haifeng Wang Baidu Inc., China {baosiqi, hehuang, xujun03, luhua05, wang.fan, Equal contribution.Corresponding authors.
Abstract
Recently, the practical deployment of open-domain dialogue systems has been plagued by the knowledge issue of information deficiency and factual inaccuracy. To this end, we introduce PLATO-K based on two-stage dialogic learning to strengthen internal knowledge memorization and external knowledge exploitation. In the first stage, PLATO-K learns through massive dialogue corpora and memorizes essential knowledge into model parameters. In the second stage, PLATO-K mimics human beings to search for external information and to leverage the knowledge in response generation. Extensive experiments reveal that the knowledge issue is alleviated significantly in PLATO-K with such comprehensive internal and external knowledge enhancement. Compared to the existing state-of-the-art Chinese dialogue model, the overall engagingness of PLATO-K is improved remarkably by 36.2% and 49.2% on chit-chat and knowledge-intensive conversations.
中文速览
开放域对话系统(open-domain dialogue system)在实际落地时普遍存在两大知识痛点:回复内容空洞缺乏信息量,以及容易产生事实性错误。为此,研究团队提出了PLATO-K,通过两阶段对话式学习(two-stage dialogic learning)来双管齐下地强化知识能力:第一阶段在海量社交媒体和网页转化的对话语料上预训练,再用高质量标注对话微调,将知识"记进"220亿参数的模型中;第二阶段则让模型学会像人一样主动搜索外部信息,并将检索结果融入回复生成。大量人工评测表明,与当时最优的中文对话模型PLATO-XL相比,PLATO-K在闲聊和知识密集型对话上的整体吸引力分别提升了36.2%和49.2%,知识性和事实准确度更是实现了数量级的跨越。这项工作证明了内部知识记忆与外部知识检索可以形成有效互补,为构建更可信、更具知识深度的实用对话系统提供了可行路径。
原文 arXiv:2211.00910;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2211.00910v1