The Second Conversational Intelligence Challenge (ConvAI2)
Emily Dinan Facebook AI Research Varvara Logacheva Moscow Institute of Physics and Technology Valentin Malykh Moscow Institute of Physics and Technology Alexander Miller Facebook AI Research Kurt Shuster Facebook AI Research Jack Urbanek Facebook AI Research Douwe Kiela Facebook AI Research Arthur Szlam Facebook AI Research Iulian Serban University of Montreal Ryan Lowe McGill University Facebook AI Research Shrimai Prabhumoye Carnegie Mellon University Alan W Black Carnegie Mellon University Alexander Rudnicky Carnegie Mellon University Jason Williams Microsoft Research Joelle Pineau Facebook AI Research McGill University Mikhail Burtsev Moscow Institute of Physics and Technology Jason Weston Facebook AI Research
Abstract
We describe the setting and results of the ConvAI2 NeurIPS competition that aims to further the state-of-the-art in open-domain chatbots. Some key takeaways from the competition are: (i) pretrained Transformer variants are currently the best performing models on this task, (ii) but to improve performance on multi-turn conversations with humans, future systems must go beyond single word metrics like perplexity to measure the performance across sequences of utterances (conversations) – in terms of repetition, consistency and balance of dialogue acts (e.g. how many questions asked vs. answered).
中文速览
开放域闲聊机器人(open-domain chatbot)长期缺乏统一的测评标准,ConvAI2竞赛正是为此而设计的——它基于Persona-Chat数据集,要求参赛模型扮演一个给定的"人设",与真实用户自然聊天,互相了解对方兴趣。竞赛同时采用了自动评估指标(困惑度、F1、候选排名准确率)和人工评估(Mechanical Turk及志愿者真实对话)三套标准来衡量模型表现。结果显示,基于预训练Transformer的模型(以Hugging Face团队为代表)在自动指标上遥遥领先,但人工评估的冠军却被Lost in Conversation夺走,说明自动指标与人类真实体验之间存在明显落差。这一差距的根源在于:现有自动评估只看单句质量,却忽略了多轮对话中的重复性、前后一致性以及提问与回答的平衡等关键因素——这正是未来闲聊系统和评估方法需要重点攻克的方向。
原文 arXiv:1902.00098;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1902.00098v1