Acute-eval: Improved dialogue evaluation with optimized questions and multi-turn comparisons
Margaret Li Affiliation: Facebook AI Research Email: Jason Weston Affiliation: Facebook AI Research Email: Stephen Roller Affiliation: Facebook AI Research Email:
Abstract
While dialogue remains an important end-goal of natural language research, the difficulty of evaluation is an oft-quoted reason why it remains troublesome to make real progress towards its solution. Evaluation difficulties are actually two-fold: not only do automatic metrics not correlate well with human judgments, but also human judgments themselves are in fact difficult to measure. The two most used human judgment tests, single-turn pairwise evaluation and multi-turn Likert scores, both have serious flaws as we discuss in this work.
原文 arXiv:1909.03087;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1909.03087v1