Towards Ecologically Valid Research on Language User Interfaces
Harm de Vries1 Dzmitry Bahdanau1 Christopher Manning2,3 1Element AI 2Stanford University 3CIFAR Fellow
Abstract
Language User Interfaces (LUIs) could improve human-machine interaction for a wide variety of tasks, such as playing music, getting insights from databases, or instructing domestic robots. In contrast to traditional hand-crafted approaches, recent work attempts to build LUIs in a data-driven way using modern deep learning methods. To satisfy the data needs of such learning algorithms, researchers have constructed benchmarks that emphasize the quantity of collected data at the cost of its naturalness and relevance to real-world LUI use cases. As a consequence, research findings on such benchmarks might not be relevant for developing practical LUIs. The goal of this paper is to bootstrap the discussion around this issue, which we refer to as the benchmarks’ low ecological validity. To this end, we describe what we deem an ideal methodology for machine learning research on LUIs and categorize five common ways in which recent benchmarks deviate from it. We give concrete examples of the five kinds of deviations and their consequences. Lastly, we offer a number of recommendations as to how to increase the ecological validity of machine learning research on LUIs.
中文速览
当前主流的语言用户界面(Language User Interfaces, LUIs)研究普遍依赖人工合成或规模优先的数据集来训练和评测模型,却牺牲了数据与真实使用场景的贴近程度。本文将这一问题定义为"生态效度(ecological validity)不足",并系统梳理出五类常见偏差:合成语言、人造任务、测试者非目标用户、脚本化与启动效应、以及单轮交互设计,通过大量具体基准数据集的案例说明每类偏差如何扭曲研究结论。作者进一步提出了一套理想的研究范式——以真实用户群体和实际任务为起点,借助"绿野仙踪(Wizard-of-Oz)"模拟收集自然交互数据,再训练和评估模型——并给出提升生态效度的若干可操作建议。这项工作的意义在于:NLP领域在基准上的性能突破若与真实用户需求脱节,便难以转化为人们愿意日常使用的实用界面,而厘清这一鸿沟是推动LUI研究真正落地的必要前提。
原文 arXiv:2007.14435;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2007.14435v1