CoAuthor: Designing a Human-AI Collaborative Writing Dataset for Exploring Language Model Capabilities
Mina Lee Stanford UniversityUnited States , Percy Liang Stanford UniversityUnited States and Qian Yang Cornell UniversityUnited States
Abstract
Large language models (LMs) offer unprecedented language generation capabilities and exciting opportunities for interaction design. However, their highly context-dependent capabilities are difficult to grasp and are often subjectively interpreted. In this paper, we argue that by curating and analyzing large interaction datasets, the HCI community can foster more incisive examinations of LMs’ generative capabilities. Exemplifying this approach, we present CoAuthor, a dataset designed for revealing GPT-3’s capabilities in assisting creative and argumentative writing. CoAuthor captures rich interactions between 63 writers and four instances of GPT-3 across 1445 writing sessions. We demonstrate that CoAuthor can address questions about GPT-3’s language, ideation, and collaboration capabilities, and reveal its contribution as a writing “collaborator” under various definitions of good collaboration. Finally, we discuss how this work may facilitate a more principled discussion around LMs’ promises and pitfalls in relation to interaction design. The dataset and an interface for replaying the writing sessions are publicly available at https://coauthor.stanford.edu.
中文速览
大型语言模型(LM)在写作辅助等交互场景中潜力巨大,但其能力极依赖具体上下文,且难以客观评估,导致设计者很难判断它到底能做什么、不能做什么。为此,研究团队构建了一个名为 CoAuthor 的大规模交互数据集,记录了 63 位写作者与四个不同配置的 GPT-3 之间共 1445 次写作会话,涵盖创意写作和议论写作两类任务。通过系统分析这些数据,研究者得以量化评估 GPT-3 在语言流畅性、创意生成和协作写作三个维度上的表现,并在多种"好协作"定义下衡量其实际贡献。这项工作的价值在于:它为 HCI 社区提供了一套可复用、可回放的真实交互数据,让研究者和设计者能够跳出主观印象,更系统、更有据可查地讨论语言模型在交互设计中的机遇与局限。
原文 arXiv:2201.06796;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2201.06796v2