RealTime QA: What’s the Answer Right Now?
Jungo Kasai♡♡{}^{\heartsuit}start_FLOATSUPERSCRIPT ♡ end_FLOATSUPERSCRIPT Keisuke Sakaguchi♣□normal-♣normal-□{}^{\clubsuit\square}start_FLOATSUPERSCRIPT ♣ □ end_FLOATSUPERSCRIPT Yoichi Takahashi♣normal-♣{}^{\clubsuit}start_FLOATSUPERSCRIPT ♣ end_FLOATSUPERSCRIPT Ronan Le Bras♢normal-♢{}^{\diamondsuit}start_FLOATSUPERSCRIPT ♢ end_FLOATSUPERSCRIPT Akari Asai♠normal-♠{}^{\spadesuit}start_FLOATSUPERSCRIPT ♠ end_FLOATSUPERSCRIPT Xinyan Velocity Yu\varheartsuit\varheartsuit{}^{\varheartsuit}start_FLOATSUPERSCRIPT end_FLOATSUPERSCRIPT Dragomir Radev\vardiamondsuit\vardiamondsuit{}^{\vardiamondsuit}start_FLOATSUPERSCRIPT end_FLOATSUPERSCRIPT Noah A. Smith♢♠normal-♢normal-♠{}^{\diamondsuit\spadesuit}start_FLOATSUPERSCRIPT ♢ ♠ end_FLOATSUPERSCRIPT Yejin Choi♢♠normal-♢normal-♠{}^{\diamondsuit\spadesuit}start_FLOATSUPERSCRIPT ♢ ♠ end_FLOATSUPERSCRIPT Kentaro Inui△♣□normal-△normal-♣normal-□{}^{\triangle\clubsuit\square}start_FLOATSUPERSCRIPT △ ♣ □ end_FLOATSUPERSCRIPT ♡♡{}^{\heartsuit}start_FLOATSUPERSCRIPT ♡ end_FLOATSUPERSCRIPTToyota Technological Institute at Chicago ♣♣{}^{\clubsuit}start_FLOATSUPERSCRIPT ♣ end_FLOATSUPERSCRIPTTohoku University □□{}^{\square}start_FLOATSUPERSCRIPT □ end_FLOATSUPERSCRIPTRIKEN ♢♢{}^{\diamondsuit}start_FLOATSUPERSCRIPT ♢ end_FLOATSUPERSCRIPTAllen Institute for AI ♠♠{}^{\spadesuit}start_FLOATSUPERSCRIPT ♠ end_FLOATSUPERSCRIPTUniversity of Washington \varheartsuit\varheartsuit{}^{\varheartsuit}start_FLOATSUPERSCRIPT end_FLOATSUPERSCRIPTUniversity of Southern California \vardiamondsuit\vardiamondsuit{}^{\vardiamondsuit}start_FLOATSUPERSCRIPT end_FLOATSUPERSCRIPTYale University △△{}^{\triangle}start_FLOATSUPERSCRIPT △ end_FLOATSUPERSCRIPTMBZUAI @realtimeqa Work was done during JK’s internship at AI2.
Abstract
We introduce RealTime QA, a dynamic question answering (QA) platform that announces questions and evaluates systems on a regular basis (weekly in this version). RealTime QA inquires about the current world, and QA systems need to answer questions about novel events or information. It therefore challenges static, conventional assumptions in open-domain QA datasets and pursues instantaneous applications. We build strong baseline models upon large pretrained language models, including GPT-3 and T5. Our benchmark is an ongoing effort, and this paper presents real-time evaluation results over the past year. Our experimental results show that GPT-3 can often properly update its generation results, based on newly-retrieved documents, highlighting the importance of up-to-date information retrieval. Nonetheless, we find that GPT-3 tends to return outdated answers when retrieved documents do not provide sufficient information to find an answer. This suggests an important avenue for future research: can an open-domain QA system identify such unanswerable cases and communicate with the user or even the retrieval module to modify the retrieval results? We hope that RealTime QA will spur progres
原文 arXiv:2207.13332;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2207.13332v2