Answering Science Exam Questions Using Query Reformulation with Background Knowledge
Ryan Musa Affiliation: IBM Research\addrIBM T.J. Watson Research Center Yorktown Heights, NY 10598 USA\email{ramusa, achille, kapanipa, krtalamad, Xiaoyan Wang Affiliation: University of Illinois at Urbana-Champaign\addrDepartment of Computer Science Urbana, IL 61801 Achille Fokoue Affiliation: IBM Research\addrIBM T.J. Watson Research Center Yorktown Heights, NY 10598 USA\email{ramusa, achille, kapanipa, krtalamad, Nicholas Mattei Affiliation: Tulane University\addrDepartment of Computer Science New Orleans, LA 70115 Maria Chang Affiliation: IBM Research\addrIBM T.J. Watson Research Center Yorktown Heights, NY 10598 USA\email{ramusa, achille, kapanipa, krtalamad, Pavan Kapanipathi Affiliation: IBM Research\addrIBM T.J. Watson Research Center Yorktown Heights, NY 10598 USA\email{ramusa, achille, kapanipa, krtalamad, Bassem Makni Affiliation: IBM Research\addrIBM T.J. Watson Research Center Yorktown Heights, NY 10598 USA\email{ramusa, achille, kapanipa, krtalamad, Kartik Talamadupula Affiliation: IBM Research\addrIBM T.J. Watson Research Center Yorktown Heights, NY 10598 USA\email{ramusa, achille, kapanipa, krtalamad, Michael Witbrock Affiliation: IBM Research\addrIBM T.J. Watson Research Center Yorktown Heights, NY 10598 USA\email{ramusa, achille, kapanipa, krtalamad,
Abstract
Open-domain question answering (QA) is an important problem in AI and NLP that is emerging as a bellwether for progress on the generalizability of AI methods and techniques. Much of the progress in open-domain QA systems has been realized through advances in information retrieval (IR) methods and corpus construction. In this paper, we focus on the recently introduced ARC Challenge dataset, which contains 2,590 multiple choice questions authored for grade-school science exams. These questions are selected to be the most challenging for current QA systems, and current state of the art performance is only slightly better than random chance. We present a system that reformulates a given question into queries that are used to retrieve supporting text from a large corpus of science-related text. Our rewriter is able to incorporate background knowledge from ConceptNet. In tandem with a generic textual entailment system trained on SciTail that identifies support in the retrieved results, our system outperforms several strong baselines on the end-to-end QA task despite only being trained to identify essential terms in the original source question. We use a generalizable decision methodology
原文 arXiv:1809.05726;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1809.05726v2