Scoring Coreference Chains with Split-Antecedent Anaphors
\nameSilviu Paun \addrSchool of Electronic Engineering and Computer Science Queen Mary University of London \AND\nameJuntao Yu∗ \addrSchool of Computer Science and Electronic Engineering University of Essex \AND\nameNafise Sadat Moosavi \addrDepartment of Computer Science University of Sheffield \AND\nameMassimo Poesio \addrSchool of Electronic Engineering and Computer Science Queen Mary University of London Equal contribution. Listed by alphabetical order
Abstract
Anaphoric reference is an aspect of language interpretation covering a variety of types of interpretation beyond the simple case of identity reference to entities introduced via nominal expressions covered by the traditional coreference task in its most recent incarnation in ontonotes and similar datasets. One of these cases that go beyond simple coreference is anaphoric reference to entities that must be added to the discourse model via accommodation, and in particular split-antecedent references to entities constructed out of other entities, as in split-antecedent plurals and in some cases of discourse deixis. Although this type of anaphoric reference is now annotated in many datasets, systems interpreting such references cannot be evaluated using the Reference coreference scorer Pradhan et al. (2014). As part of the work towards a new scorer for anaphoric reference able to evaluate all aspects of anaphoric interpretation in the coverage of the Universal Anaphora initiative, we propose in this paper a solution to the technical problem of generalizing existing metrics for identity anaphora so that they can also be used to score cases of split-antecedents. This is the first such pr
中文速览
指代消解(anaphora resolution)领域长期依赖的评分工具只能处理最简单的"同一实体"指代,却无法评价"分裂先行语"(split-antecedent)这类更复杂的情形——比如"John遇见了Mary,他们去看了电影"中,"他们"同时指向两个分别引入的实体。为了填补这一空白,本文提出了一套将现有主流共指评价指标(MUC、B³、CEAF、LEA、BLANC)统一推广到分裂先行语场景的方案,核心思路是把系统预测的"复合指代集合"与标准答案进行精细匹配,从而让单一先行语与多先行语的指代可以用完全相同的框架打分。实验表明,这套方案成功用于2021年CODI/CRAC对话指代消解共享任务的官方评测,覆盖了分裂先行语复数指代和话语回指(discourse deixis)两类场景,并在行为分析上优于此前所有同类方案。这是文献中首个能统一评价各类指代(包括需要"顺应推断"才能构造先行语的情形)的通用评分框架,为推动通用指代消解评测标准化奠定了基础。
原文 arXiv:2205.12323;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2205.12323v1