ezCoref: Towards Unifying Annotation Guidelines for Coreference Resolution

Gupta, Ankita; Karpinska, Marzena; Zhao, Wenlong; Krishna, Kalpesh; Merullo, Jack; Yeh, Luke; Iyyer, Mohit; O'Connor, Brendan

Computer Science > Computation and Language

arXiv:2210.07188 (cs)

[Submitted on 13 Oct 2022]

Title:ezCoref: Towards Unifying Annotation Guidelines for Coreference Resolution

Authors:Ankita Gupta, Marzena Karpinska, Wenlong Zhao, Kalpesh Krishna, Jack Merullo, Luke Yeh, Mohit Iyyer, Brendan O'Connor

View PDF

Abstract:Large-scale, high-quality corpora are critical for advancing research in coreference resolution. However, existing datasets vary in their definition of coreferences and have been collected via complex and lengthy guidelines that are curated for linguistic experts. These concerns have sparked a growing interest among researchers to curate a unified set of guidelines suitable for annotators with various backgrounds. In this work, we develop a crowdsourcing-friendly coreference annotation methodology, ezCoref, consisting of an annotation tool and an interactive tutorial. We use ezCoref to re-annotate 240 passages from seven existing English coreference datasets (spanning fiction, news, and multiple other domains) while teaching annotators only cases that are treated similarly across these datasets. Surprisingly, we find that reasonable quality annotations were already achievable (>90% agreement between the crowd and expert annotations) even without extensive training. On carefully analyzing the remaining disagreements, we identify the presence of linguistic cases that our annotators unanimously agree upon but lack unified treatments (e.g., generic pronouns, appositives) in existing datasets. We propose the research community should revisit these phenomena when curating future unified annotation guidelines.

Comments:	preprint (19 pages), code in this https URL
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2210.07188 [cs.CL]
	(or arXiv:2210.07188v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2210.07188

Submission history

From: Kalpesh Krishna [view email]
[v1] Thu, 13 Oct 2022 17:09:59 UTC (7,412 KB)

Computer Science > Computation and Language

Title:ezCoref: Towards Unifying Annotation Guidelines for Coreference Resolution

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:ezCoref: Towards Unifying Annotation Guidelines for Coreference Resolution

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators