A Grounded Unsupervised Universal Part-of-Speech Tagger for Low-Resource Languages

Cardenas, Ronald; Lin, Ying; Ji, Heng; May, Jonathan

Computer Science > Computation and Language

arXiv:1904.05426 (cs)

[Submitted on 10 Apr 2019]

Title:A Grounded Unsupervised Universal Part-of-Speech Tagger for Low-Resource Languages

Authors:Ronald Cardenas, Ying Lin, Heng Ji, Jonathan May

View PDF

Abstract:Unsupervised part of speech (POS) tagging is often framed as a clustering problem, but practical taggers need to \textit{ground} their clusters as well. Grounding generally requires reference labeled data, a luxury a low-resource language might not have. In this work, we describe an approach for low-resource unsupervised POS tagging that yields fully grounded output and requires no labeled training data. We find the classic method of Brown et al. (1992) clusters well in our use case and employ a decipherment-based approach to grounding. This approach presumes a sequence of cluster IDs is a `ciphertext' and seeks a POS tag-to-cluster ID mapping that will reveal the POS sequence. We show intrinsically that, despite the difficulty of the task, we obtain reasonable performance across a variety of languages. We also show extrinsically that incorporating our POS tagger into a name tagger leads to state-of-the-art tagging performance in Sinhalese and Kinyarwanda, two languages with nearly no labeled POS data available. We further demonstrate our tagger's utility by incorporating it into a true `zero-resource' variant of the Malopa (Ammar et al., 2016) dependency parser model that removes the current reliance on multilingual resources and gold POS tags for new languages. Experiments show that including our tagger makes up much of the accuracy lost when gold POS tags are unavailable.

Comments:	NAACL-HLT 2019, 12 pages, code available at this https URL
Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as:	arXiv:1904.05426 [cs.CL]
	(or arXiv:1904.05426v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.1904.05426

Submission history

From: Ronald Cardenas Acosta [view email]
[v1] Wed, 10 Apr 2019 20:22:31 UTC (163 KB)

Computer Science > Computation and Language

Title:A Grounded Unsupervised Universal Part-of-Speech Tagger for Low-Resource Languages

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:A Grounded Unsupervised Universal Part-of-Speech Tagger for Low-Resource Languages

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators