Harnessing Cross-lingual Features to Improve Cognate Detection for Low-resource Languages

Kanojia, Diptesh; Dabre, Raj; Dewangan, Shubham; Bhattacharyya, Pushpak; Haffari, Gholamreza; Kulkarni, Malhar

Computer Science > Computation and Language

arXiv:2112.08789 (cs)

[Submitted on 16 Dec 2021]

Title:Harnessing Cross-lingual Features to Improve Cognate Detection for Low-resource Languages

Authors:Diptesh Kanojia, Raj Dabre, Shubham Dewangan, Pushpak Bhattacharyya, Gholamreza Haffari, Malhar Kulkarni

View PDF

Abstract:Cognates are variants of the same lexical form across different languages; for example 'fonema' in Spanish and 'phoneme' in English are cognates, both of which mean 'a unit of sound'. The task of automatic detection of cognates among any two languages can help downstream NLP tasks such as Cross-lingual Information Retrieval, Computational Phylogenetics, and Machine Translation. In this paper, we demonstrate the use of cross-lingual word embeddings for detecting cognates among fourteen Indian Languages. Our approach introduces the use of context from a knowledge graph to generate improved feature representations for cognate detection. We, then, evaluate the impact of our cognate detection mechanism on neural machine translation (NMT), as a downstream task. We evaluate our methods to detect cognates on a challenging dataset of twelve Indian languages, namely, Sanskrit, Hindi, Assamese, Oriya, Kannada, Gujarati, Tamil, Telugu, Punjabi, Bengali, Marathi, and Malayalam. Additionally, we create evaluation datasets for two more Indian languages, Konkani and Nepali. We observe an improvement of up to 18% points, in terms of F-score, for cognate detection. Furthermore, we observe that cognates extracted using our method help improve NMT quality by up to 2.76 BLEU. We also release our code, newly constructed datasets and cross-lingual models publicly.

Comments:	Published at COLING 2020
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2112.08789 [cs.CL]
	(or arXiv:2112.08789v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2112.08789

Submission history

From: Diptesh Kanojia [view email]
[v1] Thu, 16 Dec 2021 11:17:58 UTC (200 KB)

Computer Science > Computation and Language

Title:Harnessing Cross-lingual Features to Improve Cognate Detection for Low-resource Languages

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Harnessing Cross-lingual Features to Improve Cognate Detection for Low-resource Languages

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators