Finding Universal Grammatical Relations in Multilingual BERT

Chi, Ethan A.; Hewitt, John; Manning, Christopher D.

Computer Science > Computation and Language

arXiv:2005.04511 (cs)

[Submitted on 9 May 2020 (v1), last revised 20 May 2020 (this version, v2)]

Title:Finding Universal Grammatical Relations in Multilingual BERT

Authors:Ethan A. Chi, John Hewitt, Christopher D. Manning

View PDF

Abstract:Recent work has found evidence that Multilingual BERT (mBERT), a transformer-based multilingual masked language model, is capable of zero-shot cross-lingual transfer, suggesting that some aspects of its representations are shared cross-lingually. To better understand this overlap, we extend recent work on finding syntactic trees in neural networks' internal representations to the multilingual setting. We show that subspaces of mBERT representations recover syntactic tree distances in languages other than English, and that these subspaces are approximately shared across languages. Motivated by these results, we present an unsupervised analysis method that provides evidence mBERT learns representations of syntactic dependency labels, in the form of clusters which largely agree with the Universal Dependencies taxonomy. This evidence suggests that even without explicit supervision, multilingual masked language models learn certain linguistic universals.

Comments:	To appear in ACL 2020; Farsi typo corrected
Subjects:	Computation and Language (cs.CL); Machine Learning (cs.LG)
ACM classes:	I.2.7
Cite as:	arXiv:2005.04511 [cs.CL]
	(or arXiv:2005.04511v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2005.04511

Submission history

From: Ethan Chi [view email]
[v1] Sat, 9 May 2020 20:46:02 UTC (18,222 KB)
[v2] Wed, 20 May 2020 08:32:18 UTC (18,227 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.CL

< prev | next >

new | recent | 2020-05

Change to browse by:

cs
cs.LG

References & Citations

DBLP - CS Bibliography

listing | bibtex

John Hewitt
Christopher D. Manning

export BibTeX citation

Computer Science > Computation and Language

Title:Finding Universal Grammatical Relations in Multilingual BERT

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Finding Universal Grammatical Relations in Multilingual BERT

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators