Using Document Similarity Methods to create Parallel Datasets for Code Translation

Agarwal, Mayank; Talamadupula, Kartik; Martinez, Fernando; Houde, Stephanie; Muller, Michael; Richards, John; Ross, Steven I; Weisz, Justin D.

Computer Science > Computation and Language

arXiv:2110.05423 (cs)

[Submitted on 11 Oct 2021]

Title:Using Document Similarity Methods to create Parallel Datasets for Code Translation

Authors:Mayank Agarwal, Kartik Talamadupula, Fernando Martinez, Stephanie Houde, Michael Muller, John Richards, Steven I Ross, Justin D. Weisz

View PDF

Abstract:Translating source code from one programming language to another is a critical, time-consuming task in modernizing legacy applications and codebases. Recent work in this space has drawn inspiration from the software naturalness hypothesis by applying natural language processing techniques towards automating the code translation task. However, due to the paucity of parallel data in this domain, supervised techniques have only been applied to a limited set of popular programming languages. To bypass this limitation, unsupervised neural machine translation techniques have been proposed to learn code translation using only monolingual corpora. In this work, we propose to use document similarity methods to create noisy parallel datasets of code, thus enabling supervised techniques to be applied for automated code translation without having to rely on the availability or expensive curation of parallel code datasets. We explore the noise tolerance of models trained on such automatically-created datasets and show that these models perform comparably to models trained on ground truth for reasonable levels of noise. Finally, we exhibit the practical utility of the proposed method by creating parallel datasets for languages beyond the ones explored in prior work, thus expanding the set of programming languages for automated code translation.

Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2110.05423 [cs.CL]
	(or arXiv:2110.05423v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2110.05423

Submission history

From: Mayank Agarwal [view email]
[v1] Mon, 11 Oct 2021 17:07:58 UTC (293 KB)

Computer Science > Computation and Language

Title:Using Document Similarity Methods to create Parallel Datasets for Code Translation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Using Document Similarity Methods to create Parallel Datasets for Code Translation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators