A large-scale study on research code quality and execution

Trisovic, Ana; Lau, Matthew K.; Pasquier, Thomas; Crosas, Mercè

Computer Science > Software Engineering

arXiv:2103.12793 (cs)

[Submitted on 23 Mar 2021]

Title:A large-scale study on research code quality and execution

Authors:Ana Trisovic, Matthew K. Lau, Thomas Pasquier, Mercè Crosas

View PDF

Abstract:This article presents a study on the quality and execution of research code from publicly-available replication datasets at the Harvard Dataverse repository. Research code is typically created by a group of scientists and published together with academic papers to facilitate research transparency and reproducibility. For this study, we define ten questions to address aspects impacting research reproducibility and reuse. First, we retrieve and analyze more than 2000 replication datasets with over 9000 unique R files published from 2010 to 2020. Second, we execute the code in a clean runtime environment to assess its ease of reuse. Common coding errors were identified, and some of them were solved with automatic code cleaning to aid code execution. We find that 74\% of R files crashed in the initial execution, while 56\% crashed when code cleaning was applied, showing that many errors can be prevented with good coding practices. We also analyze the replication datasets from journals' collections and discuss the impact of the journal policy strictness on the code re-execution rate. Finally, based on our results, we propose a set of recommendations for code dissemination aimed at researchers, journals, and repositories.

Comments:	30 pages
Subjects:	Software Engineering (cs.SE); Digital Libraries (cs.DL)
Cite as:	arXiv:2103.12793 [cs.SE]
	(or arXiv:2103.12793v1 [cs.SE] for this version)
	https://doi.org/10.48550/arXiv.2103.12793

Submission history

From: Ana Trisovic [view email]
[v1] Tue, 23 Mar 2021 18:55:09 UTC (1,839 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.SE

< prev | next >

new | recent | 2021-03

Change to browse by:

cs
cs.DL

References & Citations

DBLP - CS Bibliography

listing | bibtex

Matthew K. Lau
Thomas F. J.-M. Pasquier
Mercè Crosas

export BibTeX citation

Computer Science > Software Engineering

Title:A large-scale study on research code quality and execution

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Software Engineering

Title:A large-scale study on research code quality and execution

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators