Fingerprinting and Building Large Reproducible Datasets

Lefeuvre, Romain; Galasso, Jessie; Combemale, Benoit; Sahraoui, Houari; Zacchiroli, Stefano

Abstract:Obtaining a relevant dataset is central to conducting empirical studies in software engineering. However, in the context of mining software repositories, the lack of appropriate tooling for large scale mining tasks hinders the creation of new datasets. Moreover, limitations related to data sources that change over time (e.g., code bases) and the lack of documentation of extraction processes make it difficult to reproduce datasets over time. This threatens the quality and reproducibility of empirical studies.
In this paper, we propose a tool-supported approach facilitating the creation of large tailored datasets while ensuring their reproducibility. We leveraged all the sources feeding the Software Heritage append-only archive which are accessible through a unified programming interface to outline a reproducible and generic extraction process. We propose a way to define a unique fingerprint to characterize a dataset which, when provided to the extraction process, ensures that the same dataset will be extracted.
We demonstrate the feasibility of our approach by implementing a prototype. We show how it can help reduce the limitations researchers face when creating or reproducing datasets.

Subjects:	Software Engineering (cs.SE)
Cite as:	arXiv:2306.11391 [cs.SE]
	(or arXiv:2306.11391v1 [cs.SE] for this version)
	https://doi.org/10.48550/arXiv.2306.11391

Computer Science > Software Engineering

Title:Fingerprinting and Building Large Reproducible Datasets

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators