Datasheets for AI and medical datasets (DAIMS): a data validation and documentation framework before machine learning analysis in medical research

Marandi, Ramtin Zargari; Frahm, Anne Svane; Milojevic, Maja

Computer Science > Machine Learning

arXiv:2501.14094 (cs)

[Submitted on 23 Jan 2025]

Title:Datasheets for AI and medical datasets (DAIMS): a data validation and documentation framework before machine learning analysis in medical research

Authors:Ramtin Zargari Marandi (1), Anne Svane Frahm (1), Maja Milojevic (1)

View PDF

Abstract:Despite progresses in data engineering, there are areas with limited consistencies across data validation and documentation procedures causing confusions and technical problems in research involving machine learning. There have been progresses by introducing frameworks like "Datasheets for Datasets", however there are areas for improvements to prepare datasets, ready for ML pipelines. Here, we extend the framework to "Datasheets for AI and medical datasets - DAIMS." Our publicly available solution, DAIMS, provides a checklist including data standardization requirements, a software tool to assist the process of the data preparation, an extended form for data documentation and pose research questions, a table as data dictionary, and a flowchart to suggest ML analyses to address the research questions. The checklist consists of 24 common data standardization requirements, where the tool checks and validate a subset of them. In addition, we provided a flowchart mapping research questions to suggested ML methods. DAIMS can serve as a reference for standardizing datasets and a roadmap for researchers aiming to apply effective ML techniques in their medical research endeavors. DAIMS is available on GitHub and as an online app to automate key aspects of dataset evaluation, facilitating efficient preparation of datasets for ML studies.

Comments:	10 pages, 1 figure, 2 tables
Subjects:	Machine Learning (cs.LG)
Cite as:	arXiv:2501.14094 [cs.LG]
	(or arXiv:2501.14094v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2501.14094

Submission history

From: Ramtin Zargari Marandi Dr [view email]
[v1] Thu, 23 Jan 2025 21:02:56 UTC (864 KB)

Computer Science > Machine Learning

Title:Datasheets for AI and medical datasets (DAIMS): a data validation and documentation framework before machine learning analysis in medical research

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Datasheets for AI and medical datasets (DAIMS): a data validation and documentation framework before machine learning analysis in medical research

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators