Interlacing Personal and Reference Genomes for Machine Learning Disease-Variant Detection

Harries, Luke R; Zhang, Suyi; Dubourg-Felonneau, Geoffroy; Farmery, James H R; Sinai, Jonathan; Taylor, Belle; Patel, Nirmesh; Cassidy, John W; Shawe-Taylor, John; Clifford, Harry W

Quantitative Biology > Genomics

arXiv:1811.11674 (q-bio)

[Submitted on 26 Nov 2018]

Title:Interlacing Personal and Reference Genomes for Machine Learning Disease-Variant Detection

Authors:Luke R Harries, Suyi Zhang, Geoffroy Dubourg-Felonneau, James H R Farmery, Jonathan Sinai, Belle Taylor, Nirmesh Patel, John W Cassidy, John Shawe-Taylor, Harry W Clifford

View PDF

Abstract:DNA sequencing to identify genetic variants is becoming increasingly valuable in clinical settings. Assessment of variants in such sequencing data is commonly implemented through Bayesian heuristic algorithms. Machine learning has shown great promise in improving on these variant calls, but the input for these is still a standardized "pile-up" image, which is not always best suited. In this paper, we present a novel method for generating images from DNA sequencing data, which interlaces the human reference genome with personalized sequencing output, to maximize usage of sequencing reads and improve machine learning algorithm performance. We demonstrate the success of this in improving standard germline variant calling. We also furthered this approach to include somatic variant calling across tumor/normal data with Siamese networks. These approaches can be used in machine learning applications on sequencing data with the hope of improving clinical outcomes, and are freely available for noncommercial use at this http URL.

Comments:	Machine Learning for Health (ML4H) Workshop at NeurIPS 2018 arXiv:cs/0101200
Subjects:	Genomics (q-bio.GN); Machine Learning (cs.LG); Machine Learning (stat.ML)
Report number:	ML4H/2018/103
Cite as:	arXiv:1811.11674 [q-bio.GN]
	(or arXiv:1811.11674v1 [q-bio.GN] for this version)
	https://doi.org/10.48550/arXiv.1811.11674

Submission history

From: Harry Clifford [view email]
[v1] Mon, 26 Nov 2018 15:38:29 UTC (299 KB)

Quantitative Biology > Genomics

Title:Interlacing Personal and Reference Genomes for Machine Learning Disease-Variant Detection

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Quantitative Biology > Genomics

Title:Interlacing Personal and Reference Genomes for Machine Learning Disease-Variant Detection

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators