-
A Bayesian feature allocation model for tumor heterogeneity
Authors:
Juhee Lee,
Peter Müller,
Kamalakar Gulukota,
Yuan Ji
Abstract:
We develop a feature allocation model for inference on genetic tumor variation using next-generation sequencing data. Specifically, we record single nucleotide variants (SNVs) based on short reads mapped to human reference genome and characterize tumor heterogeneity by latent haplotypes defined as a scaffold of SNVs on the same homologous genome. For multiple samples from a single tumor, assuming…
▽ More
We develop a feature allocation model for inference on genetic tumor variation using next-generation sequencing data. Specifically, we record single nucleotide variants (SNVs) based on short reads mapped to human reference genome and characterize tumor heterogeneity by latent haplotypes defined as a scaffold of SNVs on the same homologous genome. For multiple samples from a single tumor, assuming that each sample is composed of some sample-specific proportions of these haplotypes, we then fit the observed variant allele fractions of SNVs for each sample and estimate the proportions of haplotypes. Varying proportions of haplotypes across samples is evidence of tumor heterogeneity since it implies varying composition of cell subpopulations. Taking a Bayesian perspective, we proceed with a prior probability model for all relevant unknown quantities, including, in particular, a prior probability model on the binary indicators that characterize the latent haplotypes. Such prior models are known as feature allocation models. Specifically, we define a simplified version of the Indian buffet process, one of the most traditional feature allocation models. The proposed model allows overlapping clustering of SNVs in defining latent haplotypes, which reflects the evolutionary process of subclonal expansion in tumor samples.
△ Less
Submitted 14 September, 2015;
originally announced September 2015.
-
Bayesian Inference for Tumor Subclones Accounting for Sequencing and Structural Variants
Authors:
Juhee Lee,
Peter Mueller,
Subhajit Sengupta,
Kamalakar Gulukota,
Yuan Ji
Abstract:
Tumor samples are heterogeneous. They consist of different subclones that are characterized by differences in DNA nucleotide sequences and copy numbers on multiple loci. Heterogeneity can be measured through the identification of the subclonal copy number and sequence at a selected set of loci. Understanding that the accurate identification of variant allele fractions greatly depends on a precise…
▽ More
Tumor samples are heterogeneous. They consist of different subclones that are characterized by differences in DNA nucleotide sequences and copy numbers on multiple loci. Heterogeneity can be measured through the identification of the subclonal copy number and sequence at a selected set of loci. Understanding that the accurate identification of variant allele fractions greatly depends on a precise determination of copy numbers, we develop a Bayesian feature allocation model for jointly calling subclonal copy numbers and the corresponding allele sequences for the same loci. The proposed method utilizes three random matrices, L, Z and w to represent subclonal copy numbers (L), numbers of subclonal variant alleles (Z) and cellular fractions of subclones in samples (w), respectively. The unknown number of subclones implies a random number of columns for these matrices. We use next-generation sequencing data to estimate the subclonal structures through inference on these three matrices. Using simulation studies and a real data analysis, we demonstrate how posterior inference on the subclonal structure is enhanced with the joint modeling of both structure and sequencing variants on subclonal genomes. Software is available at http://compgenome.org/BayClone2.
△ Less
Submitted 25 September, 2014;
originally announced September 2014.
-
MAD Bayes for Tumor Heterogeneity Feature Allocation with Non-Normal Sampling
Authors:
Yanxun Xu,
Peter Mueller,
Yuan Yuan,
Kamalakar Gulukota,
Yuan Ji
Abstract:
We propose small-variance asymptotic approximations for the inference of tumor heterogeneity (TH) using next-generation sequencing data. Understanding TH is an important and open research problem in biology. The lack of appropriate statistical inference is a critical gap in existing methods that the proposed approach aims to fill. We build on a hierarchical model with an exponential family likelih…
▽ More
We propose small-variance asymptotic approximations for the inference of tumor heterogeneity (TH) using next-generation sequencing data. Understanding TH is an important and open research problem in biology. The lack of appropriate statistical inference is a critical gap in existing methods that the proposed approach aims to fill. We build on a hierarchical model with an exponential family likelihood and a feature allocation prior. The proposed approach generalizes similar small-variance approximations proposed by Kulis and Jordan (2012) and Broderick et.al (2012) for inference with Dirichlet process mixture and Indian buffet prior models under normal sampling. We show that the new algorithm can successfully recover latent structures of different subclones and is also magnitude faster than available Markov chain Monte Carlo samplers, the latter often practically infeasible for high-dimensional genomics data. The proposed approach is scalable, simple to implement and benefits from the flexibility of Bayesian nonparametric models. More importantly, it provides a useful tool for the biological community for estimating cell subtypes in tumor samples.
△ Less
Submitted 2 December, 2014; v1 submitted 20 February, 2014;
originally announced February 2014.