-
Probabilistic Count Matrix Factorization for Single Cell Expression Data Analysis
Authors:
Ghislain Durif,
Laurent Modolo,
Jeff E. Mold,
Sophie Lambert-Lacroix,
Franck Picard
Abstract:
The development of high throughput single-cell sequencing technologies now allows the investigation of the population level diversity of cellular transcriptomes. This diversity has shown two faces. First, the expression dynamics (gene to gene variability) can be quantified more accurately, thanks to the measurement of lowly-expressed genes. Second, the cell-to-cell variability is high, with a low…
▽ More
The development of high throughput single-cell sequencing technologies now allows the investigation of the population level diversity of cellular transcriptomes. This diversity has shown two faces. First, the expression dynamics (gene to gene variability) can be quantified more accurately, thanks to the measurement of lowly-expressed genes. Second, the cell-to-cell variability is high, with a low proportion of cells expressing the same gene at the same time/level. Those emerging patterns appear to be very challenging from the statistical point of view, especially to represent and to provide a summarized view of single-cell expression data. PCA is one of the most powerful framework to provide a suitable representation of high dimensional datasets, by searching for latent directions catching the most variability in the data. Unfortunately, classical PCA is based on Euclidean distances and projections that work poorly in presence of over-dispersed counts that show drop-out events (zero-inflation) like single-cell expression data We propose a probabilistic Count Matrix Factorization (pCMF) approach for single-cell expression data analysis, that relies on a sparse Gamma-Poisson factor model. This hierarchical model is inferred using a variational EM algorithm. It is able to jointly build a low dimensional representation of cells and genes. We show how this probabilistic framework induces a geometry that is suitable for single-cell data visualization, and produces a compression of the data that is very powerful for clustering purposes. Our method is competed against other standard representation methods like t-SNE, and we illustrate its performance for the representation of single-cell data. We especially focus on publicly available single-cell RNA-seq datasets.
△ Less
Submitted 12 March, 2019; v1 submitted 30 October, 2017;
originally announced October 2017.
-
Minimax wavelet estimation for multisample heteroscedastic non-parametric regression
Authors:
Madison Giacofc,
Sophie Lambert-Lacroix,
Franck Picard
Abstract:
The problem of estimating the baseline signal from multisample noisy curves is investigated. We consider the functional mixed effects model, and we suppose that the functional fixed effect belongs to the Besov class. This framework allows us to model curves that can exhibit strong irregularities, such as peaks or jumps for instance. The lower bound for the $L_2$ minimax risk is provided, as well a…
▽ More
The problem of estimating the baseline signal from multisample noisy curves is investigated. We consider the functional mixed effects model, and we suppose that the functional fixed effect belongs to the Besov class. This framework allows us to model curves that can exhibit strong irregularities, such as peaks or jumps for instance. The lower bound for the $L_2$ minimax risk is provided, as well as the upper bound of the minimax rate, that is derived by constructing a wavelet estimator for the functional fixed effect. Our work constitutes the first theoretical functional results in multisample non parametric regression. Our approach is illustrated on realistic simulated datasets as well as on experimental data.
△ Less
Submitted 14 November, 2015;
originally announced November 2015.
-
High Dimensional Classification with combined Adaptive Sparse PLS and Logistic Regression
Authors:
G. Durif,
L. Modolo,
J. Michaelsson,
J. E. Mold,
S. Lambert-Lacroix,
F. Picard
Abstract:
Motivation: The high dimensionality of genomic data calls for the development of specific classification methodologies, especially to prevent over-optimistic predictions. This challenge can be tackled by compression and variable selection, which combined constitute a powerful framework for classification, as well as data visualization and interpretation. However, current proposed combinations lead…
▽ More
Motivation: The high dimensionality of genomic data calls for the development of specific classification methodologies, especially to prevent over-optimistic predictions. This challenge can be tackled by compression and variable selection, which combined constitute a powerful framework for classification, as well as data visualization and interpretation. However, current proposed combinations lead to instable and non convergent methods due to inappropriate computational frameworks. We hereby propose a stable and convergent approach for classification in high dimensional based on sparse Partial Least Squares (sparse PLS). Results: We start by proposing a new solution for the sparse PLS problem that is based on proximal operators for the case of univariate responses. Then we develop an adaptive version of the sparse PLS for classification, which combines iterative optimization of logistic regression and sparse PLS to ensure convergence and stability. Our results are confirmed on synthetic and experimental data. In particular we show how crucial convergence and stability can be when cross-validation is involved for calibration purposes. Using gene expression data we explore the prediction of breast cancer relapse. We also propose a multicategorial version of our method on the prediction of cell-types based on single-cell expression data. Availability: Our approach is implemented in the plsgenomics R-package.
△ Less
Submitted 30 August, 2017; v1 submitted 20 February, 2015;
originally announced February 2015.
-
The BerHu penalty and the grouped effect
Authors:
Laurent Zwald,
Sophie Lambert-Lacroix
Abstract:
The Huber's criterion is a useful method for robust regression. The adaptive least absolute shrinkage and selection operator (lasso) is a popular technique for simultaneous estimation and variable selection. In the case of small sample size and large covariables numbers, this penalty is not very satisfactory variable selection method. In this paper, we introduce an adaptive reversed version of Hub…
▽ More
The Huber's criterion is a useful method for robust regression. The adaptive least absolute shrinkage and selection operator (lasso) is a popular technique for simultaneous estimation and variable selection. In the case of small sample size and large covariables numbers, this penalty is not very satisfactory variable selection method. In this paper, we introduce an adaptive reversed version of Huber's criterion as a penalty function. We call this penalty adaptive Berhu penalty. As for elastic net penalty, small coefficients contribute their $\ell_1$ norm to this penalty while larger coefficients cause it to grow quadratically (as ridge regression). We show that the estimator associated with criterion such that ordinary least square or Huber's one combining with adaptive Berhu penalty enjoys the oracle properties. In addition, this procedure encourages a grouping effect. This approach is compared with adaptive elastic net regularization. Extensive simulation studies demonstrate satisfactory finite-sample performance of such procedure. A real example is analyzed for illustration purposes.
Keywords : Adaptive Berhu penalty; concomitant scale; elastic net penalty; Huber's criterion; oracle property; robust estimation.
△ Less
Submitted 30 July, 2012;
originally announced July 2012.