-
Assessing the impact of variance heterogeneity and misspecification in mixed-effects location-scale models
Authors:
Vincent Jeanselme,
Marco Palma,
Jessica K Barrett
Abstract:
Linear Mixed Model (LMM) is a common statistical approach to model the relation between exposure and outcome while capturing individual variability through random effects. However, this model assumes the homogeneity of the error term's variance. Breaking this assumption, known as homoscedasticity, can bias estimates and, consequently, may change a study's conclusions. If this assumption is unmet,…
▽ More
Linear Mixed Model (LMM) is a common statistical approach to model the relation between exposure and outcome while capturing individual variability through random effects. However, this model assumes the homogeneity of the error term's variance. Breaking this assumption, known as homoscedasticity, can bias estimates and, consequently, may change a study's conclusions. If this assumption is unmet, the mixed-effect location-scale model (MELSM) offers a solution to account for within-individual variability. Our work explores how LMMs and MELSMs behave when the homoscedasticity assumption is not met. Further, we study how misspecification affects inference for MELSM. To this aim, we propose a simulation study with longitudinal data and evaluate the estimates' bias and coverage. Our simulations show that neglecting heteroscedasticity in LMMs leads to loss of coverage for the estimated coefficients and biases the estimates of the standard deviations of the random effects. In MELSMs, scale misspecification does not bias the location model, but location misspecification alters the scale estimates. Our simulation study illustrates the importance of modelling heteroscedasticity, with potential implications beyond mixed effect models, for generalised linear mixed models for non-normal outcomes and joint models with survival data.
△ Less
Submitted 23 May, 2025;
originally announced May 2025.
-
A Bayesian location-scale joint model for time-to-event and multivariate longitudinal data with association based on within-individual variability
Authors:
Marco Palma,
Ruth H Keogh,
Siobhán B Carr,
Rhonda Szczesniak,
David Taylor-Robinson,
Angela M Wood,
Graciela Muniz-Terrera,
Jessica K Barrett
Abstract:
Within-individual variability of health indicators measured over time is becoming commonly used to inform about disease progression. Simple summary statistics (e.g. the standard deviation for each individual) are often used but they are not suited to account for time changes. In addition, when these summary statistics are used as covariates in a regression model for time-to-event outcomes, the est…
▽ More
Within-individual variability of health indicators measured over time is becoming commonly used to inform about disease progression. Simple summary statistics (e.g. the standard deviation for each individual) are often used but they are not suited to account for time changes. In addition, when these summary statistics are used as covariates in a regression model for time-to-event outcomes, the estimates of the hazard ratios are subject to regression dilution. To overcome these issues, a joint model is built where the association between the time-to-event outcome and multivariate longitudinal markers is specified in terms of the within-individual variability of the latter. A mixed-effect location-scale model is used to analyse the longitudinal biomarkers, their within-individual variability and their correlation. The time to event is modelled using a proportional hazard regression model, with a flexible specification of the baseline hazard, and the information from the longitudinal biomarkers is shared as a function of the random effects. The model can be used to quantify within-individual variability for the longitudinal markers and their association with the time-to-event outcome. We show through a simulation study the performance of the model in comparison with the standard joint model with constant variance. The model is applied on a dataset of adult women from the UK cystic fibrosis registry, to evaluate the association between lung function, malnutrition and mortality.
△ Less
Submitted 15 March, 2025;
originally announced March 2025.
-
HAVER: Instance-Dependent Error Bounds for Maximum Mean Estimation and Applications to Q-Learning and Monte Carlo Tree Search
Authors:
Tuan Ngo Nguyen,
Jay Barrett,
Kwang-Sung Jun
Abstract:
We study the problem of estimating the \emph{value} of the largest mean among K distributions via samples from them (rather than estimating \emph{which} distribution has the largest mean), which arises from various machine learning tasks including Q-learning and Monte Carlo Tree Search (MCTS). While there have been a few proposed algorithms, their performance analyses have been limited to their bi…
▽ More
We study the problem of estimating the \emph{value} of the largest mean among K distributions via samples from them (rather than estimating \emph{which} distribution has the largest mean), which arises from various machine learning tasks including Q-learning and Monte Carlo Tree Search (MCTS). While there have been a few proposed algorithms, their performance analyses have been limited to their biases rather than a precise error metric. In this paper, we propose a novel algorithm called HAVER (Head AVERaging) and analyze its mean squared error. Our analysis reveals that HAVER has a compelling performance in two respects. First, HAVER estimates the maximum mean as well as the oracle who knows the identity of the best distribution and reports its sample mean. Second, perhaps surprisingly, HAVER exhibits even better rates than this oracle when there are many distributions near the best one. Both of these improvements are the first of their kind in the literature, and we also prove that the naive algorithm that reports the largest empirical mean does not achieve these bounds. Finally, we confirm our theoretical findings via numerical experiments where we implement HAVER in bandit, Q-learning, and MCTS algorithms. In these experiments, HAVER consistently outperforms the baseline methods, demonstrating its effectiveness across different applications.
△ Less
Submitted 28 April, 2025; v1 submitted 1 November, 2024;
originally announced November 2024.
-
Bayesian shared parameter joint models for heterogeneous populations
Authors:
Sida Chen,
Danilo Alvares,
Marco Palma,
Jessica K. Barrett
Abstract:
Joint models (JMs) for longitudinal and time-to-event data are an important class of biostatistical models in health and medical research. When the study population consists of heterogeneous subgroups, the standard JM may be inadequate and lead to misleading results. Joint latent class models (JLCMs) and their variants have been proposed to incorporate latent class structures into JMs. JLCMs are u…
▽ More
Joint models (JMs) for longitudinal and time-to-event data are an important class of biostatistical models in health and medical research. When the study population consists of heterogeneous subgroups, the standard JM may be inadequate and lead to misleading results. Joint latent class models (JLCMs) and their variants have been proposed to incorporate latent class structures into JMs. JLCMs are useful for identifying latent subgroup structures, obtaining a more nuanced understanding of the relationships between longitudinal outcomes, and improving prediction performance. We consider the generic form of JLCM, which poses significant computational challenges for both frequentist and Bayesian approaches due to the numerical intractability and multimodality of the associated model's likelihood or posterior. Focusing on the less explored Bayesian paradigm, we propose a new Bayesian inference framework to tackle key limitations in the existing method. Our algorithm leverages state-of-the-art Markov chain Monte Carlo techniques and parallel computing for parameter estimation and model selection. Through a simulation study, we demonstrate the feasibility and superiority of our proposed method over the existing approach. Our simulations also generate important computational insights and practical guidance for implementing such complex models. We illustrate our method using data from the PAQUID prospective cohort study, where we jointly investigate the association between a repeatedly measured cognitive score and the risk of dementia and the latent class structure defined from the longitudinal outcomes.
△ Less
Submitted 29 October, 2024;
originally announced October 2024.
-
Identifying treatment response subgroups in observational time-to-event data
Authors:
Vincent Jeanselme,
Chang Ho Yoon,
Fabian Falck,
Brian Tom,
Jessica Barrett
Abstract:
Identifying patient subgroups with different treatment responses is an important task to inform medical recommendations, guidelines, and the design of future clinical trials. Existing approaches for treatment effect estimation primarily rely on Randomised Controlled Trials (RCTs), which are often limited by insufficient power, multiple comparisons, and unbalanced covariates. In addition, RCTs tend…
▽ More
Identifying patient subgroups with different treatment responses is an important task to inform medical recommendations, guidelines, and the design of future clinical trials. Existing approaches for treatment effect estimation primarily rely on Randomised Controlled Trials (RCTs), which are often limited by insufficient power, multiple comparisons, and unbalanced covariates. In addition, RCTs tend to feature more homogeneous patient groups, making them less relevant for uncovering subgroups in the population encountered in real-world clinical practice. Subgroup analyses established for RCTs suffer from significant statistical biases when applied to observational studies, which benefit from larger and more representative populations. Our work introduces a novel, outcome-guided, subgroup analysis strategy for identifying subgroups of treatment response in both RCTs and observational studies alike. It hence positions itself in-between individualised and average treatment effect estimation to uncover patient subgroups with distinct treatment responses, critical for actionable insights that may influence treatment guidelines. In experiments, our approach significantly outperforms the current state-of-the-art method for subgroup analysis in both randomised and observational treatment regimes.
△ Less
Submitted 23 February, 2025; v1 submitted 6 August, 2024;
originally announced August 2024.
-
A Bayesian joint model of multiple longitudinal and categorical outcomes with application to multiple myeloma using permutation-based variable importance
Authors:
Danilo Alvares,
Jessica K. Barrett,
François Mercier,
Jochen Schulze,
Sean Yiu,
Felipe Castro,
Spyros Roumpanis,
Yajing Zhu
Abstract:
Joint models have proven to be an effective approach for uncovering potentially hidden connections between various types of outcomes, mainly continuous, time-to-event, and binary. Typically, longitudinal continuous outcomes are characterized by linear mixed-effects models, survival outcomes are described by proportional hazards models, and the link between outcomes are captured by shared random ef…
▽ More
Joint models have proven to be an effective approach for uncovering potentially hidden connections between various types of outcomes, mainly continuous, time-to-event, and binary. Typically, longitudinal continuous outcomes are characterized by linear mixed-effects models, survival outcomes are described by proportional hazards models, and the link between outcomes are captured by shared random effects. Other modeling variations include generalized linear mixed-effects models for longitudinal data and logistic regression when a binary outcome is present, rather than time until an event of interest. However, in a clinical research setting, one might be interested in modeling the physician's chosen treatment based on the patient's medical history to identify prognostic factors. In this situation, there are often multiple treatment options, requiring the use of a multiclass classification approach. Inspired by this context, we develop a Bayesian joint model for longitudinal and categorical data. In particular, our motivation comes from a multiple myeloma study, in which biomarkers display nonlinear trajectories that are well captured through bi-exponential submodels, where patient-level information is shared with the categorical submodel. We also present a variable importance strategy to rank prognostic factors. We apply our proposal and a competing model to the multiple myeloma data, compare the variable importance and inferential results for both models, and illustrate patient-level interpretations using our joint model.
△ Less
Submitted 14 June, 2025; v1 submitted 19 July, 2024;
originally announced July 2024.
-
A Bayesian joint model of multiple nonlinear longitudinal and competing risks outcomes for dynamic prediction in multiple myeloma: joint estimation and corrected two-stage approaches
Authors:
Danilo Alvares,
Jessica K. Barrett,
François Mercier,
Spyros Roumpanis,
Sean Yiu,
Felipe Castro,
Jochen Schulze,
Yajing Zhu
Abstract:
Predicting cancer-associated clinical events is challenging in oncology. In Multiple Myeloma (MM), a cancer of plasma cells, disease progression is determined by changes in biomarkers, such as serum concentration of the paraprotein secreted by plasma cells (M-protein). Therefore, the time-dependent behaviour of M-protein and the transition across lines of therapy (LoT) that may be a consequence of…
▽ More
Predicting cancer-associated clinical events is challenging in oncology. In Multiple Myeloma (MM), a cancer of plasma cells, disease progression is determined by changes in biomarkers, such as serum concentration of the paraprotein secreted by plasma cells (M-protein). Therefore, the time-dependent behaviour of M-protein and the transition across lines of therapy (LoT) that may be a consequence of disease progression should be accounted for in statistical models to predict relevant clinical outcomes. Furthermore, it is important to understand the contribution of the patterns of longitudinal biomarkers, upon each LoT initiation, to time-to-death or time-to-next-LoT. Motivated by these challenges, we propose a Bayesian joint model for trajectories of multiple longitudinal biomarkers, such as M-protein, and the competing risks of death and transition to next LoT. Additionally, we explore two estimation approaches for our joint model: simultaneous estimation of all parameters (joint estimation) and sequential estimation of parameters using a corrected two-stage strategy aiming to reduce computational time. Our proposed model and estimation methods are applied to a retrospective cohort study from a real-world database of patients diagnosed with MM in the US from January 2015 to February 2022. We split the data into training and test sets in order to validate the joint model using both estimation approaches and make dynamic predictions of times until clinical events of interest, informed by longitudinally measured biomarkers and baseline variables available up to the time of prediction.
△ Less
Submitted 11 November, 2024; v1 submitted 30 May, 2024;
originally announced May 2024.
-
Bayesian blockwise inference for joint models of longitudinal and multistate processes
Authors:
Sida Chen,
Danilo Alvares,
Christopher Jackson,
Jessica Barrett
Abstract:
Joint models (JM) for longitudinal and survival data have gained increasing interest and found applications in a wide range of clinical and biomedical settings. These models facilitate the understanding of the relationship between outcomes and enable individualized predictions. In many applications, more complex event processes arise, necessitating joint longitudinal and multistate models. However…
▽ More
Joint models (JM) for longitudinal and survival data have gained increasing interest and found applications in a wide range of clinical and biomedical settings. These models facilitate the understanding of the relationship between outcomes and enable individualized predictions. In many applications, more complex event processes arise, necessitating joint longitudinal and multistate models. However, their practical application can be hindered by computational challenges due to increased model complexity and large sample sizes. Motivated by a longitudinal multimorbidity analysis of large UK health records, we have developed a scalable Bayesian methodology for such joint multistate models that is capable of handling complex event processes and large datasets, with straightforward implementation. We propose two blockwise inference approaches for different inferential purposes based on different levels of decomposition of the multistate processes. These approaches leverage parallel computing, ease the specification of different models for different transitions, and model/variable selection can be performed within a Bayesian framework using Bayesian leave-one-out cross-validation. Using a simulation study, we show that the proposed approaches achieve satisfactory performance regarding posterior point and interval estimation, with notable gains in sampling efficiency compared to the standard estimation strategy. We illustrate our approaches using a large UK electronic health record dataset where we analysed the coevolution of routinely measured systolic blood pressure (SBP) and the progression of multimorbidity, defined as the combinations of three chronic conditions. Our analysis identified distinct association structures between SBP and different disease transitions.
△ Less
Submitted 23 August, 2023;
originally announced August 2023.
-
Neural Fine-Gray: Monotonic neural networks for competing risks
Authors:
Vincent Jeanselme,
Chang Ho Yoon,
Brian Tom,
Jessica Barrett
Abstract:
Time-to-event modelling, known as survival analysis, differs from standard regression as it addresses censoring in patients who do not experience the event of interest. Despite competitive performances in tackling this problem, machine learning methods often ignore other competing risks that preclude the event of interest. This practice biases the survival estimation. Extensions to address this ch…
▽ More
Time-to-event modelling, known as survival analysis, differs from standard regression as it addresses censoring in patients who do not experience the event of interest. Despite competitive performances in tackling this problem, machine learning methods often ignore other competing risks that preclude the event of interest. This practice biases the survival estimation. Extensions to address this challenge often rely on parametric assumptions or numerical estimations leading to sub-optimal survival approximations. This paper leverages constrained monotonic neural networks to model each competing survival distribution. This modelling choice ensures the exact likelihood maximisation at a reduced computational cost by using automatic differentiation. The effectiveness of the solution is demonstrated on one synthetic and three medical datasets. Finally, we discuss the implications of considering competing risks when developing risk scores for medical practice.
△ Less
Submitted 11 May, 2023;
originally announced May 2023.
-
A Framework for Understanding Selection Bias in Real-World Healthcare Data
Authors:
Ritoban Kundu,
Xu Shi,
Jean Morrison,
Jessica Barrett,
Bhramar Mukherjee
Abstract:
Using administrative patient-care data such as Electronic Health Records (EHR) and medical/ pharmaceutical claims for population-based scientific research has become increasingly common. With vast sample sizes leading to very small standard errors, researchers need to pay more attention to potential biases in the estimates of association parameters of interest, specifically to biases that do not d…
▽ More
Using administrative patient-care data such as Electronic Health Records (EHR) and medical/ pharmaceutical claims for population-based scientific research has become increasingly common. With vast sample sizes leading to very small standard errors, researchers need to pay more attention to potential biases in the estimates of association parameters of interest, specifically to biases that do not diminish with increasing sample size. Of these multiple sources of biases, in this paper, we focus on understanding selection bias. We present an analytic framework using directed acyclic graphs for guiding applied researchers to dissect how different sources of selection bias may affect estimates of the association between a binary outcome and an exposure (continuous or categorical) of interest. We consider four easy-to-implement weighting approaches to reduce selection bias with accompanying variance formulae. We demonstrate through a simulation study when they can rescue us in practice with analysis of real world data. We compare these methods using a data example where our goal is to estimate the well-known association of cancer and biological sex, using EHR from a longitudinal biorepository at the University of Michigan Healthcare system. We provide annotated R codes to implement these weighted methods with associated inference.
△ Less
Submitted 17 August, 2023; v1 submitted 10 April, 2023;
originally announced April 2023.
-
Optimal risk-assessment scheduling for primary prevention of cardiovascular disease
Authors:
Francesca Gasperoni,
Christopher H. Jackson,
Angela M. Wood,
Michael J. Sweeting,
Paul J. Newcombe,
David Stevens,
Jessica K. Barrett
Abstract:
In this work, we introduce a personalised and age-specific Net Benefit function, composed of benefits and costs, to recommend optimal timing of risk assessments for cardiovascular disease prevention. We extend the 2-stage landmarking model to estimate patient-specific CVD risk profiles, adjusting for time-varying covariates. We apply our model to data from the Clinical Practice Research Datalink,…
▽ More
In this work, we introduce a personalised and age-specific Net Benefit function, composed of benefits and costs, to recommend optimal timing of risk assessments for cardiovascular disease prevention. We extend the 2-stage landmarking model to estimate patient-specific CVD risk profiles, adjusting for time-varying covariates. We apply our model to data from the Clinical Practice Research Datalink, comprising primary care electronic health records from the UK. We find that people at lower risk could be recommended an optimal risk-assessment interval of 5 years or more. Time-varying risk-factors are required to discriminate between more frequent schedules for higher-risk people.
△ Less
Submitted 9 February, 2023;
originally announced February 2023.
-
Sample Size Estimation using a Latent Variable Model for Mixed Outcome Co-Primary, Multiple Primary and Composite Endpoints
Authors:
Martina McMenamin,
Jessica K. Barrett,
Anna Berglind,
James M. S. Wason
Abstract:
Mixed outcome endpoints that combine multiple continuous and discrete components to form co-primary, multiple primary or composite endpoints are often employed as primary outcome measures in clinical trials. There are many advantages to joint modelling the individual outcomes using a latent variable framework, however in order to make use of the model in practice we require techniques for sample s…
▽ More
Mixed outcome endpoints that combine multiple continuous and discrete components to form co-primary, multiple primary or composite endpoints are often employed as primary outcome measures in clinical trials. There are many advantages to joint modelling the individual outcomes using a latent variable framework, however in order to make use of the model in practice we require techniques for sample size estimation. In this paper we show how the latent variable model can be applied to the three types of joint endpoints and propose appropriate hypotheses, power and sample size estimation methods for each. We illustrate the techniques using a numerical example based on the four dimensional endpoint in the MUSE trial and find that the sample size required for the co-primary endpoint is larger than that required for the individual endpoint with the smallest effect size. Conversely, the sample size required for the multiple primary endpoint is reduced from that required for the individual outcome with the largest effect size. We show that the analytical technique agrees with the empirical power from simulation studies. We further illustrate the reduction in required sample size that may be achieved in trials of mixed outcome composite endpoints through a simulation study and find that the sample size primarily depends on the components driving response and the correlation structure and much less so on the treatment effect structure in the individual endpoints.
△ Less
Submitted 11 December, 2019;
originally announced December 2019.
-
Selective recruitment designs for improving observational studies using electronic health records
Authors:
James E. Barrett,
Aylin Cakiroglu,
Catey Bunce,
Anoop Shah,
Spiros Denaxas
Abstract:
Large scale electronic health records (EHRs) present an opportunity to quickly identify suitable individuals in order to directly invite them to participate in an observational study. EHRs can contain data from millions of individuals, raising the question of how to optimally select a cohort of size n from a larger pool of size N. In this paper we propose a simple selective recruitment protocol th…
▽ More
Large scale electronic health records (EHRs) present an opportunity to quickly identify suitable individuals in order to directly invite them to participate in an observational study. EHRs can contain data from millions of individuals, raising the question of how to optimally select a cohort of size n from a larger pool of size N. In this paper we propose a simple selective recruitment protocol that selects a cohort in which covariates of interest tend to have a uniform distribution. We show that selectively recruited cohorts potentially offer greater statistical power and more accurate parameter estimates than randomly selected cohorts. Our protocol can be applied to studies with multiple categorical and continuous covariates. We apply our protocol to a numerically simulated prospective observational study using an EHR database of stable acute coronary disease patients from 82,089 individuals in the U.K. Selective recruitment designs require a smaller sample size, leading to more efficient and cost-effective studies.
△ Less
Submitted 13 February, 2019;
originally announced March 2019.
-
Employing latent variable models to improve efficiency in composite endpoint analysis
Authors:
Martina McMenamin,
Jessica K. Barrett,
Anna Berglind,
James M. S. Wason
Abstract:
Composite endpoints that combine multiple outcomes on different scales are common in clinical trials, particularly in chronic conditions. In many of these cases, patients will have to cross a predefined responder threshold in each of the outcomes to be classed as a responder overall. One instance of this occurs in systemic lupus erythematosus (SLE), where the responder endpoint combines two contin…
▽ More
Composite endpoints that combine multiple outcomes on different scales are common in clinical trials, particularly in chronic conditions. In many of these cases, patients will have to cross a predefined responder threshold in each of the outcomes to be classed as a responder overall. One instance of this occurs in systemic lupus erythematosus (SLE), where the responder endpoint combines two continuous, one ordinal and one binary measure. The overall binary responder endpoint is typically analysed using logistic regression, resulting in a substantial loss of information. We propose a latent variable model for the SLE endpoint, which assumes that the discrete outcomes are manifestations of latent continuous measures and can proceed to jointly model the components of the composite. We perform a simulation study and find the method to offer large efficiency gains over the standard analysis. We find that the magnitude of the precision gains are highly dependent on which components are driving response. Bias is introduced when joint normality assumptions are not satisfied, which we correct for using a bootstrap procedure. The method is applied to the Phase IIb MUSE trial in patients with moderate to severe SLE. We show that it estimates the treatment effect 2.5 times more precisely, offering a 60% reduction in required sample size.
△ Less
Submitted 19 February, 2019;
originally announced February 2019.
-
Mixed effects models for healthcare longitudinal data with an informative visiting process: a Monte Carlo simulation study
Authors:
Alessandro Gasparini,
Keith R. Abrams,
Jessica K. Barrett,
Rupert W. Major,
Michael J. Sweeting,
Nigel J. Brunskill,
Michael J. Crowther
Abstract:
Electronic health records are being increasingly used in medical research to answer more relevant and detailed clinical questions; however, they pose new and significant methodological challenges. For instance, observation times are likely correlated with the underlying disease severity: patients with worse conditions utilise health care more and may have worse biomarker values recorded. Tradition…
▽ More
Electronic health records are being increasingly used in medical research to answer more relevant and detailed clinical questions; however, they pose new and significant methodological challenges. For instance, observation times are likely correlated with the underlying disease severity: patients with worse conditions utilise health care more and may have worse biomarker values recorded. Traditional methods for analysing longitudinal data assume independence between observation times and disease severity; yet, with healthcare data such assumptions unlikely holds. Through Monte Carlo simulation, we compare different analytical approaches proposed to account for an informative visiting process to assess whether they lead to unbiased results. Furthermore, we formalise a joint model for the observation process and the longitudinal outcome within an extended joint modelling framework. We illustrate our results using data from a pragmatic trial on enhanced care for individuals with chronic kidney disease, and we introduce user-friendly software that can be used to fit the joint model for the observation process and a longitudinal outcome.
△ Less
Submitted 25 July, 2019; v1 submitted 1 August, 2018;
originally announced August 2018.
-
Estimating the association between blood pressure variability and cardiovascular disease: An application using the ARIC Study
Authors:
Jessica K. Barrett,
Raphael Huille,
Richard Parker,
Yuichiro Yano,
Michael Griswold
Abstract:
The association between visit-to-visit systolic blood pressure variability and cardiovascular events has recently received a lot of attention in the cardiovascular literature. But blood pressure variability is usually estimated on a person-by-person basis, and is therefore subject to considerable measurement error. We demonstrate that hazard ratios estimated using this approach are subject to bias…
▽ More
The association between visit-to-visit systolic blood pressure variability and cardiovascular events has recently received a lot of attention in the cardiovascular literature. But blood pressure variability is usually estimated on a person-by-person basis, and is therefore subject to considerable measurement error. We demonstrate that hazard ratios estimated using this approach are subject to bias due to regression dilution and we propose alternative methods to reduce this bias: a two-stage method and a joint model. For the two-stage method, in stage one repeated measurements are modelled using a mixed effects model with a random component on the residual standard deviation. The mixed effects model is used to estimate the blood pressure standard deviation for each individual, which in stage two is used as a covariate in a time-to-event model. For the joint model, the mixed effects sub-model and time-to-event sub-model are fitted simultaneously using shared random effects. We illustrate the methods using data from the Atherosclerosis Risk in Communities (ARIC) study.
△ Less
Submitted 23 January, 2019; v1 submitted 15 March, 2018;
originally announced March 2018.
-
Replica analysis of overfitting in regression models for time-to-event data
Authors:
ACC Coolen,
JE Barrett,
P Paga,
CJ Perez-Vicente
Abstract:
Overfitting, which happens when the number of parameters in a model is too large compared to the number of data points available for determining these parameters, is a serious and growing problem in survival analysis. While modern medicine presents us with data of unprecedented dimensionality, these data cannot yet be used effectively for clinical outcome prediction. Standard error measures in max…
▽ More
Overfitting, which happens when the number of parameters in a model is too large compared to the number of data points available for determining these parameters, is a serious and growing problem in survival analysis. While modern medicine presents us with data of unprecedented dimensionality, these data cannot yet be used effectively for clinical outcome prediction. Standard error measures in maximum likelihood regression, such as p-values and z-scores, are blind to overfitting, and even for Cox's proportional hazards model (the main tool of medical statisticians), one finds in literature only rules of thumb on the number of samples required to avoid overfitting. In this paper we present a mathematical theory of overfitting in regression models for time-to-event data, which aims to increase our quantitative understanding of the problem and provide practical tools with which to correct regression outcomes for the impact of overfitting. It is based on the replica method, a statistical mechanical technique for the analysis of heterogeneous many-variable systems that has been used successfully for several decades in physics, biology, and computer science, but not yet in medical statistics. We develop the theory initially for arbitrary regression models for time-to-event data, and verify its predictions in detail for the popular Cox model.
△ Less
Submitted 20 July, 2017; v1 submitted 4 May, 2017;
originally announced May 2017.
-
Information-adaptive clinical trials with selective recruitment and binary outcomes
Authors:
James E. Barrett
Abstract:
Selective recruitment designs preferentially recruit individuals that are estimated to be statistically informative onto a clinical trial. Individuals that are expected to contribute less information have a lower probability of recruitment. Furthermore, in an information-adaptive design recruits are allocated to treatment arms in a manner that maximises information gain. The informativeness of an…
▽ More
Selective recruitment designs preferentially recruit individuals that are estimated to be statistically informative onto a clinical trial. Individuals that are expected to contribute less information have a lower probability of recruitment. Furthermore, in an information-adaptive design recruits are allocated to treatment arms in a manner that maximises information gain. The informativeness of an individual depends on their covariate (or biomarker) values and how information is defined is a critical element of information-adaptive designs. In this paper we define and evaluate four different methods for quantifying statistical information. Using both experimental data and numerical simulations we show that selective recruitment designs can offer a substantial increase in statistical power compared to randomised designs. In trials without selective recruitment we find that allocating individuals to treatment arms according to information-adaptive protocols also leads to an increase in statistical power. Consequently, selective recruitment designs can potentially achieve successful trials using fewer recruits thereby offering economic and ethical advantages.
△ Less
Submitted 30 May, 2017; v1 submitted 3 September, 2015;
originally announced September 2015.
-
Information-adaptive clinical trials: a selective recruitment design
Authors:
James E. Barrett
Abstract:
We propose a novel adaptive design for clinical trials with time-to-event outcomes and covariates (which may consist of or include biomarkers). Our method is based on the expected entropy of the posterior distribution of a proportional hazards model. The expected entropy is evaluated as a function of a patient's covariates, and the information gained due to a patient is defined as the decrease in…
▽ More
We propose a novel adaptive design for clinical trials with time-to-event outcomes and covariates (which may consist of or include biomarkers). Our method is based on the expected entropy of the posterior distribution of a proportional hazards model. The expected entropy is evaluated as a function of a patient's covariates, and the information gained due to a patient is defined as the decrease in the corresponding entropy. Candidate patients are only recruited onto the trial if they are likely to provide sufficient information. Patients with covariates that are deemed uninformative are filtered out. A special case is where all patients are recruited, and we determine the optimal treatment arm allocation. This adaptive design has the advantage of potentially elucidating the relationship between covariates, treatments, and survival probabilities using fewer patients, albeit at the cost of rejecting some candidates. We assess the performance of our adaptive design using data from the German Breast Cancer Study group and numerical simulations of a biomarker validation trial.
△ Less
Submitted 28 March, 2016; v1 submitted 12 February, 2015;
originally announced February 2015.
-
Covariate dimension reduction for survival data via the Gaussian process latent variable model
Authors:
James E. Barrett,
Anthony C. C. Coolen
Abstract:
The analysis of high dimensional survival data is challenging, primarily due to the problem of overfitting which occurs when spurious relationships are inferred from data that subsequently fail to exist in test data. Here we propose a novel method of extracting a low dimensional representation of covariates in survival data by combining the popular Gaussian Process Latent Variable Model (GPLVM) wi…
▽ More
The analysis of high dimensional survival data is challenging, primarily due to the problem of overfitting which occurs when spurious relationships are inferred from data that subsequently fail to exist in test data. Here we propose a novel method of extracting a low dimensional representation of covariates in survival data by combining the popular Gaussian Process Latent Variable Model (GPLVM) with a Weibull Proportional Hazards Model (WPHM). The combined model offers a flexible non-linear probabilistic method of detecting and extracting any intrinsic low dimensional structure from high dimensional data. By reducing the covariate dimension we aim to diminish the risk of overfitting and increase the robustness and accuracy with which we infer relationships between covariates and survival outcomes. In addition, we can simultaneously combine information from multiple data sources by expressing multiple datasets in terms of the same low dimensional space. We present results from several simulation studies that illustrate a reduction in overfitting and an increase in predictive performance, as well as successful detection of intrinsic dimensionality. We provide evidence that it is advantageous to combine dimensionality reduction with survival outcomes rather than performing unsupervised dimensionality reduction on its own. Finally, we use our model to analyse experimental gene expression data and detect and extract a low dimensional representation that allows us to distinguish high and low risk groups with superior accuracy compared to doing regression on the original high dimensional data.
△ Less
Submitted 27 January, 2016; v1 submitted 3 June, 2014;
originally announced June 2014.
-
Gaussian process regression for survival data with competing risks
Authors:
James E. Barrett,
Anthony C. C. Coolen
Abstract:
We apply Gaussian process (GP) regression, which provides a powerful non-parametric probabilistic method of relating inputs to outputs, to survival data consisting of time-to-event and covariate measurements. In this context, the covariates are regarded as the `inputs' and the event times are the `outputs'. This allows for highly flexible inference of non-linear relationships between covariates an…
▽ More
We apply Gaussian process (GP) regression, which provides a powerful non-parametric probabilistic method of relating inputs to outputs, to survival data consisting of time-to-event and covariate measurements. In this context, the covariates are regarded as the `inputs' and the event times are the `outputs'. This allows for highly flexible inference of non-linear relationships between covariates and event times. Many existing methods, such as the ubiquitous Cox proportional hazards model, focus primarily on the hazard rate which is typically assumed to take some parametric or semi-parametric form. Our proposed model belongs to the class of accelerated failure time models where we focus on directly characterising the relationship between covariates and event times without any explicit assumptions on what form the hazard rates take. It is straightforward to include various types and combinations of censored and truncated observations. We apply our approach to both simulated and experimental data. We then apply multiple output GP regression, which can handle multiple potentially correlated outputs for each input, to competing risks survival data where multiple event types can occur. By tuning one of the model parameters we can control the extent to which the multiple outputs (the time-to-event for each risk) are dependent thus allowing the specification of correlated risks. Simulation studies suggest that in some cases assuming dependence can lead to more accurate predictions.
△ Less
Submitted 5 September, 2014; v1 submitted 5 December, 2013;
originally announced December 2013.
-
Dimensionality Detection and Integration of Multiple Data Sources via the GP-LVM
Authors:
James Barrett,
Anthony C. C. Coolen
Abstract:
The Gaussian Process Latent Variable Model (GP-LVM) is a non-linear probabilistic method of embedding a high dimensional dataset in terms low dimensional `latent' variables. In this paper we illustrate that maximum a posteriori (MAP) estimation of the latent variables and hyperparameters can be used for model selection and hence we can determine the optimal number or latent variables and the most…
▽ More
The Gaussian Process Latent Variable Model (GP-LVM) is a non-linear probabilistic method of embedding a high dimensional dataset in terms low dimensional `latent' variables. In this paper we illustrate that maximum a posteriori (MAP) estimation of the latent variables and hyperparameters can be used for model selection and hence we can determine the optimal number or latent variables and the most appropriate model. This is an alternative to the variational approaches developed recently and may be useful when we want to use a non-Gaussian prior or kernel functions that don't have automatic relevance determination (ARD) parameters. Using a second order expansion of the latent variable posterior we can marginalise the latent variables and obtain an estimate for the hyperparameter posterior. Secondly, we use the GP-LVM to integrate multiple data sources by simultaneously embedding them in terms of common latent variables. We present results from synthetic data to illustrate the successful detection and retrieval of low dimensional structure from high dimensional data. We demonstrate that the integration of multiple data sources leads to more robust performance. Finally, we show that when the data are used for binary classification tasks we can attain a significant gain in prediction accuracy when the low dimensional representation is used.
△ Less
Submitted 1 July, 2013;
originally announced July 2013.