-
SConU: Selective Conformal Uncertainty in Large Language Models
Authors:
Zhiyuan Wang,
Qingni Wang,
Yue Zhang,
Tianlong Chen,
Xiaofeng Zhu,
Xiaoshuang Shi,
Kaidi Xu
Abstract:
As large language models are increasingly utilized in real-world applications, guarantees of task-specific metrics are essential for their reliable deployment. Previous studies have introduced various criteria of conformal uncertainty grounded in split conformal prediction, which offer user-specified correctness coverage. However, existing frameworks often fail to identify uncertainty data outlier…
▽ More
As large language models are increasingly utilized in real-world applications, guarantees of task-specific metrics are essential for their reliable deployment. Previous studies have introduced various criteria of conformal uncertainty grounded in split conformal prediction, which offer user-specified correctness coverage. However, existing frameworks often fail to identify uncertainty data outliers that violate the exchangeability assumption, leading to unbounded miscoverage rates and unactionable prediction sets. In this paper, we propose a novel approach termed Selective Conformal Uncertainty (SConU), which, for the first time, implements significance tests, by developing two conformal p-values that are instrumental in determining whether a given sample deviates from the uncertainty distribution of the calibration set at a specific manageable risk level. Our approach not only facilitates rigorous management of miscoverage rates across both single-domain and interdisciplinary contexts, but also enhances the efficiency of predictions. Furthermore, we comprehensively analyze the components of the conformal procedures, aiming to approximate conditional coverage, particularly in high-stakes question-answering tasks.
△ Less
Submitted 18 April, 2025;
originally announced April 2025.
-
Analysis of Information Loss on Composition Measurement in Stiff Chemically Reacting Systems
Authors:
Yiming Lu,
Xu Zhu,
Long Zhang,
Hua Zhou
Abstract:
Gas sampling methods have been crucial for the advancement of combustion science, enabling analysis of reaction kinetics and pollutant formation. However, the measured composition can deviate from the true one because of the potential residual reactions in the sampling probes. This study formulates the initial composition estimation in stiff chemically reacting systems as a Bayesian inference prob…
▽ More
Gas sampling methods have been crucial for the advancement of combustion science, enabling analysis of reaction kinetics and pollutant formation. However, the measured composition can deviate from the true one because of the potential residual reactions in the sampling probes. This study formulates the initial composition estimation in stiff chemically reacting systems as a Bayesian inference problem, solved using the No-U-Turn Sampler (NUTS). Information loss arises from the restriction of system dynamics by low dimensional attracting manifold, where constrained evolution causes initial perturbations to decay or vanish in fast eigen-directions in composition space. This study systematically investigates the initial value inference in combustion systems and successfully validates the methodological framework in the Robertson toy system and hydrogen autoignition. Furthermore, a gas sample collected from a one-dimensional hydrogen diffusion flame is analyzed to investigate the effect of frozen temperature on information loss. The research highlights the importance of species covariance information from observations in improving estimation accuracy and identifies how the rank reduction in the sensitivity matrix leads to inference failures. Critical failure times for species inference in the Robertson and hydrogen autoignition systems are analyzed, providing insights into the limits of inference reliability and its physical significance.
△ Less
Submitted 14 March, 2025;
originally announced March 2025.
-
Quantifying Overfitting along the Regularization Path for Two-Part-Code MDL in Supervised Classification
Authors:
Xiaohan Zhu,
Nathan Srebro
Abstract:
We provide a complete characterization of the entire regularization curve of a modified two-part-code Minimum Description Length (MDL) learning rule for binary classification, based on an arbitrary prior or description language. Grunwald and Langford [2004] previously established the lack of asymptotic consistency, from an agnostic PAC (frequentist worst case) perspective, of the MDL rule with a p…
▽ More
We provide a complete characterization of the entire regularization curve of a modified two-part-code Minimum Description Length (MDL) learning rule for binary classification, based on an arbitrary prior or description language. Grunwald and Langford [2004] previously established the lack of asymptotic consistency, from an agnostic PAC (frequentist worst case) perspective, of the MDL rule with a penalty parameter of $λ=1$, suggesting that it underegularizes. Driven by interest in understanding how benign or catastrophic under-regularization and overfitting might be, we obtain a precise quantitative description of the worst case limiting error as a function of the regularization parameter $λ$ and noise level (or approximation error), significantly tightening the analysis of Grunwald and Langford for $λ=1$ and extending it to all other choices of $λ$.
△ Less
Submitted 10 March, 2025; v1 submitted 3 March, 2025;
originally announced March 2025.
-
Tight Bounds on the Binomial CDF, and the Minimum of i.i.d Binomials, in terms of KL-Divergence
Authors:
Xiaohan Zhu,
Mesrob I. Ohannessian,
Nathan Srebro
Abstract:
We provide finite sample upper and lower bounds on the Binomial tail probability which are a direct application of Sanov's theorem. We then use these to obtain high probability upper and lower bounds on the minimum of i.i.d. Binomial random variables. Both bounds are finite sample, asymptotically tight, and expressed in terms of the KL-divergence.
We provide finite sample upper and lower bounds on the Binomial tail probability which are a direct application of Sanov's theorem. We then use these to obtain high probability upper and lower bounds on the minimum of i.i.d. Binomial random variables. Both bounds are finite sample, asymptotically tight, and expressed in terms of the KL-divergence.
△ Less
Submitted 25 February, 2025;
originally announced February 2025.
-
A novel Phase I clinical trial design with unequal cohort sizes
Authors:
Xiaojun Zhu
Abstract:
This paper introduces a new Phase I design aimed at enhancing the performance of existing methods, including algorithm-based, model-based, and model-assisted designs. The design, developed by integrating the concept of Fisher information, is easily operationalized. The new design addresses the issue of the classical designs'slow dosage escalation. Simulation demonstrate that the proposed design ma…
▽ More
This paper introduces a new Phase I design aimed at enhancing the performance of existing methods, including algorithm-based, model-based, and model-assisted designs. The design, developed by integrating the concept of Fisher information, is easily operationalized. The new design addresses the issue of the classical designs'slow dosage escalation. Simulation demonstrate that the proposed design markedly enhances performance in terms of efficiency, accuracy, and reliability. Moreover, the trial duration has been notably reduced with a large sample size.
△ Less
Submitted 10 December, 2024;
originally announced December 2024.
-
Penalized Sparse Covariance Regression with High Dimensional Covariates
Authors:
Yuan Gao,
Zhiyuan Zhang,
Zhanrui Cai,
Xuening Zhu,
Tao Zou,
Hansheng Wang
Abstract:
Covariance regression offers an effective way to model the large covariance matrix with the auxiliary similarity matrices. In this work, we propose a sparse covariance regression (SCR) approach to handle the potentially high-dimensional predictors (i.e., similarity matrices). Specifically, we use the penalization method to identify the informative predictors and estimate their associated coefficie…
▽ More
Covariance regression offers an effective way to model the large covariance matrix with the auxiliary similarity matrices. In this work, we propose a sparse covariance regression (SCR) approach to handle the potentially high-dimensional predictors (i.e., similarity matrices). Specifically, we use the penalization method to identify the informative predictors and estimate their associated coefficients simultaneously. We first investigate the Lasso estimator and subsequently consider the folded concave penalized estimation methods (e.g., SCAD and MCP). However, the theoretical analysis of the existing penalization methods is primarily based on i.i.d. data, which is not directly applicable to our scenario. To address this difficulty, we establish the non-asymptotic error bounds by exploiting the spectral properties of the covariance matrix and similarity matrices. Then, we derive the estimation error bound for the Lasso estimator and establish the desirable oracle property of the folded concave penalized estimator. Extensive simulation studies are conducted to corroborate our theoretical results. We also illustrate the usefulness of the proposed method by applying it to a Chinese stock market dataset.
△ Less
Submitted 5 October, 2024;
originally announced October 2024.
-
Denoising VAE as an Explainable Feature Reduction and Diagnostic Pipeline for Autism Based on Resting state fMRI
Authors:
Xinyuan Zheng,
Orren Ravid,
Robert A. J. Barry,
Yoojean Kim,
Qian Wang,
Young-geun Kim,
Xi Zhu,
Xiaofu He
Abstract:
Autism spectrum disorders (ASDs) are developmental conditions characterized by restricted interests and difficulties in communication. The complexity of ASD has resulted in a deficiency of objective diagnostic biomarkers. Deep learning methods have gained recognition for addressing these challenges in neuroimaging analysis, but finding and interpreting such diagnostic biomarkers are still challeng…
▽ More
Autism spectrum disorders (ASDs) are developmental conditions characterized by restricted interests and difficulties in communication. The complexity of ASD has resulted in a deficiency of objective diagnostic biomarkers. Deep learning methods have gained recognition for addressing these challenges in neuroimaging analysis, but finding and interpreting such diagnostic biomarkers are still challenging computationally. Here, we propose a feature reduction pipeline using resting-state fMRI data. We used Craddock atlas and Power atlas to extract functional connectivity data from rs-fMRI, resulting in over 30 thousand features. By using a denoising variational autoencoder, our proposed pipeline further compresses the connectivity features into 5 latent Gaussian distributions, providing is a low-dimensional representation of the data to promote computational efficiency and interpretability. To test the method, we employed the extracted latent representations to classify ASD using traditional classifiers such as SVM on a large multi-site dataset. The 95% confidence interval for the prediction accuracy of SVM is [0.63, 0.76] after site harmonization using the extracted latent distributions. Without using DVAE for dimensionality reduction, the prediction accuracy is 0.70, which falls within the interval. The DVAE successfully encoded the diagnostic information from rs-fMRI data without sacrificing prediction performance. The runtime for training the DVAE and obtaining classification results from its extracted latent features was 7 times shorter compared to training classifiers directly on the raw data. Our findings suggest that the Power atlas provides more effective brain connectivity insights for diagnosing ASD than Craddock atlas. Additionally, we visualized the latent representations to gain insights into the brain networks contributing to the differences between ASD and neurotypical brains.
△ Less
Submitted 27 March, 2025; v1 submitted 30 September, 2024;
originally announced October 2024.
-
An adaptive Gaussian process method for multi-modal Bayesian inverse problems
Authors:
Zhihang Xu,
Xiaoyu Zhu,
Daoji Li,
Qifeng Liao
Abstract:
Inverse problems are prevalent in both scientific research and engineering applications. In the context of Bayesian inverse problems, sampling from the posterior distribution is particularly challenging when the forward models are computationally expensive. This challenge escalates further when the posterior distribution is multimodal. To address this, we propose a Gaussian process (GP) based meth…
▽ More
Inverse problems are prevalent in both scientific research and engineering applications. In the context of Bayesian inverse problems, sampling from the posterior distribution is particularly challenging when the forward models are computationally expensive. This challenge escalates further when the posterior distribution is multimodal. To address this, we propose a Gaussian process (GP) based method to indirectly build surrogates for the forward model. Specifically, the unnormalized posterior density is expressed as a product of an auxiliary density and an exponential GP surrogate. In an iterative way, the auxiliary density will converge to the posterior distribution starting from an arbitrary initial density. However, the efficiency of the GP regression is highly influenced by the quality of the training data. Therefore, we utilize the iterative local updating ensemble smoother (ILUES) to generate high-quality samples that are concentrated in regions with high posterior probability. Subsequently, based on the surrogate model and the mode information that is extracted by using a clustering method, MCMC with a Gaussian mixed (GM) proposal is used to draw samples from the auxiliary density. Through numerical examples, we demonstrate that the proposed method can accurately and efficiently represent the posterior with a limited number of forward simulations.
△ Less
Submitted 4 September, 2024;
originally announced September 2024.
-
Distributed quasi-Newton robust estimation under differential privacy
Authors:
Chuhan Wang,
Lixing Zhu,
Xuehu Zhu
Abstract:
For distributed computing with Byzantine machines under Privacy Protection (PP) constraints, this paper develops a robust PP distributed quasi-Newton estimation, which only requires the node machines to transmit five vectors to the central processor with high asymptotic relative efficiency. Compared with the gradient descent strategy which requires more rounds of transmission and the Newton iterat…
▽ More
For distributed computing with Byzantine machines under Privacy Protection (PP) constraints, this paper develops a robust PP distributed quasi-Newton estimation, which only requires the node machines to transmit five vectors to the central processor with high asymptotic relative efficiency. Compared with the gradient descent strategy which requires more rounds of transmission and the Newton iteration strategy which requires the entire Hessian matrix to be transmitted, the novel quasi-Newton iteration has advantages in reducing privacy budgeting and transmission cost. Moreover, our PP algorithm does not depend on the boundedness of gradients and second-order derivatives. When gradients and second-order derivatives follow sub-exponential distributions, we offer a mechanism that can ensure PP with a sufficiently high probability. Furthermore, this novel estimator can achieve the optimal convergence rate and the asymptotic normality. The numerical studies on synthetic and real data sets evaluate the performance of the proposed algorithm.
△ Less
Submitted 22 August, 2024;
originally announced August 2024.
-
A new paradigm of mortality modeling via individual vitality dynamics
Authors:
Xiaobai Zhu,
Kenneth Q. Zhou,
Zijia Wang
Abstract:
The significance of mortality modeling extends across multiple research areas, ranging from life insurance valuation to optimal lifetime decision-making. Existing approaches, such as mortality laws and factor-based models, often fall short in capturing the complexity of individual mortality, hindering their ability to address specific research needs. To overcome these limitations, this paper intro…
▽ More
The significance of mortality modeling extends across multiple research areas, ranging from life insurance valuation to optimal lifetime decision-making. Existing approaches, such as mortality laws and factor-based models, often fall short in capturing the complexity of individual mortality, hindering their ability to address specific research needs. To overcome these limitations, this paper introduces a novel approach to mortality modeling centered on the dynamics of individual vitality. A four-component framework is developed to account for initial conditions, natural aging processes, stochastic fluctuations, and accidental events over an individual's lifetime. We demonstrate the framework's analytical capabilities across various settings and explore its practical implications in solving life insurance problems and deriving optimal lifetime decisions. Our results show that the proposed framework not only encompasses existing mortality models but also provides individualized mortality outcomes and offers an intuitive explanation for survival biases.
△ Less
Submitted 21 October, 2024; v1 submitted 22 July, 2024;
originally announced July 2024.
-
Multi-relational Network Autoregression Model with Latent Group Structures
Authors:
Yimeng Ren,
Xuening Zhu,
Ganggang Xu,
Yanyuan Ma
Abstract:
Multi-relational networks among entities are frequently observed in the era of big data. Quantifying the effects of multiple networks have attracted significant research interest recently. In this work, we model multiple network effects through an autoregressive framework for tensor-valued time series. To characterize the potential heterogeneity of the networks and handle the high dimensionality o…
▽ More
Multi-relational networks among entities are frequently observed in the era of big data. Quantifying the effects of multiple networks have attracted significant research interest recently. In this work, we model multiple network effects through an autoregressive framework for tensor-valued time series. To characterize the potential heterogeneity of the networks and handle the high dimensionality of the time series data simultaneously, we assume a separate group structure for entities in each network and estimate all group memberships in a data-driven fashion. Specifically, we propose a group tensor network autoregression (GTNAR) model, which assumes that within each network, entities in the same group share the same set of model parameters, and the parameters differ across networks. An iterative algorithm is developed to estimate the model parameters and the latent group memberships simultaneously. Theoretically, we show that the group-wise parameters and group memberships can be consistently estimated when the group numbers are correctly- or possibly over-specified. An information criterion for group number estimation of each network is also provided to consistently select the group numbers. Lastly, we implement the method on a Yelp dataset to illustrate the usefulness of the method.
△ Less
Submitted 5 June, 2024;
originally announced June 2024.
-
Factor Augmented Matrix Regression
Authors:
Elynn Chen,
Jianqing Fan,
Xiaonan Zhu
Abstract:
We introduce \underline{F}actor-\underline{A}ugmented \underline{Ma}trix \underline{R}egression (FAMAR) to address the growing applications of matrix-variate data and their associated challenges, particularly with high-dimensionality and covariate correlations. FAMAR encompasses two key algorithms. The first is a novel non-iterative approach that efficiently estimates the factors and loadings of t…
▽ More
We introduce \underline{F}actor-\underline{A}ugmented \underline{Ma}trix \underline{R}egression (FAMAR) to address the growing applications of matrix-variate data and their associated challenges, particularly with high-dimensionality and covariate correlations. FAMAR encompasses two key algorithms. The first is a novel non-iterative approach that efficiently estimates the factors and loadings of the matrix factor model, utilizing techniques of pre-training, diverse projection, and block-wise averaging. The second algorithm offers an accelerated solution for penalized matrix factor regression. Both algorithms are supported by established statistical and numerical convergence properties. Empirical evaluations, conducted on synthetic and real economics datasets, demonstrate FAMAR's superiority in terms of accuracy, interpretability, and computational speed. Our application to economic data showcases how matrix factors can be incorporated to predict the GDPs of the countries of interest, and the influence of these factors on the GDPs.
△ Less
Submitted 27 May, 2024;
originally announced May 2024.
-
Bayesian Spatially Clustered Compositional Regression: Linking intersectoral GDP contributions to Gini Coefficients
Authors:
Jingcheng Meng,
Yimeng Ren,
Xuening Zhu,
Guanyu Hu
Abstract:
The Gini coefficient is an universally used measurement of income inequality. Intersectoral GDP contributions reveal the economic development of different sectors of the national economy. Linking intersectoral GDP contributions to Gini coefficients will provide better understandings of how the Gini coefficient is influenced by different industries. In this paper, a compositional regression with sp…
▽ More
The Gini coefficient is an universally used measurement of income inequality. Intersectoral GDP contributions reveal the economic development of different sectors of the national economy. Linking intersectoral GDP contributions to Gini coefficients will provide better understandings of how the Gini coefficient is influenced by different industries. In this paper, a compositional regression with spatially clustered coefficients is proposed to explore heterogeneous effects over spatial locations under nonparametric Bayesian framework. Specifically, a Markov random field constraint mixture of finite mixtures prior is designed for Bayesian log contrast regression with compostional covariates, which allows for both spatially contiguous clusters and discontinous clusters. In addition, an efficient Markov chain Monte Carlo algorithm for posterior sampling that enables simultaneous inference on both cluster configurations and cluster-wise parameters is designed. The compelling empirical performance of the proposed method is demonstrated via extensive simulation studies and an application to 51 states of United States from 2019 Bureau of Economic Analysis.
△ Less
Submitted 12 May, 2024;
originally announced May 2024.
-
Two-way Homogeneity Pursuit for Quantile Network Vector Autoregression
Authors:
Wenyang Liu,
Ganggang Xu,
Jianqing Fan,
Xuening Zhu
Abstract:
While the Vector Autoregression (VAR) model has received extensive attention for modelling complex time series, quantile VAR analysis remains relatively underexplored for high-dimensional time series data. To address this disparity, we introduce a two-way grouped network quantile (TGNQ) autoregression model for time series collected on large-scale networks, known for their significant heterogeneou…
▽ More
While the Vector Autoregression (VAR) model has received extensive attention for modelling complex time series, quantile VAR analysis remains relatively underexplored for high-dimensional time series data. To address this disparity, we introduce a two-way grouped network quantile (TGNQ) autoregression model for time series collected on large-scale networks, known for their significant heterogeneous and directional interactions among nodes. Our proposed model simultaneously conducts node clustering and model estimation to balance complexity and interpretability. To account for the directional influence among network nodes, each network node is assigned two latent group memberships that can be consistently estimated using our proposed estimation procedure. Theoretical analysis demonstrates the consistency of membership and parameter estimators even with an overspecified number of groups. With the correct group specification, estimated parameters are proven to be asymptotically normal, enabling valid statistical inferences. Moreover, we propose a quantile information criterion for consistently selecting the number of groups. Simulation studies show promising finite sample performance, and we apply the methodology to analyze connectedness and risk spillover effects among Chinese A-share stocks.
△ Less
Submitted 29 April, 2024;
originally announced April 2024.
-
A Selective Review on Statistical Methods for Massive Data Computation: Distributed Computing, Subsampling, and Minibatch Techniques
Authors:
Xuetong Li,
Yuan Gao,
Hong Chang,
Danyang Huang,
Yingying Ma,
Rui Pan,
Haobo Qi,
Feifei Wang,
Shuyuan Wu,
Ke Xu,
Jing Zhou,
Xuening Zhu,
Yingqiu Zhu,
Hansheng Wang
Abstract:
This paper presents a selective review of statistical computation methods for massive data analysis. A huge amount of statistical methods for massive data computation have been rapidly developed in the past decades. In this work, we focus on three categories of statistical computation methods: (1) distributed computing, (2) subsampling methods, and (3) minibatch gradient techniques. The first clas…
▽ More
This paper presents a selective review of statistical computation methods for massive data analysis. A huge amount of statistical methods for massive data computation have been rapidly developed in the past decades. In this work, we focus on three categories of statistical computation methods: (1) distributed computing, (2) subsampling methods, and (3) minibatch gradient techniques. The first class of literature is about distributed computing and focuses on the situation, where the dataset size is too huge to be comfortably handled by one single computer. In this case, a distributed computation system with multiple computers has to be utilized. The second class of literature is about subsampling methods and concerns about the situation, where the sample size of dataset is small enough to be placed on one single computer but too large to be easily processed by its memory as a whole. The last class of literature studies those minibatch gradient related optimization techniques, which have been extensively used for optimizing various deep learning models.
△ Less
Submitted 17 March, 2024;
originally announced March 2024.
-
Stochastic gradient descent-based inference for dynamic network models with attractors
Authors:
Hancong Pan,
Xiaojing Zhu,
Cantay Caliskan,
Dino P. Christenson,
Konstantinos Spiliopoulos,
Dylan Walker,
Eric D. Kolaczyk
Abstract:
In Coevolving Latent Space Networks with Attractors (CLSNA) models, nodes in a latent space represent social actors, and edges indicate their dynamic interactions. Attractors are added at the latent level to capture the notion of attractive and repulsive forces between nodes, borrowing from dynamical systems theory. However, CLSNA reliance on MCMC estimation makes scaling difficult, and the requir…
▽ More
In Coevolving Latent Space Networks with Attractors (CLSNA) models, nodes in a latent space represent social actors, and edges indicate their dynamic interactions. Attractors are added at the latent level to capture the notion of attractive and repulsive forces between nodes, borrowing from dynamical systems theory. However, CLSNA reliance on MCMC estimation makes scaling difficult, and the requirement for nodes to be present throughout the study period limit practical applications. We address these issues by (i) introducing a Stochastic gradient descent (SGD) parameter estimation method, (ii) developing a novel approach for uncertainty quantification using SGD, and (iii) extending the model to allow nodes to join and leave over time. Simulation results show that our extensions result in little loss of accuracy compared to MCMC, but can scale to much larger networks. We apply our approach to the longitudinal social networks of members of US Congress on the social media platform X. Accounting for node dynamics overcomes selection bias in the network and uncovers uniquely and increasingly repulsive forces within the Republican Party.
△ Less
Submitted 15 December, 2024; v1 submitted 11 March, 2024;
originally announced March 2024.
-
Uncertainty quantification for deeponets with ensemble kalman inversion
Authors:
Andrew Pensoneault,
Xueyu Zhu
Abstract:
In recent years, operator learning, particularly the DeepONet, has received much attention for efficiently learning complex mappings between input and output functions across diverse fields. However, in practical scenarios with limited and noisy data, accessing the uncertainty in DeepONet predictions becomes essential, especially in mission-critical or safety-critical applications. Existing method…
▽ More
In recent years, operator learning, particularly the DeepONet, has received much attention for efficiently learning complex mappings between input and output functions across diverse fields. However, in practical scenarios with limited and noisy data, accessing the uncertainty in DeepONet predictions becomes essential, especially in mission-critical or safety-critical applications. Existing methods, either computationally intensive or yielding unsatisfactory uncertainty quantification, leave room for developing efficient and informative uncertainty quantification (UQ) techniques tailored for DeepONets. In this work, we proposed a novel inference approach for efficient UQ for operator learning by harnessing the power of the Ensemble Kalman Inversion (EKI) approach. EKI, known for its derivative-free, noise-robust, and highly parallelizable feature, has demonstrated its advantages for UQ for physics-informed neural networks [28]. Our innovative application of EKI enables us to efficiently train ensembles of DeepONets while obtaining informative uncertainty estimates for the output of interest. We deploy a mini-batch variant of EKI to accommodate larger datasets, mitigating the computational demand due to large datasets during the training stage. Furthermore, we introduce a heuristic method to estimate the artificial dynamics covariance, thereby improving our uncertainty estimates. Finally, we demonstrate the effectiveness and versatility of our proposed methodology across various benchmark problems, showcasing its potential to address the pressing challenges of uncertainty quantification in DeepONets, especially for practical applications with limited and noisy data.
△ Less
Submitted 5 March, 2024;
originally announced March 2024.
-
Estimation of the genetic Gaussian network using GWAS summary data
Authors:
Yihe Yang,
Noah Lorincz-Comi,
Xiaofeng Zhu
Abstract:
Genetic Gaussian network of multiple phenotypes constructed through the genetic correlation matrix is informative for understanding their biological dependencies. However, its interpretation may be challenging because the estimated genetic correlations are biased due to estimation errors and horizontal pleiotropy inherent in GWAS summary statistics. Here we introduce a novel approach called Estima…
▽ More
Genetic Gaussian network of multiple phenotypes constructed through the genetic correlation matrix is informative for understanding their biological dependencies. However, its interpretation may be challenging because the estimated genetic correlations are biased due to estimation errors and horizontal pleiotropy inherent in GWAS summary statistics. Here we introduce a novel approach called Estimation of Genetic Graph (EGG), which eliminates the estimation error bias and horizontal pleiotropy bias with the same techniques used in multivariable Mendelian randomization. The genetic network estimated by EGG can be interpreted as representing shared common biological contributions between phenotypes, conditional on others, and even as indicating the causal contributions. We use both simulations and real data to demonstrate the superior efficacy of our novel method in comparison with the traditional network estimators. R package EGG is available on https://github.com/harryyiheyang/EGG.
△ Less
Submitted 15 January, 2024;
originally announced January 2024.
-
Human-in-the-loop: Towards Label Embeddings for Measuring Classification Difficulty
Authors:
Katharina Hechinger,
Christoph Koller,
Xiao Xiang Zhu,
Göran Kauermann
Abstract:
Uncertainty in machine learning models is a timely and vast field of research. In supervised learning, uncertainty can already occur in the first stage of the training process, the annotation phase. This scenario is particularly evident when some instances cannot be definitively classified. In other words, there is inevitable ambiguity in the annotation step and hence, not necessarily a "ground tr…
▽ More
Uncertainty in machine learning models is a timely and vast field of research. In supervised learning, uncertainty can already occur in the first stage of the training process, the annotation phase. This scenario is particularly evident when some instances cannot be definitively classified. In other words, there is inevitable ambiguity in the annotation step and hence, not necessarily a "ground truth" associated with each instance. The main idea of this work is to drop the assumption of a ground truth label and instead embed the annotations into a multidimensional space. This embedding is derived from the empirical distribution of annotations in a Bayesian setup, modeled via a Dirichlet-Multinomial framework. We estimate the model parameters and posteriors using a stochastic Expectation Maximization algorithm with Markov Chain Monte Carlo steps. The methods developed in this paper readily extend to various situations where multiple annotators independently label instances. To showcase the generality of the proposed approach, we apply our approach to three benchmark datasets for image classification and Natural Language Inference. Besides the embeddings, we can investigate the resulting correlation matrices, which reflect the semantic similarities of the original classes very well for all three exemplary datasets.
△ Less
Submitted 27 May, 2024; v1 submitted 15 November, 2023;
originally announced November 2023.
-
Where to serve and return in Badminton Men's Double?
Authors:
Xuelin Zhu,
Yu Sun,
Yumin Zeng,
Cong Xu
Abstract:
This study aims to analyze the service and return landing areas in badminton men's double, based on data extracted from 20 badminton matches. We find that most services land near the center-line, while returns tend to land in the crossing areas of the serving team's court. Using generalized logit models, we are able to predict the return landing area based on features of the service and return rou…
▽ More
This study aims to analyze the service and return landing areas in badminton men's double, based on data extracted from 20 badminton matches. We find that most services land near the center-line, while returns tend to land in the crossing areas of the serving team's court. Using generalized logit models, we are able to predict the return landing area based on features of the service and return round. We find that the direction of the service and the footwork and grip of the receiver could indicate his intended return landing area. Additionally, we discover that servers tend to intercept in specific areas based on their serving position. Our results offer valuable insights into the strategic decisions made by players in the service and return of a badminton rally.
△ Less
Submitted 27 October, 2023;
originally announced October 2023.
-
Categorising the World into Local Climate Zones -- Towards Quantifying Labelling Uncertainty for Machine Learning Models
Authors:
Katharina Hechinger,
Xiao Xiang Zhu,
Göran Kauermann
Abstract:
Image classification is often prone to labelling uncertainty. To generate suitable training data, images are labelled according to evaluations of human experts. This can result in ambiguities, which will affect subsequent models. In this work, we aim to model the labelling uncertainty in the context of remote sensing and the classification of satellite images. We construct a multinomial mixture mo…
▽ More
Image classification is often prone to labelling uncertainty. To generate suitable training data, images are labelled according to evaluations of human experts. This can result in ambiguities, which will affect subsequent models. In this work, we aim to model the labelling uncertainty in the context of remote sensing and the classification of satellite images. We construct a multinomial mixture model given the evaluations of multiple experts. This is based on the assumption that there is no ambiguity of the image class, but apparently in the experts' opinion about it. The model parameters can be estimated by a stochastic Expectation Maximization algorithm. Analysing the estimates gives insights into sources of label uncertainty. Here, we focus on the general class ambiguity, the heterogeneity of experts, and the origin city of the images. The results are relevant for all machine learning applications where image classification is pursued and labelling is subject to humans.
△ Less
Submitted 4 September, 2023;
originally announced September 2023.
-
Model Selection for Exposure-Mediator Interaction
Authors:
Ruiyang Li,
Xi Zhu,
Seonjoo Lee
Abstract:
In mediation analysis, the exposure often influences the mediating effect, i.e., there is an interaction between exposure and mediator on the dependent variable. When the mediator is high-dimensional, it is necessary to identify non-zero mediators (M) and exposure-by-mediator (X-by-M) interactions. Although several high-dimensional mediation methods can naturally handle X-by-M interactions, resear…
▽ More
In mediation analysis, the exposure often influences the mediating effect, i.e., there is an interaction between exposure and mediator on the dependent variable. When the mediator is high-dimensional, it is necessary to identify non-zero mediators (M) and exposure-by-mediator (X-by-M) interactions. Although several high-dimensional mediation methods can naturally handle X-by-M interactions, research is scarce in preserving the underlying hierarchical structure between the main effects and the interactions. To fill the knowledge gap, we develop the XMInt procedure to select M and X-by-M interactions in the high-dimensional mediators setting while preserving the hierarchical structure. Our proposed method employs a sequential regularization-based forward-selection approach to identify the mediators and their hierarchically preserved interaction with exposure. Our numerical experiments showed promising selection results. Further, we applied our method to ADNI morphological data and examined the role of cortical thickness and subcortical volumes on the effect of amyloid-beta accumulation on cognitive performance, which could be helpful in understanding the brain compensation mechanism.
△ Less
Submitted 2 August, 2023;
originally announced August 2023.
-
Corrected kernel principal component analysis for model structural change detection
Authors:
Luoyao Yu,
Lixing Zhu,
Ruoqing Zhu,
Xuehu Zhu
Abstract:
This paper develops a method to detect model structural changes by applying a Corrected Kernel Principal Component Analysis (CKPCA) to construct the so-called central distribution deviation subspaces. This approach can efficiently identify the mean and distribution changes in these dimension reduction subspaces. We derive that the locations and number changes in the dimension reduction data subspa…
▽ More
This paper develops a method to detect model structural changes by applying a Corrected Kernel Principal Component Analysis (CKPCA) to construct the so-called central distribution deviation subspaces. This approach can efficiently identify the mean and distribution changes in these dimension reduction subspaces. We derive that the locations and number changes in the dimension reduction data subspaces are identical to those in the original data spaces. Meanwhile, we also explain the necessity of using CKPCA as the classical KPCA fails to identify the central distribution deviation subspaces in these problems. Additionally, we extend this approach to clustering by embedding the original data with nonlinear lower dimensional spaces, providing enhanced capabilities for clustering analysis. The numerical studies on synthetic and real data sets suggest that the dimension reduction versions of existing methods for change point detection and clustering significantly improve the performances of existing approaches in finite sample scenarios.
△ Less
Submitted 15 July, 2023;
originally announced July 2023.
-
Sparsified Simultaneous Confidence Intervals for High-Dimensional Linear Models
Authors:
Xiaorui Zhu,
Yichen Qin,
Peng Wang
Abstract:
Statistical inference of the high-dimensional regression coefficients is challenging because the uncertainty introduced by the model selection procedure is hard to account for. A critical question remains unsettled; that is, is it possible and how to embed the inference of the model into the simultaneous inference of the coefficients? To this end, we propose a notion of simultaneous confidence int…
▽ More
Statistical inference of the high-dimensional regression coefficients is challenging because the uncertainty introduced by the model selection procedure is hard to account for. A critical question remains unsettled; that is, is it possible and how to embed the inference of the model into the simultaneous inference of the coefficients? To this end, we propose a notion of simultaneous confidence intervals called the sparsified simultaneous confidence intervals. Our intervals are sparse in the sense that some of the intervals' upper and lower bounds are shrunken to zero (i.e., $[0,0]$), indicating the unimportance of the corresponding covariates. These covariates should be excluded from the final model. The rest of the intervals, either containing zero (e.g., $[-1,1]$ or $[0,1]$) or not containing zero (e.g., $[2,3]$), indicate the plausible and significant covariates, respectively. The proposed method can be coupled with various selection procedures, making it ideal for comparing their uncertainty. For the proposed method, we establish desirable asymptotic properties, develop intuitive graphical tools for visualization, and justify its superior performance through simulation and real data analysis.
△ Less
Submitted 2 January, 2025; v1 submitted 14 July, 2023;
originally announced July 2023.
-
Biomass Estimation and Uncertainty Quantification from Tree Height
Authors:
Qian Song,
Conrad M Albrecht,
Zhitong Xiong,
Xiao Xiang Zhu
Abstract:
We propose a tree-level biomass estimation model approximating allometric equations by LiDAR data. Since tree crown diameters estimation is challenging from spaceborne LiDAR measurements, we develop a model to correlate tree height with biomass on the individual tree level employing a Gaussian process regressor. In order to validate the proposed model, a set of 8,342 samples on tree height, trunk…
▽ More
We propose a tree-level biomass estimation model approximating allometric equations by LiDAR data. Since tree crown diameters estimation is challenging from spaceborne LiDAR measurements, we develop a model to correlate tree height with biomass on the individual tree level employing a Gaussian process regressor. In order to validate the proposed model, a set of 8,342 samples on tree height, trunk diameter, and biomass has been assembled. It covers seven biomes globally present. We reference our model to four other models based on both, the Jucker data and our own dataset. Although our approach deviates from standard biomass-height-diameter models, we demonstrate the Gaussian process regression model as a viable alternative. In addition, we decompose the uncertainty of tree biomass estimates into the model- and fitting-based contributions. We verify the Gaussian process regressor has the capacity to reduce the fitting uncertainty down to below 5%. Exploiting airborne LiDAR measurements and a field inventory survey on the ground, a stand-level (or plot-level) study confirms a low relative error of below 1% for our model. The data used in this study are available at https://github.com/zhu-xlab/BiomassUQ .
△ Less
Submitted 17 May, 2023; v1 submitted 16 May, 2023;
originally announced May 2023.
-
Artificial intelligence to advance Earth observation: : A review of models, recent trends, and pathways forward
Authors:
Devis Tuia,
Konrad Schindler,
Begüm Demir,
Xiao Xiang Zhu,
Mrinalini Kochupillai,
Sašo Džeroski,
Jan N. van Rijn,
Holger H. Hoos,
Fabio Del Frate,
Mihai Datcu,
Volker Markl,
Bertrand Le Saux,
Rochelle Schneider,
Gustau Camps-Valls
Abstract:
Earth observation (EO) is a prime instrument for monitoring land and ocean processes, studying the dynamics at work, and taking the pulse of our planet. This article gives a bird's eye view of the essential scientific tools and approaches informing and supporting the transition from raw EO data to usable EO-based information. The promises, as well as the current challenges of these developments, a…
▽ More
Earth observation (EO) is a prime instrument for monitoring land and ocean processes, studying the dynamics at work, and taking the pulse of our planet. This article gives a bird's eye view of the essential scientific tools and approaches informing and supporting the transition from raw EO data to usable EO-based information. The promises, as well as the current challenges of these developments, are highlighted under dedicated sections. Specifically, we cover the impact of (i) Computer vision; (ii) Machine learning; (iii) Advanced processing and computing; (iv) Knowledge-based AI; (v) Explainable AI and causal inference; (vi) Physics-aware models; (vii) User-centric approaches; and (viii) the much-needed discussion of ethical and societal issues related to the massive use of ML technologies in EO.
△ Less
Submitted 16 September, 2024; v1 submitted 15 May, 2023;
originally announced May 2023.
-
The Politics of Language Choice: How the Russian-Ukrainian War Influences Ukrainians' Language Use on Twitter
Authors:
Daniel Racek,
Brittany I. Davidson,
Paul W. Thurner,
Xiao Xiang Zhu,
Göran Kauermann
Abstract:
The use of language is innately political and often a vehicle of cultural identity as well as the basis for nation building. Here, we examine language choice and tweeting activity of Ukrainian citizens based on more than 4 million geo-tagged tweets from over 62,000 users before and during the Russian-Ukrainian War, from January 2020 to October 2022. Using statistical models, we disentangle sample…
▽ More
The use of language is innately political and often a vehicle of cultural identity as well as the basis for nation building. Here, we examine language choice and tweeting activity of Ukrainian citizens based on more than 4 million geo-tagged tweets from over 62,000 users before and during the Russian-Ukrainian War, from January 2020 to October 2022. Using statistical models, we disentangle sample effects, arising from the in- and outflux of users on Twitter, from behavioural effects, arising from behavioural changes of the users. We observe a steady shift from the Russian language towards the Ukrainian language already before the war, which drastically speeds up with its outbreak. We attribute these shifts in large part to users' behavioural changes. Notably, we find that more than half of the Russian-tweeting users shift towards Ukrainian as a result of the war.
△ Less
Submitted 6 June, 2023; v1 submitted 4 May, 2023;
originally announced May 2023.
-
Improved Naive Bayes with Mislabeled Data
Authors:
Qianhan Zeng,
Yingqiu Zhu,
Xuening Zhu,
Feifei Wang,
Weichen Zhao,
Shuning Sun,
Meng Su,
Hansheng Wang
Abstract:
Labeling mistakes are frequently encountered in real-world applications. If not treated well, the labeling mistakes can deteriorate the classification performances of a model seriously. To address this issue, we propose an improved Naive Bayes method for text classification. It is analytically simple and free of subjective judgements on the correct and incorrect labels. By specifying the generatin…
▽ More
Labeling mistakes are frequently encountered in real-world applications. If not treated well, the labeling mistakes can deteriorate the classification performances of a model seriously. To address this issue, we propose an improved Naive Bayes method for text classification. It is analytically simple and free of subjective judgements on the correct and incorrect labels. By specifying the generating mechanism of incorrect labels, we optimize the corresponding log-likelihood function iteratively by using an EM algorithm. Our simulation and experiment results show that the improved Naive Bayes method greatly improves the performances of the Naive Bayes method with mislabeled data.
△ Less
Submitted 13 April, 2023;
originally announced April 2023.
-
Subsampling and Jackknifing: A Practically Convenient Solution for Large Data Analysis with Limited Computational Resources
Authors:
Shuyuan Wu,
Xuening Zhu,
Hansheng Wang
Abstract:
Modern statistical analysis often encounters datasets with large sizes. For these datasets, conventional estimation methods can hardly be used immediately because practitioners often suffer from limited computational resources. In most cases, they do not have powerful computational resources (e.g., Hadoop or Spark). How to practically analyze large datasets with limited computational resources the…
▽ More
Modern statistical analysis often encounters datasets with large sizes. For these datasets, conventional estimation methods can hardly be used immediately because practitioners often suffer from limited computational resources. In most cases, they do not have powerful computational resources (e.g., Hadoop or Spark). How to practically analyze large datasets with limited computational resources then becomes a problem of great importance. To solve this problem, we propose here a novel subsampling-based method with jackknifing. The key idea is to treat the whole sample data as if they were the population. Then, multiple subsamples with greatly reduced sizes are obtained by the method of simple random sampling with replacement. It is remarkable that we do not recommend sampling methods without replacement because this would incur a significant cost for data processing on the hard drive. Such cost does not exist if the data are processed in memory. Because subsampled data have relatively small sizes, they can be comfortably read into computer memory as a whole and then processed easily. Based on subsampled datasets, jackknife-debiased estimators can be obtained for the target parameter. The resulting estimators are statistically consistent, with an extremely small bias. Finally, the jackknife-debiased estimators from different subsamples are averaged together to form the final estimator. We theoretically show that the final estimator is consistent and asymptotically normal. Its asymptotic statistical efficiency can be as good as that of the whole sample estimator under very mild conditions. The proposed method is simple enough to be easily implemented on most practical computer systems and thus should have very wide applicability.
△ Less
Submitted 12 April, 2023;
originally announced April 2023.
-
Distributed Logistic Regression for Massive Data with Rare Events
Authors:
Xuetong Li,
Xuening Zhu,
Hansheng Wang
Abstract:
Large-scale rare events data are commonly encountered in practice. To tackle the massive rare events data, we propose a novel distributed estimation method for logistic regression in a distributed system. For a distributed framework, we face the following two challenges. The first challenge is how to distribute the data. In this regard, two different distribution strategies (i.e., the RANDOM strat…
▽ More
Large-scale rare events data are commonly encountered in practice. To tackle the massive rare events data, we propose a novel distributed estimation method for logistic regression in a distributed system. For a distributed framework, we face the following two challenges. The first challenge is how to distribute the data. In this regard, two different distribution strategies (i.e., the RANDOM strategy and the COPY strategy) are investigated. The second challenge is how to select an appropriate type of objective function so that the best asymptotic efficiency can be achieved. Then, the under-sampled (US) and inverse probability weighted (IPW) types of objective functions are considered. Our results suggest that the COPY strategy together with the IPW objective function is the best solution for distributed logistic regression with rare events. The finite sample performance of the distributed methods is demonstrated by simulation studies and a real-world Sweden Traffic Sign dataset.
△ Less
Submitted 5 April, 2023;
originally announced April 2023.
-
Efficient Bayesian Physics Informed Neural Networks for Inverse Problems via Ensemble Kalman Inversion
Authors:
Andrew Pensoneault,
Xueyu Zhu
Abstract:
Bayesian Physics Informed Neural Networks (B-PINNs) have gained significant attention for inferring physical parameters and learning the forward solutions for problems based on partial differential equations. However, the overparameterized nature of neural networks poses a computational challenge for high-dimensional posterior inference. Existing inference approaches, such as particle-based or var…
▽ More
Bayesian Physics Informed Neural Networks (B-PINNs) have gained significant attention for inferring physical parameters and learning the forward solutions for problems based on partial differential equations. However, the overparameterized nature of neural networks poses a computational challenge for high-dimensional posterior inference. Existing inference approaches, such as particle-based or variance inference methods, are either computationally expensive for high-dimensional posterior inference or provide unsatisfactory uncertainty estimates. In this paper, we present a new efficient inference algorithm for B-PINNs that uses Ensemble Kalman Inversion (EKI) for high-dimensional inference tasks. We find that our proposed method can achieve inference results with informative uncertainty estimates comparable to Hamiltonian Monte Carlo (HMC)-based B-PINNs with a much reduced computational cost. These findings suggest that our proposed approach has great potential for uncertainty quantification in physics-informed machine learning for practical applications.
△ Less
Submitted 13 March, 2023;
originally announced March 2023.
-
A General Theory of Correct, Incorrect, and Extrinsic Equivariance
Authors:
Dian Wang,
Xupeng Zhu,
Jung Yeon Park,
Mingxi Jia,
Guanang Su,
Robert Platt,
Robin Walters
Abstract:
Although equivariant machine learning has proven effective at many tasks, success depends heavily on the assumption that the ground truth function is symmetric over the entire domain matching the symmetry in an equivariant neural network. A missing piece in the equivariant learning literature is the analysis of equivariant networks when symmetry exists only partially in the domain. In this work, w…
▽ More
Although equivariant machine learning has proven effective at many tasks, success depends heavily on the assumption that the ground truth function is symmetric over the entire domain matching the symmetry in an equivariant neural network. A missing piece in the equivariant learning literature is the analysis of equivariant networks when symmetry exists only partially in the domain. In this work, we present a general theory for such a situation. We propose pointwise definitions of correct, incorrect, and extrinsic equivariance, which allow us to quantify continuously the degree of each type of equivariance a function displays. We then study the impact of various degrees of incorrect or extrinsic symmetry on model error. We prove error lower bounds for invariant or equivariant networks in classification or regression settings with partially incorrect symmetry. We also analyze the potentially harmful effects of extrinsic equivariance. Experiments validate these results in three different environments.
△ Less
Submitted 28 October, 2023; v1 submitted 8 March, 2023;
originally announced March 2023.
-
Network Autoregression for Incomplete Matrix-Valued Time Series
Authors:
Xuening Zhu,
Feifei Wang,
Zeng Li,
Yanyuan Ma
Abstract:
We study the dynamics of matrix-valued time series with observed network structures by proposing a matrix network autoregression model with row and column networks of the subjects. We incorporate covariate information and a low rank intercept matrix. We allow incomplete observations in the matrices and the missing mechanism can be covariate dependent. To estimate the model, a two-step estimation p…
▽ More
We study the dynamics of matrix-valued time series with observed network structures by proposing a matrix network autoregression model with row and column networks of the subjects. We incorporate covariate information and a low rank intercept matrix. We allow incomplete observations in the matrices and the missing mechanism can be covariate dependent. To estimate the model, a two-step estimation procedure is proposed. The first step aims to estimate the network autoregression coefficients, and the second step aims to estimate the regression parameters, which are matrices themselves. Theoretically, we first separately establish the asymptotic properties of the autoregression coefficients and the error bounds of the regression parameters. Subsequently, a bias reduction procedure is proposed to reduce the asymptotic bias and the theoretical property of the debiased estimator is studied. Lastly, we illustrate the usefulness of the proposed method through a number of numerical studies and an analysis of a Yelp data set.
△ Less
Submitted 6 February, 2023;
originally announced February 2023.
-
Unbiased estimation and asymptotically valid inference in multivariable Mendelian randomization with many weak instrumental variables
Authors:
Yihe Yang,
Noah Lorincz-Comi,
Xiaofeng Zhu
Abstract:
Mendelian randomization (MR) is an instrumental variable (IV) approach to infer causal relationships between exposures and outcomes with genome-wide association studies (GWAS) summary data. However, the multivariable inverse-variance weighting (IVW) approach, which serves as the foundation for most MR approaches, cannot yield unbiased causal effect estimates in the presence of many weak IVs. To ad…
▽ More
Mendelian randomization (MR) is an instrumental variable (IV) approach to infer causal relationships between exposures and outcomes with genome-wide association studies (GWAS) summary data. However, the multivariable inverse-variance weighting (IVW) approach, which serves as the foundation for most MR approaches, cannot yield unbiased causal effect estimates in the presence of many weak IVs. To address this problem, we proposed the MR using Bias-corrected Estimating Equation (MRBEE) that can infer unbiased causal relationships with many weak IVs and account for horizontal pleiotropy simultaneously. While the practical significance of MRBEE was demonstrated in our parallel work (Lorincz-Comi (2023)), this paper established the statistical theories of multivariable IVW and MRBEE with many weak IVs. First, we showed that the bias of the multivariable IVW estimate is caused by the error-in-variable bias, whose scale and direction are inflated and influenced by weak instrument bias and sample overlaps of exposures and outcome GWAS cohorts, respectively. Second, we investigated the asymptotic properties of multivariable IVW and MRBEE, showing that MRBEE outperforms multivariable IVW regarding unbiasedness of causal effect estimation and asymptotic validity of causal inference. Finally, we applied MRBEE to examine myopia and revealed that education and outdoor activity are causal to myopia whereas indoor activity is not.
△ Less
Submitted 10 February, 2024; v1 submitted 12 January, 2023;
originally announced January 2023.
-
Matrix-valued Network Autoregression Model with Latent Group Structure
Authors:
Yimeng Ren,
Xuening Zhu,
Yanyuan Ma
Abstract:
Matrix-valued time series data are frequently observed in a broad range of areas and have attracted great attention recently. In this work, we model network effects for high dimensional matrix-valued time series data in a matrix autoregression framework. To characterize the potential heterogeneity of the subjects and handle the high dimensionality simultaneously, we assume that each subject has a…
▽ More
Matrix-valued time series data are frequently observed in a broad range of areas and have attracted great attention recently. In this work, we model network effects for high dimensional matrix-valued time series data in a matrix autoregression framework. To characterize the potential heterogeneity of the subjects and handle the high dimensionality simultaneously, we assume that each subject has a latent group label, which enables us to cluster the subject into the corresponding row and column groups. We propose a group matrix network autoregression (GMNAR) model, which assumes that the subjects in the same group share the same set of model parameters. To estimate the model, we develop an iterative algorithm. Theoretically, we show that the group-wise parameters and group memberships can be consistently estimated when the group numbers are correctly or possibly over-specified. An information criterion for group number estimation is also provided to consistently select the group numbers. Lastly, we implement the method on a Yelp dataset to illustrate the usefulness of the method.
△ Less
Submitted 5 December, 2022;
originally announced December 2022.
-
Online Linearized LASSO
Authors:
Shuoguang Yang,
Yuhao Yan,
Xiuneng Zhu,
Qiang Sun
Abstract:
Sparse regression has been a popular approach to perform variable selection and enhance the prediction accuracy and interpretability of the resulting statistical model. Existing approaches focus on offline regularized regression, while the online scenario has rarely been studied. In this paper, we propose a novel online sparse linear regression framework for analyzing streaming data when data poin…
▽ More
Sparse regression has been a popular approach to perform variable selection and enhance the prediction accuracy and interpretability of the resulting statistical model. Existing approaches focus on offline regularized regression, while the online scenario has rarely been studied. In this paper, we propose a novel online sparse linear regression framework for analyzing streaming data when data points arrive sequentially. Our proposed method is memory efficient and requires less stringent restricted strong convexity assumptions. Theoretically, we show that with a properly chosen regularization parameter, the $\ell_2$-norm statistical error of our estimator diminishes to zero in the optimal order of $\tilde{O}({\sqrt{s/t}})$, where $s$ is the sparsity level, $t$ is the streaming sample size, and $\tilde{O}(\cdot)$ hides logarithmic terms. Numerical experiments demonstrate the practical efficiency of our algorithm.
△ Less
Submitted 1 January, 2023; v1 submitted 11 November, 2022;
originally announced November 2022.
-
Distributed Estimation and Inference for Spatial Autoregression Model with Large Scale Networks
Authors:
Yimeng Ren,
Zhe Li,
Xuening Zhu,
Yuan Gao,
Hansheng Wang
Abstract:
The rapid growth of online network platforms generates large-scale network data and it poses great challenges for statistical analysis using the spatial autoregression (SAR) model. In this work, we develop a novel distributed estimation and statistical inference framework for the SAR model on a distributed system. We first propose a distributed network least squares approximation (DNLSA) method. T…
▽ More
The rapid growth of online network platforms generates large-scale network data and it poses great challenges for statistical analysis using the spatial autoregression (SAR) model. In this work, we develop a novel distributed estimation and statistical inference framework for the SAR model on a distributed system. We first propose a distributed network least squares approximation (DNLSA) method. This enables us to obtain a one-step estimator by taking a weighted average of local estimators on each worker. Afterwards, a refined two-step estimation is designed to further reduce the estimation bias. For statistical inference, we utilize a random projection method to reduce the expensive communication cost. Theoretically, we show the consistency and asymptotic normality of both the one-step and two-step estimators. In addition, we provide theoretical guarantee of the distributed statistical inference procedure. The theoretical findings and computational advantages are validated by several numerical simulations implemented on the Spark system. Lastly, an experiment on the Yelp dataset further illustrates the usefulness of the proposed methodology.
△ Less
Submitted 27 November, 2023; v1 submitted 29 October, 2022;
originally announced October 2022.
-
Understanding Edge-of-Stability Training Dynamics with a Minimalist Example
Authors:
Xingyu Zhu,
Zixuan Wang,
Xiang Wang,
Mo Zhou,
Rong Ge
Abstract:
Recently, researchers observed that gradient descent for deep neural networks operates in an ``edge-of-stability'' (EoS) regime: the sharpness (maximum eigenvalue of the Hessian) is often larger than stability threshold $2/η$ (where $η$ is the step size). Despite this, the loss oscillates and converges in the long run, and the sharpness at the end is just slightly below $2/η$. While many other wel…
▽ More
Recently, researchers observed that gradient descent for deep neural networks operates in an ``edge-of-stability'' (EoS) regime: the sharpness (maximum eigenvalue of the Hessian) is often larger than stability threshold $2/η$ (where $η$ is the step size). Despite this, the loss oscillates and converges in the long run, and the sharpness at the end is just slightly below $2/η$. While many other well-understood nonconvex objectives such as matrix factorization or two-layer networks can also converge despite large sharpness, there is often a larger gap between sharpness of the endpoint and $2/η$. In this paper, we study EoS phenomenon by constructing a simple function that has the same behavior. We give rigorous analysis for its training dynamics in a large local region and explain why the final converging point has sharpness close to $2/η$. Globally we observe that the training dynamics for our example has an interesting bifurcating behavior, which was also observed in the training of neural nets.
△ Less
Submitted 21 February, 2023; v1 submitted 6 October, 2022;
originally announced October 2022.
-
Simultaneous Estimation and Group Identification for Network Vector Autoregressive Model with Heterogeneous Nodes
Authors:
Xuening Zhu,
Ganggang Xu,
Jianqing Fan
Abstract:
Individuals or companies in a large social or financial network often display rather heterogeneous behaviors for various reasons. In this work, we propose a network vector autoregressive model with a latent group structure to model heterogeneous dynamic patterns observed from network nodes, for which group-wise network effects and timeinvariant fixed-effects can be naturally incorporated. In our f…
▽ More
Individuals or companies in a large social or financial network often display rather heterogeneous behaviors for various reasons. In this work, we propose a network vector autoregressive model with a latent group structure to model heterogeneous dynamic patterns observed from network nodes, for which group-wise network effects and timeinvariant fixed-effects can be naturally incorporated. In our framework, the model parameters and network node memberships can be simultaneously estimated by minimizing a least-squares type objective function. In particular, our theoretical investigation allows the number of latent groups G to be over-specified when achieving the estimation consistency of the model parameters and group memberships, which significantly improves the robustness of the proposed approach. When G is correctly specified, valid statistical inference can be made for model parameters based on the asymptotic normality of the estimators. A data-driven criterion is developed to consistently identify the true group number for practical use. Extensive simulation studies and two real data examples are used to demonstrate the effectiveness of the proposed methodology.
△ Less
Submitted 11 August, 2023; v1 submitted 25 September, 2022;
originally announced September 2022.
-
Consistent Selection of the Number of Groups in Panel Models via Cross-Validation
Authors:
Zhe Li,
Xuening Zhu,
Changliang Zou
Abstract:
Group number selection is a key problem for group panel data modeling. In this work, we develop a cross-validation (CV) method to tackle this problem. Specifically, we split the panel data into two data folds on the time span, with group structure preserved for individuals. We first estimate the group memberships and parameters on one data fold, then we plug in the estimates and utilize the other…
▽ More
Group number selection is a key problem for group panel data modeling. In this work, we develop a cross-validation (CV) method to tackle this problem. Specifically, we split the panel data into two data folds on the time span, with group structure preserved for individuals. We first estimate the group memberships and parameters on one data fold, then we plug in the estimates and utilize the other data fold to evaluate a designed criterion. Subsequently, the group number is estimated by minimizing the average criterion across all data folds. The proposed CV method has two advantages compared to existing approaches. First, the method is totally data-driven, thus no further tuning parameters are involved. Second, the method can be flexibly applied to a wide range of panel data models. Theoretically, we establish the estimation consistency by taking advantage of the optimization property of the estimation algorithm. Experiments are carried out with a variety of synthetic datasets and panel models to further illustrate the advantages of the proposed method. Lastly, the CV method is employed to analyze the heterogeneous patterns of stock volatilities in the Chinese stock market through the financial crisis.
△ Less
Submitted 16 May, 2025; v1 submitted 12 September, 2022;
originally announced September 2022.
-
Seismic fragility analysis using stochastic polynomial chaos expansions
Authors:
X. Zhu,
M. Broccardo,
B. Sudret
Abstract:
Within the performance-based earthquake engineering (PBEE) framework, the fragility model plays a pivotal role. Such a model represents the probability that the engineering demand parameter (EDP) exceeds a certain safety threshold given a set of selected intensity measures (IMs) that characterize the earthquake load. The-state-of-the art methods for fragility computation rely on full non-linear ti…
▽ More
Within the performance-based earthquake engineering (PBEE) framework, the fragility model plays a pivotal role. Such a model represents the probability that the engineering demand parameter (EDP) exceeds a certain safety threshold given a set of selected intensity measures (IMs) that characterize the earthquake load. The-state-of-the art methods for fragility computation rely on full non-linear time-history analyses. Within this perimeter, there are two main approaches: the first relies on the selection and scaling of recorded ground motions; the second, based on random vibration theory, characterizes the seismic input with a parametric stochastic ground motion model (SGMM). The latter case has the great advantage that the problem of seismic risk analysis is framed as a forward uncertainty quantification problem. However, running classical full-scale Monte Carlo simulations is intractable because of the prohibitive computational cost of typical finite element models. Therefore, it is of great interest to define fragility models that link an EDP of interest with the SGMM parameters -- which are regarded as IMs in this context. The computation of such fragility models is a challenge on its own and, despite few recent studies, there is still an important research gap in this domain. This study tackles this computational challenge by using stochastic polynomial chaos expansions to represent the statistical dependence of EDP on IMs. More precisely, this surrogate model estimates the full conditional probability distribution of EDP conditioned on IMs. We compare the proposed approach with some state-of-the-art methods in two case studies. The numerical results show that the new method prevails its competitors in estimating both the conditional distribution and the fragility functions.
△ Less
Submitted 1 February, 2023; v1 submitted 16 August, 2022;
originally announced August 2022.
-
Multiple change point detection in tensors
Authors:
Jiaqi Huang,
Junhui Wang,
Xuehu Zhu,
Lixing Zhu
Abstract:
This paper proposes a criterion for detecting change structures in tensor data. To accommodate tensor structure with structural mode that is not suitable to be equally treated and summarized in a distance to measure the difference between any two adjacent tensors, we define a mode-based signal-screening Frobenius distance for the moving sums of slices of tensor data to handle both dense and sparse…
▽ More
This paper proposes a criterion for detecting change structures in tensor data. To accommodate tensor structure with structural mode that is not suitable to be equally treated and summarized in a distance to measure the difference between any two adjacent tensors, we define a mode-based signal-screening Frobenius distance for the moving sums of slices of tensor data to handle both dense and sparse model structures of the tensors. As a general distance, it can also deal with the case without structural mode. Based on the distance, we then construct signal statistics using the ratios with adaptive-to-change ridge functions. The number of changes and their locations can then be consistently estimated in certain senses, and the confidence intervals of the locations of change points are constructed. The results hold when the size of the tensor and the number of change points diverge at certain rates, respectively. Numerical studies are conducted to examine the finite sample performances of the proposed method. We also analyze two real data examples for illustration.
△ Less
Submitted 18 March, 2023; v1 submitted 26 June, 2022;
originally announced June 2022.
-
Byzantine-Robust Online and Offline Distributed Reinforcement Learning
Authors:
Yiding Chen,
Xuezhou Zhang,
Kaiqing Zhang,
Mengdi Wang,
Xiaojin Zhu
Abstract:
We consider a distributed reinforcement learning setting where multiple agents separately explore the environment and communicate their experiences through a central server. However, $α$-fraction of agents are adversarial and can report arbitrary fake information. Critically, these adversarial agents can collude and their fake data can be of any sizes. We desire to robustly identify a near-optimal…
▽ More
We consider a distributed reinforcement learning setting where multiple agents separately explore the environment and communicate their experiences through a central server. However, $α$-fraction of agents are adversarial and can report arbitrary fake information. Critically, these adversarial agents can collude and their fake data can be of any sizes. We desire to robustly identify a near-optimal policy for the underlying Markov decision process in the presence of these adversarial agents. Our main technical contribution is Weighted-Clique, a novel algorithm for the robust mean estimation from batches problem, that can handle arbitrary batch sizes. Building upon this new estimator, in the offline setting, we design a Byzantine-robust distributed pessimistic value iteration algorithm; in the online setting, we design a Byzantine-robust distributed optimistic value iteration algorithm. Both algorithms obtain near-optimal sample complexities and achieve superior robustness guarantee than prior works.
△ Less
Submitted 31 May, 2022;
originally announced June 2022.
-
So2Sat POP -- A Curated Benchmark Data Set for Population Estimation from Space on a Continental Scale
Authors:
Sugandha Doda,
Yuanyuan Wang,
Matthias Kahl,
Eike Jens Hoffmann,
Kim Ouan,
Hannes Taubenböck,
Xiao Xiang Zhu
Abstract:
Obtaining a dynamic population distribution is key to many decision-making processes such as urban planning, disaster management and most importantly helping the government to better allocate socio-technical supply. For the aspiration of these objectives, good population data is essential. The traditional method of collecting population data through the census is expensive and tedious. In recent y…
▽ More
Obtaining a dynamic population distribution is key to many decision-making processes such as urban planning, disaster management and most importantly helping the government to better allocate socio-technical supply. For the aspiration of these objectives, good population data is essential. The traditional method of collecting population data through the census is expensive and tedious. In recent years, statistical and machine learning methods have been developed to estimate population distribution. Most of the methods use data sets that are either developed on a small scale or not publicly available yet. Thus, the development and evaluation of new methods become challenging. We fill this gap by providing a comprehensive data set for population estimation in 98 European cities. The data set comprises a digital elevation model, local climate zone, land use proportions, nighttime lights in combination with multi-spectral Sentinel-2 imagery, and data from the Open Street Map initiative. We anticipate that it would be a valuable addition to the research community for the development of sophisticated approaches in the field of population estimation.
△ Less
Submitted 10 November, 2022; v1 submitted 7 April, 2022;
originally announced April 2022.
-
Quantifying Uncertainty for Temporal Motif Estimation in Graph Streams under Sampling
Authors:
Xiaojing Zhu,
Eric D. Kolaczyk
Abstract:
Dynamic networks, a.k.a. graph streams, consist of a set of vertices and a collection of timestamped interaction events (i.e., temporal edges) between vertices. Temporal motifs are defined as classes of (small) isomorphic induced subgraphs on graph streams, considering both edge ordering and duration. As with motifs in static networks, temporal motifs are the fundamental building blocks for tempor…
▽ More
Dynamic networks, a.k.a. graph streams, consist of a set of vertices and a collection of timestamped interaction events (i.e., temporal edges) between vertices. Temporal motifs are defined as classes of (small) isomorphic induced subgraphs on graph streams, considering both edge ordering and duration. As with motifs in static networks, temporal motifs are the fundamental building blocks for temporal structures in dynamic networks. Several methods have been designed to count the occurrences of temporal motifs in graph streams, with recent work focusing on estimating the count under various sampling schemes along with concentration properties. However, little attention has been given to the problem of uncertainty quantification and the asymptotic statistical properties for such count estimators. In this work, we establish the consistency and the asymptotic normality of a certain Horvitz-Thompson type of estimator in an edge sampling framework for deterministic graph streams, which can be used to construct confidence intervals and conduct hypothesis testing for the temporal motif count under sampling. We also establish similar results under an analogous stochastic model. Our results are relevant to a wide range of applications in social, communication, biological, and brain networks, for tasks involving pattern discovery.
△ Less
Submitted 21 February, 2022;
originally announced February 2022.
-
Stochastic polynomial chaos expansions to emulate stochastic simulators
Authors:
X. Zhu,
B. Sudret
Abstract:
In the context of uncertainty quantification, computational models are required to be repeatedly evaluated. This task is intractable for costly numerical models. Such a problem turns out to be even more severe for stochastic simulators, the output of which is a random variable for a given set of input parameters. To alleviate the computational burden, surrogate models are usually constructed and e…
▽ More
In the context of uncertainty quantification, computational models are required to be repeatedly evaluated. This task is intractable for costly numerical models. Such a problem turns out to be even more severe for stochastic simulators, the output of which is a random variable for a given set of input parameters. To alleviate the computational burden, surrogate models are usually constructed and evaluated instead. However, due to the random nature of the model response, classical surrogate models cannot be applied directly to the emulation of stochastic simulators. To efficiently represent the probability distribution of the model output for any given input values, we develop a new stochastic surrogate model called stochastic polynomial chaos expansions. To this aim, we introduce a latent variable and an additional noise variable, on top of the well-defined input variables, to reproduce the stochasticity. As a result, for a given set of input parameters, the model output is given by a function of the latent variable with an additive noise, thus a random variable. In this paper, we propose an adaptive algorithm which does not require repeated runs of the simulator for the same input parameters. The performance of the proposed method is compared with the generalized lambda model and a state-of-the-art kernel estimator on two case studies in mathematical finance and epidemiology and on an analytical example whose response distribution is bimodal. The results show that the proposed method is able to accurately represent general response distributions, i.e., not only normal or unimodal ones. In terms of accuracy, it generally outperforms both the generalized lambda model and the kernel density estimator.
△ Less
Submitted 26 November, 2022; v1 submitted 7 February, 2022;
originally announced February 2022.
-
An Asymptotic Analysis of Minibatch-Based Momentum Methods for Linear Regression Models
Authors:
Yuan Gao,
Xuening Zhu,
Haobo Qi,
Guodong Li,
Riquan Zhang,
Hansheng Wang
Abstract:
Momentum methods have been shown to accelerate the convergence of the standard gradient descent algorithm in practice and theory. In particular, the minibatch-based gradient descent methods with momentum (MGDM) are widely used to solve large-scale optimization problems with massive datasets. Despite the success of the MGDM methods in practice, their theoretical properties are still underexplored.…
▽ More
Momentum methods have been shown to accelerate the convergence of the standard gradient descent algorithm in practice and theory. In particular, the minibatch-based gradient descent methods with momentum (MGDM) are widely used to solve large-scale optimization problems with massive datasets. Despite the success of the MGDM methods in practice, their theoretical properties are still underexplored. To this end, we investigate the theoretical properties of MGDM methods based on the linear regression models. We first study the numerical convergence properties of the MGDM algorithm and further provide the theoretically optimal tuning parameters specification to achieve faster convergence rate. In addition, we explore the relationship between the statistical properties of the resulting MGDM estimator and the tuning parameters. Based on these theoretical findings, we give the conditions for the resulting estimator to achieve the optimal statistical efficiency. Finally, extensive numerical experiments are conducted to verify our theoretical results.
△ Less
Submitted 2 November, 2021;
originally announced November 2021.
-
Graphical Assistant Grouped Network Autoregression Model: a Bayesian Nonparametric Recourse
Authors:
Yimeng Ren,
Xuening Zhu,
Guanyu Hu
Abstract:
Vector autoregression model is ubiquitous in classical time series data analysis. With the rapid advance of social network sites, time series data over latent graph is becoming increasingly popular. In this paper, we develop a novel Bayesian grouped network autoregression model to simultaneously estimate group information (number of groups and group configurations) and group-wise parameters. Speci…
▽ More
Vector autoregression model is ubiquitous in classical time series data analysis. With the rapid advance of social network sites, time series data over latent graph is becoming increasingly popular. In this paper, we develop a novel Bayesian grouped network autoregression model to simultaneously estimate group information (number of groups and group configurations) and group-wise parameters. Specifically, a graphically assisted Chinese restaurant process is incorporated under framework of the network autoregression model to improve the statistical inference performance. An efficient Markov chain Monte Carlo sampling algorithm is used to sample from the posterior distribution. Extensive studies are conducted to evaluate the finite sample performance of our proposed methodology. Additionally, we analyze two real datasets as illustrations of the effectiveness of our approach.
△ Less
Submitted 11 October, 2021;
originally announced October 2021.
-
A Sequential Addressing Subsampling Method for Massive Data Analysis under Memory Constraint
Authors:
Rui Pan,
Yingqiu Zhu,
Baishan Guo,
Xuening Zhu,
Hansheng Wang
Abstract:
The emergence of massive data in recent years brings challenges to automatic statistical inference. This is particularly true if the data are too numerous to be read into memory as a whole. Accordingly, new sampling techniques are needed to sample data from a hard drive. In this paper, we propose a sequential addressing subsampling (SAS) method, that can sample data directly from the hard drive. T…
▽ More
The emergence of massive data in recent years brings challenges to automatic statistical inference. This is particularly true if the data are too numerous to be read into memory as a whole. Accordingly, new sampling techniques are needed to sample data from a hard drive. In this paper, we propose a sequential addressing subsampling (SAS) method, that can sample data directly from the hard drive. The newly proposed SAS method is time saving in terms of addressing cost compared to that of the random addressing subsampling (RAS) method. Estimators (e.g., the sample mean) based on the SAS subsamples are constructed, and their properties are studied. We conduct a series of simulation studies to verify the finite sample performance of the proposed SAS estimators. The time cost is also compared between the SAS and RAS methods. An analysis of the airline data is presented for illustration purpose.
△ Less
Submitted 3 October, 2021;
originally announced October 2021.
-
Disentangling positive and negative partisanship in social media interactions using a coevolving latent space network with attractors model
Authors:
Xiaojing Zhu,
Cantay Caliskan,
Dino P. Christenson,
Konstantinos Spiliopoulos,
Dylan Walker,
Eric D. Kolaczyk
Abstract:
We develop a broadly applicable class of coevolving latent space network with attractors (CLSNA) models, where nodes represent individual social actors assumed to lie in an unknown latent space, edges represent the presence of a specified interaction between actors, and attractors are added in the latent level to capture the notion of attractive and repulsive forces. We apply the CLSNA models to u…
▽ More
We develop a broadly applicable class of coevolving latent space network with attractors (CLSNA) models, where nodes represent individual social actors assumed to lie in an unknown latent space, edges represent the presence of a specified interaction between actors, and attractors are added in the latent level to capture the notion of attractive and repulsive forces. We apply the CLSNA models to understand the dynamics of partisan polarization on social media, where we expect Republicans and Democrats to increasingly interact with their own party and disengage with the opposing party. Using longitudinal social networks from the social media platforms Twitter and Reddit, we investigate the relative contributions of positive (attractive) and negative (repulsive) forces among political elites and the public, respectively. Our goals are to disentangle the positive and negative forces within and between parties and explore if and how they change over time. Our analysis confirms the existence of partisan polarization in social media interactions among both political elites and the public. Moreover, while positive partisanship is the driving force of interactions across the full periods of study for both the public and Democratic elites, negative partisanship has come to dominate Republican elites' interactions since the run-up to the 2016 presidential election.
△ Less
Submitted 13 August, 2022; v1 submitted 27 September, 2021;
originally announced September 2021.