-
Estimating Propensities of Selection for Big Datasets via Data Integration
Authors:
Lyndon Ang,
Robert Clark,
Bronwyn Loong,
Anders Holmberg
Abstract:
Big data presents potential but unresolved value as a source for analysis and inference. However,selection bias, present in many of these datasets, needs to be accounted for so that appropriate inferences can be made on the target population. One way of approaching the selection bias issue is to first estimate the propensity of inclusion in the big dataset for each member of the big dataset, and t…
▽ More
Big data presents potential but unresolved value as a source for analysis and inference. However,selection bias, present in many of these datasets, needs to be accounted for so that appropriate inferences can be made on the target population. One way of approaching the selection bias issue is to first estimate the propensity of inclusion in the big dataset for each member of the big dataset, and then to apply these propensities in an inverse probability weighting approach to produce population estimates. In this paper, we provide details of a new variant of existing propensity score estimation methods that takes advantage of the ability to integrate the big data with a probability sample. We compare the ability of this method to produce efficient inferences for the target population with several alternative methods through an empirical study.
△ Less
Submitted 7 January, 2025;
originally announced January 2025.
-
An Empirical Comparison of Methods to Produce Business Statistics Using Non-Probability Data
Authors:
Lyndon Ang,
Robert Clark,
Bronwyn Loong,
Anders Holmberg
Abstract:
There is a growing trend among statistical agencies to explore non-probability data sources for producing more timely and detailed statistics, while reducing costs and respondent burden. Coverage and measurement error are two issues that may be present in such data. The imperfections may be corrected using available information relating to the population of interest, such as a census or a referenc…
▽ More
There is a growing trend among statistical agencies to explore non-probability data sources for producing more timely and detailed statistics, while reducing costs and respondent burden. Coverage and measurement error are two issues that may be present in such data. The imperfections may be corrected using available information relating to the population of interest, such as a census or a reference probability sample.
In this paper, we compare a wide range of existing methods for producing population estimates using a non-probability dataset through a simulation study based on a realistic business population. The study was conducted to examine the performance of the methods under different missingness and data quality assumptions. The results confirm the ability of the methods examined to address selection bias. When no measurement error is present in the non-probability dataset, a screening dual-frame approach for the probability sample tends to yield lower sample size and mean squared error results. The presence of measurement error and/or nonignorable missingness increases mean squared errors for estimators that depend heavily on the non-probability data. In this case, the best approach tends to be to fall back to a model-assisted estimator based on the probability sample.
△ Less
Submitted 17 September, 2024; v1 submitted 23 May, 2024;
originally announced May 2024.
-
A note on the optimum allocation of resources to follow up unit nonrespondents in probability
Authors:
Su-Ming Tam,
Anders Holmberg,
Summer Wang
Abstract:
Common practice to address nonresponse in probability surveys in National Statistical Offices is to follow up every nonrespondent with a view to lifting response rates. As response rate is an insufficient indicator of data quality, it is argued that one should follow up nonrespondents with a view to reducing the mean squared error (MSE) of the estimator of the variable of interest. In this paper,…
▽ More
Common practice to address nonresponse in probability surveys in National Statistical Offices is to follow up every nonrespondent with a view to lifting response rates. As response rate is an insufficient indicator of data quality, it is argued that one should follow up nonrespondents with a view to reducing the mean squared error (MSE) of the estimator of the variable of interest. In this paper, we propose a method to allocate the nonresponse follow-up resources in such a way as to minimise the MSE under a quasi-randomisation framework. An example to illustrate the method using the 2018/19 Rural Environment and Agricultural Commodities Survey from the Australian Bureau of Statistics is provided.
△ Less
Submitted 6 June, 2023;
originally announced June 2023.
-
Linking Administrative Data: An Evolutionary Schema
Authors:
Jack Lothian,
Anders Holmberg,
Allyson Seyb
Abstract:
Statistics New Zealand (Stats NZ) has committed unreservedly to an administrative data first policy. Thus, all new methods used at Stats NZ are to be viewed within this context and discussing strategies for using administrative data is an integral part of every working day. As statistical methodologists, the three authors were drawn into these discussions. Like most methodologists, the authors see…
▽ More
Statistics New Zealand (Stats NZ) has committed unreservedly to an administrative data first policy. Thus, all new methods used at Stats NZ are to be viewed within this context and discussing strategies for using administrative data is an integral part of every working day. As statistical methodologists, the three authors were drawn into these discussions. Like most methodologists, the authors see surveys and the publications of their results as a process where estimation is the key tool to achieve the final goal of an accurate statistical output. Randomness and sampling exists to support this goal, and early on it was clear to us that the incoming it-is-what-it-is data sources were not randomly selected. These sources were obviously biased and thus would produce biased estimates. So, we set out to design a strategy to deal with this issue. This led us to the concept of representativeness which is closely related to statistical bias but has a wider context invoking both randomness and judgement. The representativeness issue was the principal question that we set out to answer. The necessary components that we gathered for our solution are summarized in the paper.
Keywords: Representativeness, Timeline Databases, Statistical Registers, Estimation
△ Less
Submitted 20 December, 2017;
originally announced December 2017.
-
Mark-Recapture with Multiple Non-invasive Marks
Authors:
Simon J. Bonner,
Jason A. Holmberg
Abstract:
Non-invasive marks, including pigmentation patterns, acquired scars,and genetic mark- ers, are often used to identify individuals in mark-recapture experiments. If animals in a population can be identified from multiple, non-invasive marks then some individuals may be counted twice in the observed data. Analyzing the observed histories without accounting for these errors will provide incorrect inf…
▽ More
Non-invasive marks, including pigmentation patterns, acquired scars,and genetic mark- ers, are often used to identify individuals in mark-recapture experiments. If animals in a population can be identified from multiple, non-invasive marks then some individuals may be counted twice in the observed data. Analyzing the observed histories without accounting for these errors will provide incorrect inference about the population dynamics. Previous approaches to this problem include modeling data from only one mark and combining estimators obtained from each mark separately assuming that they are independent. Motivated by the analysis of data from the ECOCEAN online whale shark (Rhincodon typus) catalog, we describe a Bayesian method to analyze data from multiple, non-invasive marks that is based on the latent-multinomial model of Link et al. (2010). Further to this, we describe a simplification of the Markov chain Monte Carlo algorithm of Link et al. (2010) that leads to more efficient computation. We present results from the analysis of the ECOCEAN whale shark data and from simulation studies comparing our method with the previous approaches.
△ Less
Submitted 8 March, 2013; v1 submitted 21 August, 2012;
originally announced August 2012.