-
Model-assisted estimation in high-dimensional settings for survey data
Authors:
Mehdi Dagdoug,
Camelia Goga,
David Haziza
Abstract:
Model-assisted estimators have attracted a lot of attention in the last three decades. These estimators attempt to make an efficient use of auxiliary information available at the estimation stage. A working model linking the survey variable to the auxiliary variables is specified and fitted on the sample data to obtain a set of predictions, which are then incorporated in the estimation procedures.…
▽ More
Model-assisted estimators have attracted a lot of attention in the last three decades. These estimators attempt to make an efficient use of auxiliary information available at the estimation stage. A working model linking the survey variable to the auxiliary variables is specified and fitted on the sample data to obtain a set of predictions, which are then incorporated in the estimation procedures. A nice feature of model-assisted procedures is that they maintain important design properties such as consistency and asymptotic unbiasedness irrespective of whether or not the working model is correctly specified. In this article, we examine several model-assisted estimators from a design-based point of view and in a high-dimensional setting, including penalized estimators and tree-based estimators. We conduct an extensive simulation study using data from the Irish Commission for Energy Regulation Smart Metering Project, in order to assess the performance of several model-assisted estimators in terms of bias and efficiency in this high-dimensional data set.
△ Less
Submitted 7 April, 2021; v1 submitted 14 December, 2020;
originally announced December 2020.
-
Efficient multiply robust imputation in the presence of influential units in surveys
Authors:
Sixia Chen,
David Haziza,
Victoire Michal
Abstract:
Item nonresponse is a common issue in surveys. Because unadjusted estimators may be biased in the presence of nonresponse, it is common practice to impute the missing values with the objective of reducing the nonresponse bias as much as possible. However, commonly used imputation procedures may lead to unstable estimators of population totals/means when influential units are present in the set of…
▽ More
Item nonresponse is a common issue in surveys. Because unadjusted estimators may be biased in the presence of nonresponse, it is common practice to impute the missing values with the objective of reducing the nonresponse bias as much as possible. However, commonly used imputation procedures may lead to unstable estimators of population totals/means when influential units are present in the set of respondents. In this article, we consider the class of multiply robust imputation procedures that provide some protection against the failure of underlying model assumptions. We develop an efficient version of multiply robust estimators based on the concept of conditional bias, a measure of influence. We present the results of a simulation study to show the benefits of the proposed method in terms of bias and efficiency.
△ Less
Submitted 4 October, 2020;
originally announced October 2020.
-
Imputation procedures in surveys using nonparametric and machine learning methods: an empirical comparison
Authors:
Mehdi Dagdoug,
Camelia Goga,
David Haziza
Abstract:
Nonparametric and machine learning methods are flexible methods for obtaining accurate predictions. Nowadays, data sets with a large number of predictors and complex structures are fairly common. In the presence of item nonresponse, nonparametric and machine learning procedures may thus provide a useful alternative to traditional imputation procedures for deriving a set of imputed values. In this…
▽ More
Nonparametric and machine learning methods are flexible methods for obtaining accurate predictions. Nowadays, data sets with a large number of predictors and complex structures are fairly common. In the presence of item nonresponse, nonparametric and machine learning procedures may thus provide a useful alternative to traditional imputation procedures for deriving a set of imputed values. In this paper, we conduct an extensive empirical investigation that compares a number of imputation procedures in terms of bias and efficiency in a wide variety of settings, including high-dimensional data sets. The results suggest that a number of machine learning procedures perform very well in terms of bias and efficiency.
△ Less
Submitted 14 December, 2020; v1 submitted 13 July, 2020;
originally announced July 2020.
-
Model-assisted estimation through random forests in finite population sampling
Authors:
Mehdi Dagdoug,
Camelia Goga,
David Haziza
Abstract:
In surveys, the interest lies in estimating finite population parameters such as population totals and means. In most surveys, some auxiliary information is available at the estimation stage. This information may be incorporated in the estimation procedures to increase their precision. In this article, we use random forests to estimate the functional relationship between the survey variable and th…
▽ More
In surveys, the interest lies in estimating finite population parameters such as population totals and means. In most surveys, some auxiliary information is available at the estimation stage. This information may be incorporated in the estimation procedures to increase their precision. In this article, we use random forests to estimate the functional relationship between the survey variable and the auxiliary variables. In recent years, random forests have become attractive as National Statistical Offices have now access to a variety of data sources, potentially exhibiting a large number of observations on a large number of variables. We establish the theoretical properties of model-assisted procedures based on random forests and derive corresponding variance estimators. A model-calibration procedure for handling multiple survey variables is also discussed. The results of a simulation study suggest that the proposed point and estimation procedures perform well in term of bias, efficiency, and coverage of normal-based confidence intervals, in a wide variety of settings. Finally, we apply the proposed methods using data on radio audiences collected by Médiamétrie, a French audience company.
△ Less
Submitted 17 June, 2021; v1 submitted 22 February, 2020;
originally announced February 2020.
-
Joint imputation procedures for categorical variables
Authors:
Hélène Chaput,
Guillaume Chauvet,
David Haziza,
Laurianne Salembier,
Julie Solard
Abstract:
Marginal imputation, which consists of imputing each item requiring imputation separately, is often used in surveys. This type of imputation procedures leads to asymptotically unbiased estimators of simple parameters such as population totals (or means), but tends to distort relationships between variables. As a result, it generally leads to biased estimators of bivariate parameters such as coeffi…
▽ More
Marginal imputation, which consists of imputing each item requiring imputation separately, is often used in surveys. This type of imputation procedures leads to asymptotically unbiased estimators of simple parameters such as population totals (or means), but tends to distort relationships between variables. As a result, it generally leads to biased estimators of bivariate parameters such as coefficients of correlation or odd-ratios. Household and social surveys typically collect categorical variables, for which missing values are usually handled by nearest-neighbour imputation or random hot-deck imputation. In this paper, we propose a simple random imputation procedure, closely related to random hot-deck imputation, which succeeds in preserving the relationship between categorical variables. Also, a fully efficient version of the latter procedure is proposed. A limited simulation study compares several estimation procedures in terms of relative bias and relative efficiency.
△ Less
Submitted 3 November, 2015;
originally announced November 2015.