Sunday, Aug 2: 4:00 PM - 5:50 PM
6447
Contributed Papers
Thomas M. Menino Convention & Exhibition Center
Room: CC-156B
Main Sponsor
Survey Research Methods Section
Presentations
Agencies conducting large-scale educational assessments recognize that the problem of student non-response is a fundamental threat to the validity of the reported findings. Response rates are thus often reported as quantitative measures of data quality. These studies, however, typically do not distinguish between the statistical and the identification problems arising from non-response when estimating and reporting response rates. The statistical problem is a consequence of a reduced sample size and manifests itself as sampling error. The identification problem is a consequence of the unobservability of some students in the population, which forces analysts to make untestable distributional assumptions. The identification problem manifests as non-sampling error.
We develop a coherent strategy for estimating and reporting response rates that considers an orderly decomposition of these two inferential problems. Importantly, our study accounts for the subtleties arising from the multi-stage sample design governing the data collection, which makes this decomposition non-trivial. We exemplify our strategy with application to the International Computer and Information Literacy Study
Keywords
Survey methodology
Non-response in stratified, mult-stage random sample designs
Probability samples
Large-scale educational assessments
The apportionment of House seats to states after each decennial census can be viewed as a probability proportional to size (pps) sample allocation to the states. The sample size is 435, and the measure of size is the state's apportionment population (resident population plus U.S. federal employees and dependents living overseas, allocated to their home state), subject to the constraint that all states must receive at least one seat. An investigation of a group of sample allocations to areas that were supposed to be pps led to the discovery that the current method (Huntington-Hill) used for apportionment assigned excessive sample to areas with larger populations in the early sample assignment stages. An alternative algorithm was constructed that assigned sample to areas by minimizing the sum of the absolute values of the difference between the population proportion and the sample proportion across the areas at each stage. Subsequently the alternative algorithm was used to do a hypothetical apportionment following each census from 1790 and 2020. The results are similar or identical to Webster's method (used after the 1840, 1910, 1930 censuses).
Keywords
sample allocation
Balanced sampling uses auxiliary information to enhance sample representativeness in multipurpose survey settings. The USDA's National Agricultural Statistics Service (NASS) has traditionally implemented this through the Multivariate Probability Proportional to Size (MPPS) design, an extension of Brewer sampling. In collaboration with the National Opinion Research Center (NORC) at the University of Chicago, we assessed whether the cube method can further improve MPPS by incorporating balanced sampling principles. This presentation also evaluates integer‑calibration (INCA) algorithms, which apply discrete optimization over a constrained integer lattice, as a potential alternative to the cube method. We compare the statistical performance and computational requirements of INCA, originally developed for the U.S. Census of Agriculture, with those of the cube method. Using data from the 2017 Census of Agriculture, we quantify and contrast the relative errors produced by each approach, highlighting their practical implications for large‑scale survey design.
Keywords
Auxiliary information
Balanced sampling
Discrete optimization
Multipurpose survey
MPPS sampling
Relative errors
Speaker
Yang Cheng, National Agricultural Statistics Service
Co-Author(s)
Luca Sartore, National Institute of Statistical Sciences
Valbona Bejleri, United States Department of Agriculture – National Agricultural Statistics Service
The purpose of this paper is to 1) develop a three-term Edgeworth expansion to the order n^(-1) for the Studentized sample mean under stratified simple random sampling from a finite population and 2) use this expansion to establish minimum sample size rules for STSRS that measure the normal approximation of the Studentized sample mean with a user-specified coverage tolerance for the traditional one- or two-sided confidence interval of the population mean. The new rules improve the extended Cochran's sample size rule previously established by Qing and Valliant (2025) from a two-term Edgeworth expansion of the same statistic, yielding less conservative sample size thresholds across common allocations. In simulations spanning a wide range of skewed finite populations, the performance of the new rules is evaluated.
Keywords
Confidence interval
Coverage probability
Edgeworth expansion
Minimum sample size
Stratified simple random sampling
We present design-based Horvitz-Thompson and Multiplicity estimators of the size, total and mean of a response variable associated with the elements of a hidden population, such as drug users, to be used with the link-tracing sampling variant proposed by Félix-Medina and Thompson (Jour. Official Stat., 2004). In this sampling variant a frame of venues where the elements of the population tend to gather is constructed. The frame does not need to cover the whole population. An initial sample of venues is selected and people in those sites are asked to name other members of the population. Since the computation of the design-based estimators require to know the number of venues in the frame that are linked to each sampled person and this information is not observable, we consider a Bayesian model for the distribution of those numbers which allows us to estimate them by means of a Metropolis-Hastings within Gibbs sampling procedure and consequently to compute the design-based estimators. Inference about the parameters of interest is carried out under the design-based approach. The results of a numerical study indicate that the performance of the proposed estimators is acceptable.
Keywords
Bayesian inference
Design-based inference
Hard-to-detect population
Hidden-population
Markov-Chain Monte Carlo
Snowball sampling
Traditional fixed-size confidence region (FSCR) methods for estimating the mean of a multivariate normal distribution often fix the region's maximum diameter in advance, without regard to the quality of available data. We propose a new approach that incorporates data quality into determining the region's size. Starting with a modified FSCR method where the structure of the variance-covariance matrix is known, we introduce a minimum risk FSCR (MRFSCR) framework inspired by point estimation methods that balance estimation accuracy and sampling cost. We develop a unified multistage sampling strategy to construct these regions, ensuring desirable asymptotic properties. The methodology is illustrated through practical sampling strategies, simulation studies, and real-data examples.
Keywords
sequential sampling
multivariate analysis
confidence region construction
statistical inference