Advances in AI and Machine Learning for Official Statistics and Survey Methodology

Scott Holan Chair
University of Missouri/U.S. Census Bureau
 
Scott Holan Organizer
University of Missouri/U.S. Census Bureau
 
Monday, Aug 3: 8:30 AM - 10:20 AM
1330 
Invited Paper Session 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-156C 

Applied

Yes

Main Sponsor

Survey Research Methods Section

Co Sponsors

Government Statistics Section
Section on Statistical Learning and Data Science

Presentations

Developing a workflow for AI-based image interpretation of forest inventory plots

Large Language Models (LLMs) show great promise for increasing the efficiency of repetitive processes implemented by organizations.  In collaboration with the USFS Forest Inventory and Analysis program (FIA), we attempted to create a LLM-based tool for a common forest inventory task: image interpretation.  This exploratory project provided FIA with a concrete example of how to incorporate LLMs into their work but also raised a host of questions around quality assurance and uncertainty quantification.  In my talk, I will go through our process, show how to use the ELLMER package in R for leveraging LLMs, and discuss the lessons learned about the advantages and disadvantages of utilizing LLMs for image interpretation. 

Keywords

large language models

ELLMER

forest inventory 

Speaker

Kelly McConville, Bucknell University

Co-Author(s)

Andrew Lister, US Forest Service
Grayson White
Odilon Ligan, Bucknell University
Jean Marie Ngabonziza, Bucknell University

New Approaches for Handling Nonlinearity in Area-level Small Area Estimation

Small area estimation models are critical for dissemination and understanding of important population characteristics within sub-domains that often have limited sample size. The classic Fay-Herriot model is perhaps the most widely used approach to generate such estimates. However, a limiting assumption of this approach is that the latent true population quantity has a linear relationship with the given covariates. We introduce two new approaches that allow for estimation of nonlinear relationships between the true population quantity and the covariates. First, through the use of random weight neural networks, we develop a Bayesian hierarchical extension of the Fay-Herriot model. Second, we consider the use of Bayesian additive regression trees (BART) embedded within a Fay-Herriot model. We illustrate our approaches through an empirical simulation study as well as an analysis of median household income for census tracts in the state of California.  

Speaker

Paul Parker, University of California Santa Cruz

Random Forest Models for Data from an Informative Sample

Despite the fact that random forests are effective and flexible nonparametric models with many potential applications on survey data, until recently there were no methods for estimating them that accounted for data from an informative sample design. Since survey data is usually collected using an informative sample design, it is necessary to have an algorithm for creating random forest models that account for this design during model estimation. Because random forests are the result of averaging a large number tree models obtained from bootstrapped samples of the original data and obtaining bootstrap samples can be hard to produce for a general set o non-independent data, the sample design has been ignore.d. We investigate a recently proposed method to predict response burden to the US Bureau of Labor Statistics Consumer Expenditure Survey. 

Keywords

Sample Design

Surveys

Machine Learning

rpms

Response Burden

Consumer Expenditure Survey 

Speaker

Daniell Toth, US Bureau of Labor Statistics