Monday, Aug 3: 8:30 AM - 10:20 AM
1330
Invited Paper Session
Thomas M. Menino Convention & Exhibition Center
Room: CC-156C
Applied
Yes
Main Sponsor
Survey Research Methods Section
Co Sponsors
Government Statistics Section
Section on Statistical Learning and Data Science
Presentations
Large Language Models (LLMs) show great promise for increasing the efficiency of repetitive processes implemented by organizations. In collaboration with the USFS Forest Inventory and Analysis program (FIA), we attempted to create a LLM-based tool for a common forest inventory task: image interpretation. This exploratory project provided FIA with a concrete example of how to incorporate LLMs into their work but also raised a host of questions around quality assurance and uncertainty quantification. In my talk, I will go through our process, show how to use the ELLMER package in R for leveraging LLMs, and discuss the lessons learned about the advantages and disadvantages of utilizing LLMs for image interpretation.
Keywords
large language models
ELLMER
forest inventory
Small area estimation models are critical for dissemination and understanding of important population characteristics within sub-domains that often have limited sample size. The classic Fay-Herriot model is perhaps the most widely used approach to generate such estimates. However, a limiting assumption of this approach is that the latent true population quantity has a linear relationship with the given covariates. We introduce two new approaches that allow for estimation of nonlinear relationships between the true population quantity and the covariates. First, through the use of random weight neural networks, we develop a Bayesian hierarchical extension of the Fay-Herriot model. Second, we consider the use of Bayesian additive regression trees (BART) embedded within a Fay-Herriot model. We illustrate our approaches through an empirical simulation study as well as an analysis of median household income for census tracts in the state of California.
Speaker
Paul Parker, University of California Santa Cruz
Despite the fact that random forests are effective and flexible nonparametric models with many potential applications on survey data, until recently there were no methods for estimating them that accounted for data from an informative sample design. Since survey data is usually collected using an informative sample design, it is necessary to have an algorithm for creating random forest models that account for this design during model estimation. Because random forests are the result of averaging a large number tree models obtained from bootstrapped samples of the original data and obtaining bootstrap samples can be hard to produce for a general set o non-independent data, the sample design has been ignore.d. We investigate a recently proposed method to predict response burden to the US Bureau of Labor Statistics Consumer Expenditure Survey.
Keywords
Sample Design
Surveys
Machine Learning
rpms
Response Burden
Consumer Expenditure Survey