59: Evaluating LLM‑Assisted Forecasting and Data Driven Conclusions within the Insurance Industry

Philip Wong Speaker
CSAA IG
 
Nathan Cook Co-Author
CSAA
 
Patrick Hare Co-Author
CSAA
 
Sean McCarthy Co-Author
 
Gabe Cotapos Jr Co-Author
CSAA
 
eric lenz Co-Author
csaa
 
Monday, Aug 3: 2:00 PM - 3:50 PM
2693 
Contributed Posters 
Thomas M. Menino Convention & Exhibition Center 
We assess whether large language model (LLM) assistants can build credible forecasting and risk‑flagging pipelines from multi‑year monthly insurance data without external sources or disclosure of proprietary levels. Using product–state–club–policy‑characteristic segments over ~3 years, we compare multiple LLM workflows under controlled conditions, strict time‑based splits, and leakage controls, with a naïve seasonal reference. Primary outcomes are 1‑month forecasts of loss ratio; 3‑month is secondary and 12‑month exploratory. Evaluation uses rolling‑origin cross‑validation and scale‑free metrics (MASE primary; sMAPE, 80/95% coverage secondary). Adverse selection is defined prospectively: segments whose next‑quarter loss ratio and PIF each increase ≥10% versus trailing‑12 baselines. We score probabilistic flags with a variety of statistical tests. Model differences are tested via paired comparisons. Results quantify accuracy, calibration, and reproducibility of LLM‑assisted analytics while preserving confidentiality through scale‑free reporting only.

Keywords

Large Language Models

Insurance

Forecasting

LLM 

Main Sponsor

Business Analytics/Statistics Education Interest Group