59: Evaluating LLM‑Assisted Forecasting and Data Driven Conclusions within the Insurance Industry
Monday, Aug 3: 2:00 PM - 3:50 PM
2693
Contributed Posters
Thomas M. Menino Convention & Exhibition Center
We assess whether large language model (LLM) assistants can build credible forecasting and risk‑flagging pipelines from multi‑year monthly insurance data without external sources or disclosure of proprietary levels. Using product–state–club–policy‑characteristic segments over ~3 years, we compare multiple LLM workflows under controlled conditions, strict time‑based splits, and leakage controls, with a naïve seasonal reference. Primary outcomes are 1‑month forecasts of loss ratio; 3‑month is secondary and 12‑month exploratory. Evaluation uses rolling‑origin cross‑validation and scale‑free metrics (MASE primary; sMAPE, 80/95% coverage secondary). Adverse selection is defined prospectively: segments whose next‑quarter loss ratio and PIF each increase ≥10% versus trailing‑12 baselines. We score probabilistic flags with a variety of statistical tests. Model differences are tested via paired comparisons. Results quantify accuracy, calibration, and reproducibility of LLM‑assisted analytics while preserving confidentiality through scale‑free reporting only.
Large Language Models
Insurance
Forecasting
LLM
Main Sponsor
Business Analytics/Statistics Education Interest Group
You have unsaved changes.