A196-04
Using Machine Learning to Predict the Accuracy of Thunderstorm Forecasts from a Warn-on-Forecast Ensemble

Tuesday, 15 December 2020: 10:12
Virtual
Corey Potvin, National Severe Storms Lab, Norman, OK, United States, Montgomery Flora, University of Oklahoma, School of Meteorology, Norman, OK, United States and Patrick Skinner, University of Oklahoma and NOAA/National Severe Storms Laboratory, Cooperative Institute for Mesoscale Meteorological Studies, Norman, OK, United States
Abstract:
Ensembles facilitate the assessment of forecast uncertainty by providing a set of theoretically equally likely realizations of the future atmospheric state. In practice, however, the ensemble forecast distribution often seriously mischaracterizes the true forecast uncertainty, due primarily to biases in the initial conditions or forecast model(s). Ensemble under-dispersion is particularly problematic since it implies spuriously high forecast certainty and therefore could lead users to mistakenly discount potential outcomes, especially when the ensemble forecast is biased. In a convection-allowing ensemble such as that examined herein, under-dispersion could manifest as every ensemble member weakening a storm that in reality intensifies and produces severe weather. The misalignment between ensemble spread and true forecast uncertainty motivates the development of additional metrics to estimate forecast confidence so that forecasters, emergency managers, and other users can better utilize forecast guidance in taking appropriate action.

To address this topic, we leverage the several hundred ensemble forecasts generated by the NOAA National Severe Storms Laboratory Warn-on-Forecast System (WoFS) during the 2017-2020 warm seasons. We begin by developing an ensemble forecast accuracy score that is heavily informed by object-based verification metrics of storm occurrence and location. This score is computed for every WoFS forecast for every observed storm at lead times of [0, 30, ..., 180] minutes. We then examine relationships between the forecast scores and various characteristics of the observed storms and of near-storm sounding parameters. Upon identifying features that substantially correlate with WoFS forecast accuracy, we develop and evaluate machine learning models that input subsets of these features valid at 0 min or 30 min into the forecast and output predictions of whether the forecast at later lead times will score in the lower, middle, or upper tercile of all forecasts having the same lead time (i.e., have below average, near average, or above average accuracy). Preliminary results indicate that WoFS forecast performance varies non-monotonically with near-storm sounding parameters, is higher for mesoscale convective systems than for discrete storms, and is higher in the evening than the afternoon. A trained random forest model exhibits substantial skill in predicting forecast performance, with much of the skill deriving from the storm environment features and the ensemble forecast accuracy valid at the 30-min lead time.