H159-05
Leveraging Applied Machine Learning for Flood Risk Assessment Using Open Source Data: A Stacked Ensemble Approach for Predicting Residential Flood Insurance Claims Across the State of Indiana
Leveraging Applied Machine Learning for Flood Risk Assessment Using Open Source Data: A Stacked Ensemble Approach for Predicting Residential Flood Insurance Claims Across the State of Indiana
Monday, 14 December 2020: 20:46
Virtual
Abstract:
Demand for low-cost, performant flood risk assessment and prediction modeling is rising in tandem with the widespread availability of sophisticated machine learning (ML) tools. When paired with projected growth in the severity and scope of hazard susceptibility due to climate change, the market for increasingly accurate and efficient risk assessment is expected to be robust among decision makers. Many existing flood risk models incorporate extensive hydrologic modeling to generate dependable forecasts. However, this approach tends to feature extended computational runtimes and high engineering labor costs. As ML frameworks mature, their potential to facilitate efficiencies for modeling anticipated losses from flooding while achieving comparable results is compelling. The goal of this research is to develop a statistical model that leverages open source data to predict mean residential building flood insurance claims within Indiana, USA. The target variable for the predictive model is FEMA National Flood Insurance Program (NFIP) redacted insurance claim amounts in USD collected from approximately 1985-2015. The predictor variables are approximately 65 attributes associated with each subwatershed, including geophysical characteristics, population demographics, discharge summary statistics, and mean NFIP insurance policy amounts. Redacted policy and claim insurance data were aggregated at the census tract level then proportionally assigned to subwatersheds via geographic union. The performance of four models – GLM, Random Forest, GBN, and ANN – were tested and compared using a k-fold train-test-validate protocol with gridded hyperparameter tuning using the H2O modeling framework. Predictions from hundreds of models generated via random grid search were supplied to the H2O stacked ensemble method, which is a supervised ensemble ML algorithm. Initial results indicate that stacked ensemble modeling using a GBM metalearner yielded predictions on par with the best performing ML models when comparing RMSE metrics using min-max scaled data (about 0.08), with an adjusted R2 of 0.22 calculated on the validation set. While feature engineering and comparison to hydrologic models is ongoing, near-term goals include achieving improved adjusted R2 results and generalization to additional US states.