H166-0032
Sensitivity Analysis of Input Feature Selection in Multi-Layer Perceptron Neural Network to Predict Groundwater Levels

Tuesday, 15 December 2020
Poster
Reetik Sahu1, Juliane Müller2, Jangho Park2, Charuleka Varadharajan3, Bhavna Arora4, Boris Faybishenko3 and Deb Agarwal5, (1)Lawrence Berkeley National Laboratory, Computational Sciences Area, Berkeley, CA, United States, (2)Lawrence Berkeley National Laboratory, Center for Computational Science and Engineering, Berkeley, CA, United States, (3)Lawrence Berkeley National Laboratory, Berkeley, CA, United States, (4)Lawrence Berkeley National Laboratory, Energy Geosciences Division, Berkeley, CA, United States, (5)LBNL, Berkeley, CA, United States
Abstract:
Despite the increasing application of Machine Learning (ML) techniques in hydrology, there is a critical gap in analyzing the robustness and reliability of the predictions made by ML models. In our study, we addressed this issue by analyzing groundwater level predictions made using a Multilayer Perceptron (MLP) neural network with optimized hyperparameters for different types and lengths of training data. The MLP model is trained on historical observations of different hydrological and meteorological features, such as groundwater level, precipitation, river flow, and ambient temperature in various combinations, different training periods and temporal frequency. The numerical experiments were performed for three different locations in California with varying degrees of groundwater use. The sensitivity analysis-based numerical experiments revealed that the use of all available features in the training process did not always lead to the best prediction performance. In particular, river flow and precipitation were important features only to some locations. With the optimal selection of features, the MLP models can make reliable predictions for up to one year. We also observed that models built solely using groundwater and temperature measurements (without additional hydrological information) performed the worst at all locations and were unreliable. Such analysis can help derive recommendations for selecting the optimal training features for ML models to obtain accurate short- and long-term predictions in hydrological applications.