H085-0001
A data-based approach for the estimation of streamflow components

Thursday, 10 December 2020
Poster
Yaser Kishawi, University of Nebraska Lincoln, Biological Systems Engineering, Lincoln, NE, United States, Chuyang Liu, University of Nebraska Lincoln, Civil and Environmental Engineering, Lincoln, NE, United States, Pham Huong, University of Nebraska-Lincoln, Civil and Environmental Engineering, Lincoln, United States, Antonio Alves Meira Neto, UFES Federal University of EspĂ­rito Santo, Environmental Engineering, Vitoria, Brazil, Paulo Tarso S. Oliveira, UFMS Federal University of Mato Grosso do Sul, Campo Grande, Brazil and Tirthankar Roy, University of Nebraska-Lincoln, Civil and Environmental Engineering, Omaha, NE, United States
Abstract:
Direct runoff and baseflow are essential streamflow components for various hydrologic applications. To balance the tradeoff between water resources sustainability and increasing water demand, efficient and accurate estimation of streamflow components at different spatial and temporal scales becomes indispensable. In this study, we first evaluated simple regression models from a previous study (Meira-Neto et al., under review) for estimating streamflow components across different spatial scales and hydroclimatic zones using the coefficient of determination, Kling-Gupta efficiency, and Nash-Sutcliffe efficiency. Our results revealed that regression models provide reasonable estimates in most cases; however, the performance of these models deteriorates in certain areas (e.g., arid regions). Analysis of the time series of flow components showed that the general flow patterns were well captured, but there were instances of both under- and over-predictions. The correlation between the predicted and observed flows was lower in the case of baseflow as compared to direct runoff and total flow. Next, we investigated how the estimation could be improved by implementing five machine learning models, including support vector machine, random forest, multilayer perceptron, XGBoost, and lasso regression. Initially, thirty features from the CAMELS (Catchment Attributes and Meteorology for Large-sample Studies) dataset, including topographic characteristics, climate indices, hydrological signatures, land cover, soil, and geology characteristics, were used to train the models to vote for the most important features. Various combinations of the voted features were then used to train a new set of models to select the optimal features for each streamflow components. We discuss how the different machine-learning models perform relative to each other and whether or not the shortcomings of the simple regression models can be addressed by more sophisticated machine learning models. Our data-based approach can be seen as a step forward towards the prediction of catchment fluxes at ungagged basins, while at the same time, it shows an interesting avenue to explore the climatic and landscape controls on different streamflow components.