H103-08
Transcending the uniqueness of places with large-sample multi-physics catchment modeling based on machine learning

Thursday, 10 December 2020: 19:34
Virtual
Chaopeng Shen1, Farshid Rahmani1, Wei Zhi2, Kuai Fang3, Wen-Ping Tsai1, Li Li1 and Kathryn Lawson1, (1)Pennsylvania State University Main Campus, Department of Civil and Environmental Engineering, University Park, PA, United States, (2)Pennsylvania State University Main Campus, Department of Energy and Mineral Engineering, University Park, PA, United States, (3)Pennsylvania State University Main Campus, Civil and Environmental Engineering, University Park, PA, United States
Abstract:
Watersheds in the world are often perceived as being unique from each other, requiring customized study for each basin. Models uniquely built for each watershed, in general, cannot be leveraged for other watersheds. It is also a customary practice in hydrology and related geoscientific disciplines to divide the whole domain into multiple regimes and study each region separately, in an approach sometimes called regionalization or stratification. However, in the era of big-data machine learning, models can learn across regions and identify commonalities and differences. In this presentation, we first show that machine learning can derive highly functional continental-scale models for streamflow, evapotranspiration, and water quality variables. Next, we investigate the optimal training data sampling scheme (how training data should be selected) for different regimes of data availability, and explore when it is a good idea to group data together for training. We show that in many cases there is an effect we call data synergy, where data from a diversity of regions joins together to train a better model for all regions. However, stratification by attributes can in some cases be useful in reducing noise during training. Overall, however, it can be said that these machine learning models have learned to transcend the uniqueness of places.