H224-05
Capturing continental-scale dissolved oxygen patterns using deep learning and big data

Thursday, 17 December 2020: 05:50
Virtual
Wei Zhi1, Dapeng Feng1, Wen-Ping Tsai1, Gary Sterle2, Adrian Adam Harpold2, Chaopeng Shen1 and Li Li1, (1)Pennsylvania State University Main Campus, Department of Civil and Environmental Engineering, University Park, PA, United States, (2)University of Nevada Reno, Department of Natural Resources and Environmental Science, Reno, NV, United States
Abstract:
Dissolved oxygen (DO) is one of the most important indicators of water quality yet our understanding on controls of stream DO is limited. DO concentrations typically peak in winter at low temperature with maximum oxygen solubility and reach minimum in summer at high temperature when aquatic ecosystems actively use DO. We developed a deep-learning model of Long Short-Term Memory (LSTM) for the spatial-temporal dynamics of stream DO concentrations in the conterminous US using CAMELS-chem, a new stream chemistry database. The dataset includes DO concentration record with an average length of 26 years for over 200 reference, minimally disturbed watersheds. The dataset was split between training (1980-01-01 to 2000-12-31, 21 years) and testing (2001-01-01 to 2014-12-31, 14 years). The model generally captures spatial pattern and temporal dynamics of stream DO with mean and median NSE of 0.51 and 0.59 respectively. A fraction of 36% and 40% basins with mean NSE of 0.78 and 0.54 achieved good (NSE >=0.7) and fair (0.4 <= NSE < 0.7) performance, respectively, indicating a robust model for stream DO at continental scale. A total of 24 basins lacking data for the training period was used to evaluate model predictive power for ungauged basins. The model achieved a mean and median NSE of 0.6 and 0.77 for these ungauged basins. Only 5 basins with low DO concentrations shows unsatisfactory model performance (NSE < 0.4) as low DO levels were underrepresented in the model training process. A correlation analysis shows that model performance is largely influenced by standard deviation (R2 = 0.52) and minimum DO concentrations with (R2 = 0.40), respectively. The model performs better in high runoff-ratio (> 0.45) regions where precipitation peaks in winter. In conclusion, the deep-learning model with a national-wide CAMELS-chem dataset show the potential to model water quality at the continental scale and to predict water quality in ungauged basins.