H114-0006
Machine Learning for Early Warning of Cyanobacteria Blooms in Vermont’s Lake Champlain

Friday, 11 December 2020
Poster
Mahalia Clark, Timothy Laracy, Wilton Burns, Safwan Wshah and Gillian L Galford, University of Vermont, Burlington, VT, United States
Abstract:
Cyanobacteria blooms are a major problem for Lake Champlain, one of the USA’s larger lakes. Recurring each summer, they produce toxins that disrupt aquatic ecosystems, endanger public health, and impact home values. Measuring cyanobacteria levels directly is time and resource intensive, making them difficult to track over time. Blooms are also difficult to model, since they are driven by the complex interactions of many factors. We applied a variety of machine learning methods to a unique, long-term, public data set produced by the Vermont Department of Environmental Conservation, with the goal of predicting elevated cyanobacteria levels from seven other water quality variables.

We used 10 years of data from 15 sites around Lake Champlain to train two regression models and four classification models. A cyanobacteria biovolume threshold of 4e8 𝜇𝑚3/𝐿 was used to label ‘bloom’ conditions for classification. Support vector machine (SVM) classifiers with linear, polynomial, and RBF kernels were the most successful models, distinguishing blooms from non-blooms with perfect sensitivity and high specificity (recall = 1, ROC AUC > 0.93). Logistic regression was moderately sensitive and specific (recall = 0.75, ROC AUC = 0.84), while random forest classification had lower sensitivity and precision (recall = 0.5, F1 = 0.67). A fully connected artificial neural network had relatively high sensitivity but very low precision (recall = 0.8, PR AUC = 0.35). Regression was less successful than classification: linear regression and random forest regression explained a low portion of variance, with R2 values of 0.30 and 0.45 respectively. We show the potential for SVM to predict cyanobacteria levels indirectly. Machine learning methods tend to be even more accurate on larger datasets. In the long term, they could be applied to high-frequency data from buoy-mounted data loggers or satellite imagery, reducing the need for direct cyanobacteria measurements.