H114-0006
Machine Learning for Early Warning of Cyanobacteria Blooms in Vermont’s Lake Champlain
Abstract:
We used 10 years of data from 15 sites around Lake Champlain to train two regression models and four classification models. A cyanobacteria biovolume threshold of 4e8 𝜇𝑚3/𝐿 was used to label ‘bloom’ conditions for classification. Support vector machine (SVM) classifiers with linear, polynomial, and RBF kernels were the most successful models, distinguishing blooms from non-blooms with perfect sensitivity and high specificity (recall = 1, ROC AUC > 0.93). Logistic regression was moderately sensitive and specific (recall = 0.75, ROC AUC = 0.84), while random forest classification had lower sensitivity and precision (recall = 0.5, F1 = 0.67). A fully connected artificial neural network had relatively high sensitivity but very low precision (recall = 0.8, PR AUC = 0.35). Regression was less successful than classification: linear regression and random forest regression explained a low portion of variance, with R2 values of 0.30 and 0.45 respectively. We show the potential for SVM to predict cyanobacteria levels indirectly. Machine learning methods tend to be even more accurate on larger datasets. In the long term, they could be applied to high-frequency data from buoy-mounted data loggers or satellite imagery, reducing the need for direct cyanobacteria measurements.