A060-0001
A land use model for high-resolution black carbon estimation in Oakland, CA: A comparison of different machine learning models’ performance in spatial prediction
A land use model for high-resolution black carbon estimation in Oakland, CA: A comparison of different machine learning models’ performance in spatial prediction
Wednesday, 9 December 2020
Poster
Abstract:
With the development of high accuracy portable pollution sensing instruments and Global Positioning System (GPS) technology, the use of vehicles for mobile air pollution monitoring has been proposed to tackle some of the challenges of stationary monitoring sites. These mobile sensors can typically achieve high spatial resolution air pollution measurement at the cost of reduced temporal resolution, which have the inherently different characteristic with stationary monitors. However, there are limited studies evaluating different machine learning models’ ability to handle the mobile sensors’ measurement. In this paper, we couple land use model with three different machine learning models including Random Forest (RF), Support Vector Machine (SVM), and Neural Network (NN) to estimate black carbon (BC) concentrations in Oakland, CA. which is measured by Google street view vehicles. Since SVM is sensitive to input features, we apply LASSO, principle component analysis, conditional independence feature ordering, and genetic algorithm for feature selections from the output of the land use model; while no feature selection method is used for RF and NN. We use Bayesian Optimization method to automatically tune RF and SVM; while NN is manually tuned because of the complexity of the model’s structure. Based on the results, NN has the highest estimation accuracy among all three models, which is also the most time consuming one. For SVM with different feature selection methods and different number of input features, the optimized SVM model has R2 about 65%, while the RF model without any feature selections has R2 at 69%. In general for mobile sensors’ measurement, if the estimation accuracy is the only concern, NN is recommended, while extra computational resources like GPU may be necessary. But if there are other concerns like computational speed or extra work required after concentration estimation, RF is better because of its short tuning time, automatic tuning process, and relatively high estimation accuracy.