A061-0009
Using Statistical and Machine Learning Models with Remotely Sensed Data to Estimate PM2.5 in the San Francisco Bay Area
Using Statistical and Machine Learning Models with Remotely Sensed Data to Estimate PM2.5 in the San Francisco Bay Area
Wednesday, 9 December 2020
Poster
Abstract:
Ambient fine particulate matter (PM2.5) is associated with significant adverse health impacts. Continuous, high quality and high resolution PM2.5 data has the potential to be greatly useful in public health research and mitigation efforts, but PM2.5 monitors are few and unevenly distributed over the landscape. In California, this is of particular concern because catastrophic wildfires have caused and are projected to continue causing episodes of very high levels of PM2.5. Previous studies have shown the potential for Aerosol Optical Depth (AOD), meteorological data, emissions, and land cover/land use (LCLU) data to estimate PM2.5 using a variety of models. However, the most recent research has yet to be applied in the San Francisco Bay Area, where high density episodes of PM2.5 were observed in 2017 and 2018. In addition, few studies have taken advantage of flexible and powerful machine learning algorithms to estimate PM2.5 levels, especially considering the variety of parameters known to improve such models. This study aims to apply the state of the art PM2.5 estimation techniques, including a proven two-stage model trained on AOD, meteorological, and LCLU data, and compare it to promising ML algorithms including random forests, and gradient boosted decision trees. We envision that this approach will lead to greatly improved estimation of PM2.5 in California, and that more flexible ML techniques will allow for improved results when predicting extreme PM2.5 events, such as resulting from a wildfire, which are particularly important for public health research.