IN041-11
Assessing Extreme Regional Water Scarcity with Open Source Data and Cloud Computing
Abstract:
Data was collected from a variety of industry standard open source data providers (NASA, ESA, NOAA, etc.). The seven initial datasets included indicators for soil moisture, precipitation, groundwater storage, vegetation, and population. To facilitate varying areas of interest, the processing of source data was performed at a cell level. Further processing was performed to establish a per-cell historical baseline, thus allowing for dynamic aggregation of variance over time.
Several data quantity issues stemming from identifying deviances from historical trends were detected. Retrieval/storage and historical trend processing of the data, and aggregation at each pixel/cell across all available data, initially proved challenging, but both were mitigated upon transition to a series of scalable cloud-based environments. At this point in the project, the results indicate a moderately strong inference can be drawn from our current indicators as a measurement of Water Scarcity in a given region. As the model progressed, correlations were strengthened by back-testing the resulting Risk Scores to a well curated inventory of conflict data as provided by Uppsala Conflict Data Program (UCDP). Careful considerations were given to mitigate the overweighting of externalities, but further work is being done to identify additional data sources to allow for more precision and broader historical context.
Using unique data processing and analytical methods, the team interrogated openly available and remote sensing datasets in an effort to understand fluctuations in regional water scarcity. Further refinement of the data processing pipeline will enable improvements in accuracy.