IN003-14
SpatioTemporal Adaptive Resolution Encoding as a Framework for Integrative Analysis with Volume and Variety Scalability

Monday, 7 December 2020: 07:39
Virtual
Michael Lee Rilee, NASA Goddard Space Flight Center, Greenbelt, MD, United States; Rilee Systems Technologies, Derwood, MD, United States, Kwo-Sen Kuo, NASA Goddard SFC, Greenbelt, MD, United States; Bayesics, LLC, Bowie, MD, United States, James H R Gallagher, OPeNDAP, Inc., Butte, MT, United States, Niklas Fabian Griessbaum, University of California Santa Barbara, Bren School, Santa Barbara, CA, United States, James Frew, University of California, Santa Barbara, CA, United States, Edward Hartnett, Ed Hartnett Consulting, Boulder, CO, United States and Robert Edward Wolfe, NASA GSFC, Greenbelt, MD, United States
Abstract:
The SpatioTemporal Adaptive Resolution Encoding (STARE) has been developed to enable integrative Earth Science Data (ESD) analysis to scale in both volume and variety, by co-aligning data both geo-spatiotemporally and on storage/computing resources. Encoding geolocations as floating-point numbers alongside data as arrays is not amenable to partitioning for scalable parallel processing, as there is no organizing principle connecting the data's spatiotemporal layout to its native arrays, particularly for low-level, non-gridded data. STARE provides that organizing principle, using hierarchically constructed integer indices to optimize processing, minimizing intermediate data products, duplication, and costly data transfers. Essential for ESD analysis, resolution is explicitly incorporated in STARE. These innovative features of STARE have proved advantageous for scaling parallel processing over voluminous, diverse ESD, and also for any analysis requiring the harmonization and integration of dissimilar data sources. STARE’s benefits are not restricted to large-scale analyses over extensive areas (e.g. global) or durations; spatiotemporally targeted analyses also enjoy the unified data indexing. Thus, STARE is meant for large-scale ESD production, archive, and analysis operations looking for better scalability with variety and volume, as well as the day-to-day work of individuals integrating diverse data. A promise of STARE is mitigating the need to resample data from multiple sources into a common grid as a prerequisite for analysis, especially for swath data, whose varying resolutions and idiosyncratic coordinate systems are barriers to their exploitation. STARE is uniquely positioned to “unlock” broader access to swath data. STARE’s hierarchical integer indices with embedded resolution have the ESD-harmonizing power to improve interoperability, to optimize parallel scalability by minimizing data movement, and thus to take science productivity of ESD analysis to an unprecedented height. In this presentation, we show examples of STARE-based analyses and integrations of diverse low-level data using such tools as OPeNDAP's STARE-aware server functions, Dask-accelerated PySTARE, and STAREPandas in the context of Cloud, archive, and high-end computing environments.