IN002-05
Jupyter meets the Earth: advancing an open ecosystem that supports science

Monday, 7 December 2020: 07:16
Virtual
Fernando Perez1, Joseph Hamman2, Laurel Larsen3, Kevin Paul4, Lindsey Justine Heagy5, Edom Moges6, Anderson Banihirwe7, Alice Cima5, Facundo Sapienza5, Erik Sundell8 and Chris Holdgraf9, (1)University of California Berkeley, Statistics, Berkeley, CA, United States, (2)CarbonPlan, Seattle, United States, (3)University of California, Berkeley, CA, United States, (4)National Center for Atmospheric Research, Boulder, CO, United States, (5)University of California, Berkeley, Statistics, Berkeley, CA, United States, (6)University of California Berkeley, Berkeley, CA, United States, (7)National Center for Atmospheric Research, Boulder, United States, (8)Sundell Open Source, Stockholm, Sweden, (9)University of California, Berkeley, Berkeley Institute for Data Science, Berkeley, CA, United States
Abstract:
Jupyter Notebooks are ubiquitous in exploratory computational research and data analysis. Though the Notebook is the most well-known component of the Jupyter ecosystem, Jupyter technologies span the spectrum from JupyterHub, which manages users and Jupyter sessions on cloud or high-performance computing infrastructure, to JupyterLab, the modular, extensible interface for interactive computing, to Widgets and Voilà dashboards that allow researchers to easily build graphical interactive controls connected to computational tools.

The Pangeo Project has demonstrated that deploying Jupyter infrastructure instrumented with Python packages for efficiently working with large multidimensional datasets (Xarray) and performing parallel computation (Dask) on the cloud can enable researchers to work interactively with Tera- and Peta-byte scale data sets. Jupyter meets the Earth, a project under the NSF EarthCube program (awards 1928406 & 1928374), uses research use-cases from domains across the geosciences to motivate technical developments within the Jupyter and Pangeo ecosystems.

This project is a close partnership between domain geoscientists and software developers, where both teams work as co-equal partners. We will provide an overview of the motivating research in cryosphere science, hydrology, climate and geophysical imaging, and highlight how technical developments in the Jupyter ecosystem facilitate data discovery, interactive large-scale computations on cloud and HPC infrastructure, and enable researchers to share work in a reproducible, engaging manner. This infrastructure has been deployed at NCAR’s Cheyenne supercomputer, enabling laptop-like interactive workflows at HPC scale that can be translated to commercial cloud environments with limited (sometimes zero) need for modification.

Finally, we will discuss avenues for getting involved and contributing to this ecosystem of open source software tools that serve researchers across the geosciences.