IN031-0005
CMIP6 in the Cloud - Why?
Abstract:
Here at LDEO, we started collecting the new CMIP6 data for our own server-side analysis platform in early 2019. As with prior CMIPx, we pooled our resources to support a full time data scientist to collect and pre-process the data as it was needed. At the same time, the Pangeo community was inspiring many of us with their vision of promoting open, reproducible, and scalable science. The Pangeo principles of 'taking the computations to the data' and 'reducing toil' also resonated with us. As a proof-of-concept activity, we started converting our own 'dark repository' of CMIP6 data into the new (parallel I/O friendly) zarr format and uploading to Google Cloud.
Thanks to hosting support from Google Cloud, this effort became the CMIP6 Google Public Dataset, accessible by anyone with an internet connection (no registration, passwords, cryptocards needed) using common tools (web browser, python, etc) together with the standard Pangeo collection of open source python packages.
We believe that the existence of cloud-based CMIP6 data has the potential to seriously accelerate climate science. Our attempts to fully automate the data transfer is outlined. We refer those interested in the many details (e.g, the handling of errata, support of data handles and licensing) to our Github repo. Instead, we will present aspects which might be of more general interest to the community: sustainability of the effort, lessons learned and how you can help.