A022-12
Challenges of Running Global Climate Simulations at 2.8 km on GPU with the ICON Model
Monday, 7 December 2020: 16:33
Virtual
Xavier Lapillonne1, William Sawyer2, Remo Dietlicher3, Luis Kornblueh4, Sebastian Rast4, Reiner Schnur4, Monika Esch4, Giorgetta A. Marco4, Dmitry Alexeev5 and Robert Pincus6, (1)MeteoSwiss Federal Office of Meteorology and Climatology, Zurich, Switzerland, (2)CSCS Swiss National Supercomputing Centre, Lugano, Switzerland, (3)ETH Swiss Federal Institute of Technology Zurich, Zurich, Switzerland, (4)Max Planck Institute for Meteorology, Hamburg, Germany, (5)Nvidia, Santa Clara, CA, United States, (6)University of Colorado at Boulder, Boulder, CO, United States
Abstract:
The ICON modelling framework is a unified numerical weather and climate model used for operational numerical weather prediction as well as climate projection. In view of further pushing the frontier of possible applications and to make use of the latest evolution in hardware technologies, parts of the model were recently adapted to run on heterogeneous GPU system. This initial GPU port focus on components required for high-resolution climate application, and allow considering multi-years simulations at 2.8 km on the Piz Daint heterogeneous supercomputer. These simulations are planned as part of the QUIBICC project “The Quasi-Biennial Oscillation (QBO) in a changing climate”, which propose to investigate effects of climate change on the dynamics of the QBO.
In order to achieve optimal performance, all components required for the simulations are ported to GPU. For the dynamics, most of the physical parameterizations and infrastructure code the OpenACC compiler directives are used. For the soil parameterization, a Fortran based domain specific language (DSL) the CLAW-DSL has been considered. We discuss the challenges associated to port a large community code, about 1 million lines of code, as well as to run simulations on large-scale system at 2.8 km horizontal resolution in terms of run time and I/O constraints. We also present many of the optimizations implemented for GPUs and high-resolution simulations, such as asynchronous I/O, GPU-to-GPU communication, asynchronous execution of kernels, data compression on GPUs. We finally show performance comparison of the full model on CPU and GPU, achieving a speed up factor of approximately 5x, as well as scaling results on up to 2000 GPU nodes.