Benchmarking for Machine Learning Development and Enhancement of Earth System Models
Benchmarking for Machine Learning Development and Enhancement of Earth System Models
Session ID#: 283019
Session Description:
Machine learning (ML) is rapidly transforming Earth system science, enabling advances in weather prediction, hydrology, cryosphere modeling, and beyond. However, evaluation remains fragmented, often relying on limited metrics that do not fully capture scientific validity, generalizability, or suitability for real-world applications.
This session focuses on the quantitative evaluation, benchmarking, and scientific interpretation of ML models across scientific domains. We emphasize frameworks that assess not only model skill, but also whether models are fit for specific scientific and decision-making contexts.
Topics include benchmarking across domains; evaluation across spatial, temporal, and variable scales; generalization across regions, datasets, and users; bias and error diagnostics; comparison with physics-based models; and reproducibility and standardized protocols.
We encourage contributions that apply common evaluation frameworks, examine failure modes and uncertainty, and connect research to operational and user needs.
This session aims to advance community standards for rigorous, interpretable, and application-relevant evaluation of ML in Earth system science.
This session focuses on the quantitative evaluation, benchmarking, and scientific interpretation of ML models across scientific domains. We emphasize frameworks that assess not only model skill, but also whether models are fit for specific scientific and decision-making contexts.
Topics include benchmarking across domains; evaluation across spatial, temporal, and variable scales; generalization across regions, datasets, and users; bias and error diagnostics; comparison with physics-based models; and reproducibility and standardized protocols.
We encourage contributions that apply common evaluation frameworks, examine failure modes and uncertainty, and connect research to operational and user needs.
This session aims to advance community standards for rigorous, interpretable, and application-relevant evaluation of ML in Earth system science.
Co-Sponsor(s):
- A - Atmospheric Sciences
- EP - Earth and Planetary Surface Processes
- H - Hydrology
- OS - Ocean Sciences
Index Terms:
1622 Earth system modeling [GLOBAL CHANGE]
1942 Machine learning [INFORMATICS]
1982 Standards [INFORMATICS]
9820 Techniques applicable in three or more fields [GENERAL OR MISCELLANEOUS]
Primary Convener: Katherine Breen, Morgan State University, Baltimore, United States
Conveners: Lesley Ott, NASA Goddard Space Flight Center, Global Modeling and Assimilation Office, Greenbelt, United States, Mark Carroll, NASA Goddard Space Flight Center, Data Science Group, Greenbelt, United States, Olya Skulovich, Columbia University, Department of Earth and Environmental Engineering, New York, United States and Sujay V Kumar, NASA Goddard Space Flight Center, Greenbelt, MD, United States
See more of: Science and Society