IN047-04
Risks of Automated Data Discovery: Recommendations for Data Stewardship in Data-Centric Climate Science

Thursday, 17 December 2020: 04:09
Virtual
Stuart Murray Gluck, Indiana University Bloomington, Bloomington, IN, United States, Elisabeth A Lloyd, Indiana University Bloomington, History and Philosophy of Science and Medicine, Bloomington, IN, United States and Greg Lusk, Michigan State University, Philosophy, East Lansing, MI, United States
Abstract:
To understand data-centrism in climate science and develop recommendations for best practices moving forward, we compared the production, management, and dissemination of big data and associated metadata in climate science to that in model organism biology. Our interdisciplinary team of philosophers and data and climate scientists embedded with a regional-climate modeling group at NCAR to analyze both community practices and the underlying information architectures of data systems. We found that the forms of data-centrism in these two fields are quite distinct. In biology, the primary goal is integration of data for use across heterogeneous subfields, and curators employ ontologies to classify data into stable, large-scale phenomena to allow effective searches. In climate science, the primary goal is overcoming the challenges of massive datasets to share crucial data within an ecosystem of similarly trained researchers, and modeling teams prepare and share data without ontologizing it and labeling phenomena (e.g., atmospheric rivers, hurricanes) within it. Users rely on shared expertise rather than labels of “phenomena-terms” to evaluate the applicability of datasets. Whereas large consortia develop and maintain complex relational databases for dissemination of datasets in biology, in climate science the modeling teams internally prepare netCDF files, with its lean architecture supporting computational tractability. Today’s data-centrism in climate science thus avoids the inadvertent, and potentially pernicious, steering of downstream research by data curators that can arise in model organism biology. However, the looming adoption of machine learning models to automate identification of phenomena in RCM output datasets could introduce such unintended steering. We provide recommendations for best practices moving forward, including especially for ameliorating the risks associated with adoption of big data techniques such as machine learning and AI.