IN016-07
Incorporating Data Management Best Practices into Scientific Workflows

Wednesday, 9 December 2020: 20:48
Virtual
Zarine Kakalia1, Charuleka Varadharajan1, Madison Burrus1, Danielle S Christianson1, Robert Crystal-Ornelas1, Joan E Damerow1, Dipankar Dwivedi1, Boris Faybishenko1, Valerie C Hendrix1, Emily Robles1, Roelof Versteeg2, Karen Whitenack1 and Deb Agarwal1, (1)Lawrence Berkeley National Laboratory, Berkeley, CA, United States, (2)Subsurface Insights, Hanover, NH, United States
Abstract:
The U.S. Department of Energy's Watershed Function Scientific Focus Area (SFA) in the East River, Colorado generates and uses interdisciplinary data from hydrological, geochemical, geophysical, microbiological and remote sensing observations. The project has developed an end-to-end infrastructure to acquire the SFA’s multi-scale data, generate data products, and enable internal and public data access. Maintaining FAIR data throughout this pipeline is challenging due to the diversity of the data and scientific workflows. To ensure data pipelines generate integratable products and meet repository standards, the SFA Data Management Team engages with field scientists to incorporate best data management practices throughout the scientific workflow. SFA data is published through the DOE’s Environmental Systems Science Data Infrastructure for a Virtual Ecosystem (ESS-DIVE) data repository, and thus adopts standards and metadata quality criteria required for publication through ESS-DIVE.

To overcome the challenge of acquiring critical metadata from diverse data streams, the SFA Data Management Team developed an integrated field-data workflow. Field scientists are required to use persistent location identifiers for long-term sites and register field locations prior to site creation. Scientists are encouraged to use International Geo Sample Numbers (IGSNs), which are persistent identifiers for their samples that are recommended by the ESS-DIVE repository. The SFA has completed two IGSN pilot tests with ESS-DIVE to begin incorporating sample tracking into the end-to-end data pipeline. This required extensive time and education on behalf of the field team, proving that shifting scientists’ processes to curate better data requires substantial effort. Finally, scientists are asked to provide sensor data, following practices adopted by the DOE’s Ameriflux network. Datasets are reviewed and compiled internally, and final data products and the associated metadata are published on ESS-DIVE.

This integrated workflow makes it easier to apply data to downstream analysis, synthesis and models. We found that developing project data/metadata standards and workflows in line with repository requirements is an effective way to develop FAIR and transparent data practices throughout the field-data pipeline.