IN047-09
Optimizing the Efficiency of Metadata Curation in Large Scale Data Repositories

Thursday, 17 December 2020: 04:24
Virtual
Emily Robles1, Charuleka Varadharajan1, Shreyas Cholia1, Valerie C Hendrix1, Joan E Damerow1, Madison Burrus1, Robert Crystal-Ornelas1, Hesham Elbashandy1, Zarine Kakalia1, Mario Melara2, Fianna O'Brien1, Makayla Shepherd2, Maegen Simmonds3, Karen Whitenack1, Matthew B. Jones4, Christopher S. Jones4, Peter Slaughter4 and Deb Agarwal5, (1)Lawrence Berkeley National Laboratory, Berkeley, CA, United States, (2)Lawrence Berkeley National Laboratory, Berkeley, United States, (3)University of California Davis, Davis, CA, United States, (4)National Center for Ecological Analysis and Synthesis, Santa Barbara, CA, United States, (5)LBNL, Berkeley, CA, United States
Abstract:
The Environmental System Science Data Infrastructure for a Virtual Ecosystem (ESS-DIVE) data repository stores highly diverse Earth and environmental science data generated by projects funded by the U.S. Department of Energy (DOE). A system of metadata quality standards was developed through extensive community collaboration to ensure the data submitted to ESS-DIVE remain findable, accessible, interoperable, and reproducible (FAIR) for data users. However, ongoing implementation of these checks requires a metadata review process capable of scaling with the growth of the repository as increasing emphasis is placed on the importance of data archival within the environmental sciences.

To address this challenge, ESS-DIVE created a robust data package review workflow incorporating both automated and manual checks for each data package submitted for publication. A suite of automated metadata quality FAIR checks was developed by the National Center for Ecological Analysis and Synthesis (NCEAS) and tailored to fit ESS-DIVE's needs through research into metadata best practices, review of journal metadata requirements, and community feedback. The results are compiled into Metadata Quality Reports, which provide instantaneous feedback to both the data contributor and ESS-DIVE reviewers on problem areas within the metadata. Reviewers then carry out manual checks focused on metadata content and complete post-review assessments that collect the length of time each review takes. Standardized feedback responses are generated by both series of checks and are used by the reviewer to collaborate 1:1 with contributors until all standards are met and the data package is eligible for publication.

This system has improved the quality of ESS-DIVE data while decreasing review time by ~60% from the start of implementation. The integration of automation allows our team members to focus efforts on the content-oriented manual metadata checks, which are the most commonly failed metadata requirements. Post-review assessments inform future automation efforts to continuously increase efficiency. This system of metadata review will sustain and support higher volumes of publication requests, ensuring that metadata quality standards are enforced throughout the continued growth of the ESS-DIVE repository.