IN008-03
The ESS-DIVE repository and next steps toward a usable, trusted, and FAIR repository

Tuesday, 8 December 2020: 10:36
Virtual
Deb Agarwal1, Shreyas Cholia1, Charuleka Varadharajan1, Valerie C Hendrix1, Joan E Damerow1, Madison Burrus1, Robert Crystal-Ornelas1, Hesham Elbashandy1, Emily Robles1, Fianna O'Brien1, Zarine Kakalia1, Mario Melara2, William J Riley1, Cory Snavely3, Makayla Shepherd2, Maegen Simmonds4, Karen Whitenack1, Matthew B. Jones5, Christopher S. Jones5 and Peter Slaughter5, (1)Lawrence Berkeley National Laboratory, Berkeley, CA, United States, (2)Lawrence Berkeley National Laboratory, Berkeley, United States, (3)Lawrence Berkeley National Laboratory, NERSC, Berkeley, CA, United States, (4)University of California Davis, Davis, CA, United States, (5)National Center for Ecological Analysis and Synthesis, Santa Barbara, CA, United States
Abstract:
The US Department of Energy’s (DOE) Environmental Systems Science Data Infrastructure for a Virtual Ecosystem (ESS-DIVE) data repository is in its third year of operation. The repository focus is on three areas of development: expanding adoption and use by ESS Users, standardization of data, and support for projects providing data to the repository. Our approach is designed around user experience methods and involves significant discussion and involvement of the community. The priorities of the repository are continually revised and refined based on input from the community.

Our current focus is on expanding the user-base and functionality of ESS-DIVE through five key innovations: (1) understand user needs; (2) support for early data archiving by projects; (3) reaching a broader portion of the ESS community; (4) support search of extracted ESS-DIVE data with a fusion database; and (5) federation with other repositories. We are focused on providing a scalable, robust repository and long-term curation of ESS data that adhere to Findable, Accessible, Interoperable, and Reusable (FAIR) principles, with the goal of increasing the ease and capacity of storing data in the repository. A key goal is enhancing usability of the data. For example, many of the projects contributing data to ESS-DIVE have large teams, last many years, and generate a large number of data packages. We are working with our community to evaluate the available methods of providing usable citations for large subsets of the data from a project.

Our end goal is to have a repository that is trusted by the community and that is the preferred storage facility for data generated by the DOE ESS program and the preferred provider of ESS data. One challenge is that FAIR principles are designed to address the needs of the data user, and largely ignore the needs of the data provider. The CoreTrustSeal is not yet well known so there is no pressure from our user community or funders to complete the application process. However, now that at least one repository based on the same software, MetaCat, has been certified the process might be less work for ESS-DIVE. As publishers move to require CoreTrustSeal certification, we expect to see increased pressure to obtain the certification.