IN015-05
Challenges in Assessing Data Citation and Reuse in Arctic Research Repositories

Wednesday, 9 December 2020: 17:46
Virtual
Maya Samet1,2, Jeanette Clark1,2, Matthew B. Jones2, Amber E Budden2 and Erin McLean2, (1)Arctic Data Center, Santa Barbara, CA, United States, (2)National Center for Ecological Analysis and Synthesis, Santa Barbara, CA, United States
Abstract:
As journals, funding agencies, and researchers increasingly acknowledge the importance of making data publicly available, the ability to track the impact and use of published datasets is an important step in quantifying the effect of open data practices, and imperative to the duly crediting of impactful data creators and repositories in an open science landscape. The aim of the Data Citation and Reuse project at the NSF Arctic Data Center is to produce accurate citation and research impact metrics for data housed in the Arctic Data Center that can inform researchers about data reuse, as well as to assess the culture and condition of data citation practices in the Arctic science community. The project takes a multi-pronged approach to capturing citations, including programmatic queries to publisher APIs, text mining dataset abstracts for references to related academic work, and investigative case studies of known impactful datasets. We have also developed tools to automate discovery of data citations, including the `scythe` R package. These methods capture a higher number of citations and references than existing citation aggregators do, since researchers do not consistently cite datasets in a way that is captured by these services, and publishers often do not report dataset citations in the same way as they report article citations. In this contribution we will present results comparing the number of citations of Arctic Data Center datasets captured by different methods and discuss future directions to continue capturing data citations and improving data citation practices.