IN003-03
Designing for 2030: The Impact and Potential of Virtual Laboratories

Monday, 7 December 2020: 07:06
Virtual
Shweta Narkar, Brenda L Thomson and Peter Arthur Fox, Rensselaer Polytechnic Institute, Tetherless World Constellation, Troy, NY, United States
Abstract:
Since the early 2000s, virtual observatories have grown in quantity and scientific impact. Virtual observatories are research environments that provide data and software tools to facilitate research over the internet. An example is the Deep Carbon Observatory (DCO) (www.deepcarbon.net) developed in 2009 using Drupal, an open-source web content management framework, Comprehensive Knowledge Archive Network (CKAN) ,an open-source open-data portal, the then-nascent VIVO, a semantic web application for data-driven profiles, and the Global Handle Registry. Underneath this architecture, DCO implemented a knowledge graph that brought the human-human knowledge network to a virtual global network. The underlying knowledge graph also provides scalability to the project. DCO’s continued growth led to it becoming a virtual laboratory by implementing Jupyter Notebooks(www.jupyter.org), which provides scientists with an interactive, collaborative research environment to design and run experiments. DCO currently has over 1,200 scientists from 55 countries and has produced almost 2,800 publications in its initial decade.

A key lesson from DCO is that cyberinfrastructure design requires more than consideration of the amount of data; it requires careful consideration of data structures, data quality, and, since culture eats strategy for breakfast, the people involved. Building upon our experience with DCO and in preparation for Earth and space science exascale computational needs for 2030, we propose a discussion of how the impact of provenance, interoperability, completeness of data, and scalability inform and shape cyberinfrastructure. We also touch on other considerations such as data formats, data availability, and data management.