IN014-0007
Developing a structured seafloor sediment database from disparate datasets using SmartSearch

Wednesday, 9 December 2020
Poster
MacKenzie Mark-Moser, Leidos Research Support Team, Albany, OR, United States, Kelly Rose, National Energy Technology Laboratory Albany, Albany, OR, United States and David Vic Baker, Matric Mid-Atlantic Technology, Research & Innovation Center, Morgantown, WV, United States
Abstract:
Disparate collections of datasets that contain information characterizing the seafloor and subsurface can be systematically categorized and analyzed using artificial intelligence tools. We present a data and knowledge discovery method, using a custom, big-data, NLP-driven algorithm, SmartSearch, and demonstrate the results of its use to rapidly identify, parse, label, and explore datasets containing seafloor sediment information in the northern Gulf of Mexico. This effort largely focuses on the exploration of open-source, publicly available data exploration and discovery, and includes integration of previously unpublished data from the authors as well. These tools employ machine learning, including deep learning techniques, to identify relevant data from a variety of data formats and resources, including spatial and nonspatial.

SmartSearch addresses the need for scalable, advanced data searching through a supervised, machine learning tool in both online and local data servers. In online, open-source searches, SmartSearch crawls the entire worldwide web to identify data resources relevant to a given user’s needs. Implemented using Spark and Apache Tika, SmartSearch can perform deep contextual analysis and correlate contextually similar content to automate the process of integrating complex data. The current instance of SmartSearch is leveraging Cloud capabilities from Google Cloud Platform and Azure, to help scale and optimize search and discovery for data resources. Parsing, labeling and exploration of the data is handled on local, on-premise computing assets. These local assets are also used in the on-premise data search, labeling and parsing for the previously unpublished seafloor sediment data, and to help label those data for public release on DOE’s Energy Data eXchange (EDX).

This big data search and labeling effort has produced a multi-source, -resolution, and -variate dataset, integrating information from datasets such as seafloor surveys, geophysical surveys, core, geological literature, and well logs to provide a more complete account of the spatial distribution of seafloor sediments in the northern Gulf of Mexico. This dataset is being analyzed for a geohazards detection project, and is relevant to gas hydrates, environmental surveys, geomorphic process studies and more.