IN014-0007
Developing a structured seafloor sediment database from disparate datasets using SmartSearch
Abstract:
SmartSearch addresses the need for scalable, advanced data searching through a supervised, machine learning tool in both online and local data servers. In online, open-source searches, SmartSearch crawls the entire worldwide web to identify data resources relevant to a given user’s needs. Implemented using Spark and Apache Tika, SmartSearch can perform deep contextual analysis and correlate contextually similar content to automate the process of integrating complex data. The current instance of SmartSearch is leveraging Cloud capabilities from Google Cloud Platform and Azure, to help scale and optimize search and discovery for data resources. Parsing, labeling and exploration of the data is handled on local, on-premise computing assets. These local assets are also used in the on-premise data search, labeling and parsing for the previously unpublished seafloor sediment data, and to help label those data for public release on DOE’s Energy Data eXchange (EDX).
This big data search and labeling effort has produced a multi-source, -resolution, and -variate dataset, integrating information from datasets such as seafloor surveys, geophysical surveys, core, geological literature, and well logs to provide a more complete account of the spatial distribution of seafloor sediments in the northern Gulf of Mexico. This dataset is being analyzed for a geohazards detection project, and is relevant to gas hydrates, environmental surveys, geomorphic process studies and more.