IN046-02
Caltech/USGS Southern California Earthquake Data Available in the Amazon Cloud (AWS)

Wednesday, 16 December 2020: 11:33
Virtual
Ellen Yu1, Shang-lin Chen1, Aparna Bhaskaran1, Rayomand Bhadha1, Zachary E. Ross1, Egill Hauksson1 and Robert W Clayton2, (1)California Institute of Technology, Seismological Laboratory, Pasadena, CA, United States, (2)California Institute of Technology, Pasadena, CA, United States
Abstract:
The Southern California Earthquake Data Center (SCEDC) has made the Caltech/USGS Southern California Seismic Network (SCSN) data archive available in the cloud as part of the Amazon Open Dataset Program. The AWS bucket name is s3://scedc-pds and it is hosted in the US West (Oregon) region. In this poster, we describe the contents of this dataset and show that cloud based archives reduce time and efforts needed for completing research as compared to traditional data gathering from a data center and research data processing.

The main contents of the SCEDC/SCSN public data set are:

  1. The SCSN event catalog (1932-present) and phase picks for these events in ascii format.
  2. Continuous recorded waveforms (1999 to present) from 603 seismic stations recorded by the SCSN. Each file contains one channel day in mSEED format. The total size of the dataset is approximately 105 Tbytes.
  3. Event-windowed waveforms (1977-present) in mSEED format.
  4. Metadata from CI stations in FDSN StationXML format

Users that process large volumes of data in ambient noise correlations, template matching, and machine-learning studies for example, will find that the I/O time is considerably reduced when the processing is done in the cloud in the same AWS region (us-west-2). I/O costs from the AWS public dataset are no-cost. We have put some simple scripts and examples at https://github.com/SCEDC/cloud that can be used as templates to get started with scientific processing. We are experimenting with indexing the waveform archive to allow more efficient sorting of the data.

The poster will also present cost estimates for a variety of research activities to give users an idea of the processing speed, ease of operating in the cloud, and costs incurred working with a cloud archive. Such costs can be compared with the costs of purchasing a computer server and a disk array, and weeks or months spent on downloading and processing data.