S051-06
An unsupervised automatic classification for continuous seismic records: introducing an anomaly detection algorithm to solve the imbalanced data problem

Tuesday, 15 December 2020: 04:22
Virtual
Yuki Kodera, Meteorological Research Institute, Japan Meteorological Agency, Tsukuba, Japan and Shin'ichi Sakai, Univ Tokyo, Bunkyo-Ku Tokyo, Japan
Abstract:
Continuous seismic waveforms record various kinds of signals such as earthquakes, human activities, and instrumental noises. An automatic classification algorithm for continuous records would allow us to understand geophysical phenomena around the target seismometers. The classification technique could be also used for automatic monitoring of seismometers incorporated in a real-time processing. In this study, we have developed an unsupervised continuous waveform classification method applicable to various seismometers deployed in different observation environments. Our method uses running spectra as features and classifies waveforms by the k-means clustering algorithm in the frequency domain and the spectral clustering algorithm in the time domain.

A considerable challenge in the automatic classification for continuous records is that the dataset is imbalanced (e.g., He and Garcia, 2009); that is, a major part of the waveforms consists of stationary signals (e.g., background noises), and non-stationary signals (e.g., earthquakes) are relatively rare. Without solving the imbalanced data problem, it would be possible for classification algorithms to pay attention only to stationary signals and classify them into excessively small fragments, neglecting non-stationary signals. We therefore introduced an anomaly detection algorithm based on a sampling technique (Sugiyama and Borgwardt, 2013). Before clustering data in the frequency domain, the anomaly detection algorithm extracts inliers and outliners (1.0% of the dataset). Then, the k-means clustering algorithm is applied to those extracted data points, not the entire dataset.

We tested the proposed algorithm using a continuous waveform recorded at station E.JDJM of MeSO-net (Kawakita and Sakai, 2009; a seismometer near a subway) from March 1 to 7, 2017. When applied without the anomaly detection process, the algorithm subdivided train noises into several different classes, which was due to the imbalanced data problem; on the other hand, the algorithm with the anomaly detection successfully classified the train noises into a single unique class. This result indicates that introducing the anomaly detection process is an effective measure to solve the imbalanced data problem occurring with continuous waveform classification.