S051-06
An unsupervised automatic classification for continuous seismic records: introducing an anomaly detection algorithm to solve the imbalanced data problem
Abstract:
A considerable challenge in the automatic classification for continuous records is that the dataset is imbalanced (e.g., He and Garcia, 2009); that is, a major part of the waveforms consists of stationary signals (e.g., background noises), and non-stationary signals (e.g., earthquakes) are relatively rare. Without solving the imbalanced data problem, it would be possible for classification algorithms to pay attention only to stationary signals and classify them into excessively small fragments, neglecting non-stationary signals. We therefore introduced an anomaly detection algorithm based on a sampling technique (Sugiyama and Borgwardt, 2013). Before clustering data in the frequency domain, the anomaly detection algorithm extracts inliers and outliners (1.0% of the dataset). Then, the k-means clustering algorithm is applied to those extracted data points, not the entire dataset.
We tested the proposed algorithm using a continuous waveform recorded at station E.JDJM of MeSO-net (Kawakita and Sakai, 2009; a seismometer near a subway) from March 1 to 7, 2017. When applied without the anomaly detection process, the algorithm subdivided train noises into several different classes, which was due to the imbalanced data problem; on the other hand, the algorithm with the anomaly detection successfully classified the train noises into a single unique class. This result indicates that introducing the anomaly detection process is an effective measure to solve the imbalanced data problem occurring with continuous waveform classification.