Analysis of Seismic Data in Sumatra using Robust K-Means Clustering

Ulfasari Rafflesia, Dedi Rosadi, Devni Prima Sari, Pepi Novianti

Abstract


Indonesia is located within the Pacific Ring of Fire and frequently experiences significant seismic activities, rendering the region susceptible to hazards. Specifically, Sumatra is an island in the western part of the country, near the Eurasian and Indo-Australian tectonic plates. Over the past five years, an observable uptick in seismic events has been recorded in Sumatra. This research aimed to cluster the Sumatra region’s seismic data using the k-means algorithm and its extensions, including trimmed and robust sparse k-means, to determine the characteristics and patterns of seismic events. The k-means clustering algorithm operates effectively on many data but needs to work better in the presence of outliers. Meanwhile, the data identification reports the presence of outliers in the seismic data. The clustering analysis identified two main clusters, supported by multivariate and spatial outlier detection during preprocessing. The first cluster, encompassing 62% of seismic events, is located offshore near the Mentawai seismic gap, characterized by shallow depths (33–41 km) and magnitudes of 4.5–5.0 Ms. The second cluster, representing 28% of events, includes both mainland and offshore regions, associated with the Sumatran Fault system and slab deformation zones, at moderate depths (54–154 km) with magnitudes of 4.3–4.4 Ms. Rare deep-focus events exceeding depths of 214 km were identified as outliers. Evaluation using Silhouette, Davies-Bouldin, and Dunn indices determined that k=2 was the optimal number of clusters. This study contributes by integrating robust clustering methods to handle outliers, enhancing the reliability of seismic data analysis. This study demonstrates the value of applying trimmed and robust sparse k-means algorithms to improve clustering performance in regions with complex tectonic activity.


Keywords


K-means; Outliers; Robust; Clustering; Earthquake

Full Text:

PDF

References


M. Irsyam et al., “Development of the 2017 national seismic hazard maps of Indonesia,” Earthq. Spectra, vol. 36, no. 1_suppl, pp. 112–136, 2020, doi: 10.1177/8755293020951206.

S. J. Hutchings and W. D. Mooney, “The Seismicity of Indonesia and Tectonic Implications,” Geochemistry, Geophys. Geosystems, vol. 22, no. 9, pp. 1–42, 2021, doi: 10.1029/2021GC009812.

A. A. S. Putra, B. A. D. Nugraha, C. N. T. Puspito, and D. D. P. Sahara, “PRELIMINARY RESULT: SOURCE PARAMETERS FOR SMALL-MODERATE EARTHQUAKES IN ACEH SEGMENT, SUMATRAN FAULT ZONE (NORTHERN SUMATRA),” in 18th Annual Meeting of the Asia Oceania Geosciences Society, Aug. 2022, no. Aogs 2021, pp. 224–226. doi: 10.1142/9789811260100_0076.

P. Novianti, D. Setyorini, and U. Rafflesia, “K-means cluster analysis in earthquake epicenter clustering,” Int. J. Adv. Intell. Informatics, vol. 3, no. 2, pp. 81–89, Jul. 2017, doi: 10.26555/ijain.v3i2.100.

A. E. Ezugwu et al., “A comprehensive survey of clustering algorithms: State-of-the-art machine learning applications, taxonomy, challenges, and future research prospects,” Eng. Appl. Artif. Intell., vol. 110, no. December 2021, p. 104743, 2022, doi: 10.1016/j.engappai.2022.104743.

S. Askari, “Fuzzy C-Means clustering algorithm for data with unequal cluster sizes and contaminated with noise and outliers: Review and development,” Expert Syst. Appl., vol. 165, no. August 2020, p. 113856, 2021, doi: 10.1016/j.eswa.2020.113856.

R. Mussabayev, N. Mladenovic, B. Jarboui, and R. Mussabayev, “How to Use K-means for Big Data Clustering?,” Pattern Recognit., vol. 137, 2023, doi: 10.1016/j.patcog.2022.109269.

N. H. M. M. Shrifan, M. F. Akbar, and N. A. M. Isa, “An adaptive outlier removal aided k-means clustering algorithm,” J. King Saud Univ. - Comput. Inf. Sci., vol. 34, no. 8, pp. 6365–6376, 2022, doi: 10.1016/j.jksuci.2021.07.003.

A. M. Ikotun, A. E. Ezugwu, L. Abualigah, B. Abuhaija, and J. Heming, “K-means clustering algorithms: A comprehensive review, variants analysis, and advances in the era of big data,” Inf. Sci. (Ny)., vol. 622, pp. 178–210, 2023, doi: 10.1016/j.ins.2022.11.139.

Y. Kondo, M. Salibian-Barrera, and R. Zamar, “RSKC: An R package for a robust and sparse k-means clustering algorithm,” J. Stat. Softw., vol. 72, no. 5, 2016, doi: 10.18637/jss.v072.i05.

Š. Brodinová, P. Filzmoser, T. Ortner, C. Breiteneder, and M. Rohm, “Robust and sparse k-means clustering for high-dimensional data,” Adv. Data Anal. Classif., vol. 13, no. 4, pp. 905–932, 2019, doi: 10.1007/s11634-019-00356-9.

O. Dorabiala, J. N. Kutz, and A. Y. Aravkin, “Robust trimmed k-means,” Pattern Recognit. Lett., vol. 161, pp. 9–16, Sep. 2022, doi: 10.1016/j.patrec.2022.07.007.

Y. DING, Y. PENG, and J. LI, “Cluster Analysis of Earthquake Ground-Motion Records and Characteristic Period of Seismic Response Spectrum,” J. Earthq. Eng., vol. 24, no. 6, pp. 1012–1033, 2020, doi: 10.1080/13632469.2018.1453420.

R. Yuan, “An improved K-means clustering algorithm for global earthquake catalogs and earthquake magnitude prediction,” J. Seismol., vol. 25, no. 3, pp. 1005–1020, 2021, doi: 10.1007/s10950-021-09999-8.

S. Dhole and S. Bakre, “An updated homogeneous earthquake catalogue and earthquake recurrence parameters of Maharashtra state, an Indian stable continental region,” J. Earth Syst. Sci., vol. 133, no. 1, 2024, doi: 10.1007/s12040-023-02220-z.

A. Smiti, “A critical overview of outlier detection methods,” Comput. Sci. Rev., vol. 38, p. 100306, 2020, doi: 10.1016/j.cosrev.2020.100306.

H. Ghorbani, “Mahalanobis Distance and Its Application for,” Facta Univ., vol. 34, no. 3, pp. 583–595, 2019.

C.-T. Lu, D. Chen, and Y. Kou, “MULTIVARIATE SPATIAL OUTLIER DETECTION,” Int. J. Artif. Intell. Tools, vol. 13, no. 04, pp. 801–811, Dec. 2004, doi: 10.1142/S021821300400182X.

S. Shukla, S. Lalitha, and S. Lalitha, “Spatial Analysis of Water Quality Data Using Multivariate Spatial Outlier Detection Algorithms Spatial data analysis View project Spatial Analysis of Water Quality Data Using Multivariate Spatial Outlier Detection Algorithms,” Ganita, vol. 70, no. 2, pp. 87–96, 2021, [Online]. Available: https://www.researchgate.net/publication/369541756

S. S. Yu, S. W. Chu, C. M. Wang, Y. K. Chan, and T. C. Chang, “Two improved k-means algorithms,” Appl. Soft Comput. J., vol. 68, pp. 747–755, 2018, doi: 10.1016/j.asoc.2017.08.032.

R. Garcia-Dias, C. A. Prieto, J. S. Almeida, and I. Ordovás-Pascual, “Machine learning in APOGEE,” Astron. Astrophys., vol. 612, p. A98, Apr. 2018, doi: 10.1051/0004-6361/201732134.

E. U. Oti, M. O. Olusola, F. C. Eze, and S. U. Enogwe, “Comprehensive Review of K-Means Clustering Algorithms,” Int. J. Adv. Sci. Res. Eng., vol. 07, no. 08, pp. 64–69, 2021, doi: 10.31695/ijasre.2021.34050.

Q. Xu, Q. Zhang, J. Liu, and B. Luo, “Efficient synthetical clustering validity indexes for hierarchical clustering,” Expert Syst. Appl., vol. 151, p. 113367, Aug. 2020, doi: 10.1016/j.eswa.2020.113367.

F. Ros, R. Riad, and S. Guillaume, “PDBI: A partitioning Davies-Bouldin index for clustering evaluation,” Neurocomputing, vol. 528, pp. 178–199, 2023, doi: 10.1016/j.neucom.2023.01.043.

M. Gagolewski, M. Bartoszuk, and A. Cena, “Are cluster validity measures (in) valid?,” Inf. Sci. (Ny)., vol. 581, pp. 620–636, Dec. 2021, doi: 10.1016/j.ins.2021.10.004.

N. Wiroonsri, “Clustering performance analysis using a new correlation-based cluster validity index,” Pattern Recognit., vol. 145, p. 109910, Jan. 2024, doi: 10.1016/j.patcog.2023.109910.

G. Vardakas, I. Papakostas, and A. Likas, “Deep Clustering Using the Soft Silhouette Score: Towards Compact and Well-Separated Clusters,” 2024, [Online]. Available: http://arxiv.org/abs/2402.00608

P. J. Rousseeuw, “Silhouettes: A graphical aid to the interpretation and validation of cluster analysis,” J. Comput. Appl. Math., vol. 20, no. C, pp. 53–65, 1987, doi: 10.1016/0377-0427(87)90125-7.

D. L. Davies and D. W. Bouldin, “A Cluster Separation Measure,” IEEE Trans. Pattern Anal. Mach. Intell., vol. PAMI-1, no. 2, pp. 224–227, Apr. 1979, doi: 10.1109/TPAMI.1979.4766909.

F. Informa, W. R. Number, M. House, M. Street, and J. C. Dunn, “Well-Separated Clusters and Optimal Fuzzy Partitions,” no. June 2013, pp. 37–41, 2008, [Online]. Available: https://www.tandfonline.com/doi/abs/10.1080/01969727408546059




DOI: https://doi.org/10.47738/jads.v6i1.523

Refbacks

  • There are currently no refbacks.



Barcode

Journal of Applied Data Sciences

ISSN:2723-6471 (Online)
Publisher:Bright Publisher
Website:http://bright-journal.org/JADS
Email:taqwa@amikompurwokerto.ac.id (principal contact)
  support@bright-journal.org (technical issues)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0