K-Cube Consensus Clustering with Centroid Improvement and Variance-Based Metrics on High-Dimensional Data
Abstract
High-dimensional and multidimensional cube data structures (K-Cube) are posing a significant challenge for conventional clustering algorithms due to the effect of dimensionality, uniform feature weight assumptions, and loss of hierarchical information. Therefore, this study aimed to propose K-Cube Consensus Clustering framework, which integrates Variance-Based Centroid Refinement, Weighted Distance Metrics, and consensus voting mechanism to overcome the challenges of high-dimensional cube data. The proposed method systematically clustered all dimensions and sub-dimensions of cube data, refined centroid by emphasizing more stable low-variance attributes, and applied adaptive distance weighting based on variance-derived feature weights integrated into the distance metric to improve cluster assignment. The final clusters were obtained through majority voting of the clustering results for each dimension. Unlike existing consensus clustering methods that operate on flat data representations or combine independent clustering results, the proposed framework explicitly exploits the hierarchical structure of multidimensional cube data by clustering dimensions and sub-dimensions prior to consensus integration. Moreover, variance-based centroid refinement and weighted distance metrics are jointly embedded within each cube dimension rather than applied as isolated enhancements. This hierarchy-aware design preserves cube semantics while simultaneously improving centroid stability and distance adaptivity, resulting in a distinct and scalable clustering framework for complex high-dimensional cube data. The framework processes cube dimensions independently with iterative convergence control, enabling scalable application to large-scale cube data. The results of synthetic and real-world high-dimensional datasets, including cube data with approximately 2.2 million instances, showed that the proposed method consistently outperformed K-Means, K-Medoids, and Hamiltonian formulations. The method produced lower SSE such as 3,179,328 on Arcene and 1,422.21 on Lung Cancer, higher Silhouette Score of approximately 0.5718 and 0.4905 for consensus results, better cluster stability of 0.9947, and faster convergence. These results confirmed the effectiveness of K-Cube Consensus Clustering in producing stable and meaningful clusters in large-scale high-dimensional data applications.
Keywords
Full Text:
PDFReferences
X. Chen, Y. Ye, X. Xu, and J. Z. Huang, “A Feature Group Weighting Method for Subspace Clustering of High‐Dimensional Data,” Pattern Recognit., vol. 45, no. 1, pp. 434–446, 2012, doi: 10.1016/j.patcog.2011.06.004.
R. R. Isnanto and B. Warsito, “Applied Data Science and Artificial Intelligence for Tourism and Hospitality Industry in Society 5 . 0 : A Review,” J. Appl. Data Sci., vol. 5, no. 4, pp. 1566–1578, 2024, doi: 10.47738/jads.v5i4.300.
W. Qu, X. Xiu, H. Chen, and L. Kong, “A Survey on High-Dimensional Subspace Clustering A Survey on High-Dimensional Subspace Clustering,” Mathematics, vol. 11, no. 2, pp. 1–39, 2023, doi: 10.3390/math11020436.
S. Chakraborty, P. Dey, D. Das, N. Shah, H. Wang, and S. Sra, “Clustering High-dimensional Data with Ordered Weighted Regularization,” in Proceedings of The 26th International Conference on Artificial Intelligence and Statistics (AISTATS), PMLR, 2023, pp. 10483–10496. [Online]. Available: https://proceedings.mlr.press/v206/chakraborty23a.html
X. Du, “A Robust and High-Dimensional Clustering Algorithm Based on Feature Weight and Entropy,” entropy, vol. 245, no. 3, pp. 1–7, 2023, doi: 10.3390/e25030510.
A. G. Oskouei et al., “Feature-weighted fuzzy clustering methods: An experimental review,” Neurocomputing, vol. 619, p. 129176, 2024, doi: 10.1016/j.neucom.2024.129176.
J. Leprince, C. Miller, and W. Zeiler, “Data mining cubes for buildings: a generic framework for multidimensional analytics of building performance data,” Energy Build., vol. 248, p. 111195, 2021, doi: 10.1016/j.enbuild.2021.111195.
M. Ahmed, R. Seraj, and S. M. S. Islam, “The k-means algorithm: A comprehensive survey and performance evaluation,” Electronics, vol. 9, no. 8, p. 1295, 2020, doi: 10.3390/electronics9081295.
A. M. Ikotun, A. E. Ezugwu, O. N. Oyelade, and others, “K-means clustering algorithms: A comprehensive review, variants and their applications,” Inf. Sci. (Ny)., vol. 622, pp. 178–210, 2023, doi: 10.1016/j.ins.2022.11.139.
J. Zamora and J. Sublime, “An Ensemble and Multi-View Clustering Method Based on Kolmogorov Complexity,” Entropy, vol. 25, no. 2, p. 371, 2023, doi: 10.3390/e25020371.
X. Zhao, X. Niu, Y. Ma, and J. Zhang, “A Multi-View Ensemble Clustering Approach Using Joint Entropy,” Expert Syst. Appl., vol. 255, p. 124683, 2024, doi: 10.1016/j.eswa.2024.124683.
Z. Liu, H. Qiu, T. Senapati, M. Lin, L. Abualigah, and M. Deveci, “Enhancements of Evidential C-Means Algorithms: A Clustering Framework via Feature-Weight Learning,” Expert Syst. Appl., vol. 259, p. 125246, 2025, doi: 10.1016/j.eswa.2024.125246.
F. Galis and D. Onchis, “Refining Filter Global Feature Weighting for Fully-Unsupervised Clustering,” 2025. [Online]. Available: https://arxiv.org/abs/2503.11277
Z. Ning, Z. Dai, H. Zhang, Y. Chen, and Z. Yuan, “A Clustering Method for Small scRNA-seq Data Based on Subspace and Weighted Distance (SSWD),” PeerJ, p. e14706, 2023, doi: 10.7717/peerj.14706.
Z. Liu, H. Qiu, T. Senapati, M. Lin, L. Abualigah, and M. Deveci, “A Clustering Framework via Feature-Weight Learning,” Appl. Soft Comput., vol. 150, p. 111015, 2025, doi: 10.1016/j.asoc.2023.111015.
W. Qu, Y. Li, B. Liu, and Z. Zhang, “A survey on high-dimensional subspace clustering,” Mathematics, vol. 11, no. 2, p. 436, 2023, doi: 10.3390/math11020436.
T. ALASALI and Y. ORTAKCI, “Clustering Techniques in Data Mining: A Survey of Methods, Challenges, and Applications,” Comput. Sci., no. 1, pp. 32–50, 2024, doi: 10.53070/bbd.1421527.
C. C. Aggarwal, A. Hinneburg, and D. A. Keim, “On the surprising behavior of distance metrics in high dimensional space,” in Lecture Notes in Computer Science, 2001, pp. 420–434. doi: 10.1007/3-540-44794-6_27.
I. Hussain et al., “Weighted Multiview K-Means Clustering with L2 Regularization,” Symmetry (Basel)., vol. 16, no. 12, p. 1646, 2024, doi: 10.3390/sym16121646.
L. Liu, Y. Jiang, and X. He, “High-dimensional data clustering: A survey on subspace clustering, pattern-based clustering, and correlation clustering,” ACM Trans. Knowl. Discov. from Data, vol. 14, no. 3, pp. 1–34, 2020, doi: 10.1145/1497577.1497578.
A. Cuzzocrea, I.-Y. Song, and K. C. Davis, “Analytics over big multidimensional data: The big cube analytics,” Inf. Syst., vol. 103, p. 101706, 2021, doi: 10.1016/j.is.2021.101706.
M. Sabri et al., “A Novel Classification Algorithm Based on the Synergy Between Dynamic Clustering with Adaptive Distances and K-Nearest Neighbors,” J. Classif., vol. 41, pp. 264–288, 2024, doi: 10.1007/s00357-024-09471-5.
X. Du and others, “A robust and high-dimensional clustering algorithm based on non-Euclidean distance combining feature weights and entropy weights,” Entropy, vol. 25, no. 3, p. 510, 2023, doi: 10.3390/e25030510.
W. Li, P. Chen, and others, “An Ensemble Clustering Framework Based on Hierarchical Clustering Ensemble Selection,” Int. J. Pattern Recognit. Artif. Intell., vol. 54, no. 5, pp. 741–766, 2023, doi: 10.1080/01969722.2022.2073704.
M. Filippone, F. Camastra, F. Masulli, and S. Rovetta, “A survey of clustering algorithms for big data: Taxonomy and empirical analysis,” ACM Comput. Surv., vol. 54, no. 7, pp. 1–36, 2021, doi: 10.1109/TETC.2014.2330519.
W. Authors, “WeDIV – An improved k-means clustering algorithm with a weighted distance and a novel internal validation index,” Egypt. Informatics J., vol. 23, no. 4, pp. 133–144, 2022, doi: 10.1016/j.eij.2022.09.002.
J. Sun, Y. Zhang, and B. Liu, “Adaptive distance metrics for improved clustering in high-dimensional spaces,” Knowledge-Based Syst., vol. 264, p. 110229, 2023, doi: 10.1016/j.knosys.2023.110229.
Q. Wang, X. Li, and Y. Wang, “Multi-view clustering via graph-based consensus learning for high-dimensional data,” Knowledge-Based Syst., vol. 265, p. 110182, 2023, doi: 10.1016/j.knosys.2023.110182.
Henderi, T. Wahyuningsih, and E. Rahwanto, “Comparison of Min-Max normalization and Z-Score Normalization in the K-nearest neighbor ( kNN ) Algorithm to Test the Accuracy of Types of Breast Cancer,” IJIIS Int. J. Informatics Inf. Syst., vol. 4, no. 1, pp. 13–20, 2021, doi: 10.47738/ijiis.v4i1.73.
DOI: https://doi.org/10.47738/jads.v7i2.1209
Refbacks
- There are currently no refbacks.

Journal of Applied Data Sciences
| ISSN | : | 2723-6471 (Online) |
| Publisher | : | Bright Publisher |
| Website | : | http://bright-journal.org/JADS |
| : | taqwa@amikompurwokerto.ac.id (principal contact) | |
| support@bright-journal.org (technical issues) |
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0




.png)