Clustering Methods for Chemical Profile Segmentation
Isi Artikel Utama
Abstrak
Cluster analysis is widely used to reveal latent structure in unlabeled multivariate data, yet its practical value depends on whether the selected algorithm is compatible with the geometry, scale, and noise pattern of the data. This article develops a journal-style empirical study of clustering and its application to chemical profile segmentation using the UCI Wine Recognition data set, which contains 178 observations, 13 continuous chemical variables, and three reference cultivars. The method combines statistical preprocessing, principal component analysis, K-Means, Ward hierarchical clustering, Gaussian mixture models, spectral clustering, DBSCAN, and HDBSCAN. Performance is evaluated with internal indices, external agreement measures, cluster profiles, and visual diagnostics. The two leading principal components explain 55.41% of standardized variation, indicating a meaningful low-dimensional representation without eliminating multivariate complexity. K-Means produces the highest adjusted Rand index (0.8975) and normalized mutual information (0.8759), followed by Gaussian mixture and spectral clustering (ARI = 0.8804). Density-based methods detect possible noisy observations but recover only two dense groups and show lower agreement with reference cultivars. The findings demonstrate that clustering is not a universal algorithmic choice but a decision process linking data geometry, statistical assumptions, validation evidence, and domain interpretation. The article contributes a reproducible applied framework for selecting, validating, and explaining clustering results in chemical and quality-segmentation problems
##plugins.themes.bootstrap3.displayStats.downloads##
Rincian Artikel
Bagian
Cara Mengutip
Referensi
J. MacQueen, "Some methods for classification and analysis of multivariate observations," in Proc. Fifth Berkeley Symp. Math. Statist. Prob., vol. 1, pp. 281-297, 1967.
S. Lloyd, "Least squares quantization in PCM," IEEE Trans. Inf. Theory, vol. 28, no. 2, pp. 129-137, 1982, doi: 10.1109/TIT.1982.1056489.
J. H. Ward, "Hierarchical grouping to optimize an objective function," J. Am. Stat. Assoc., vol. 58, no. 301, pp. 236-244, 1963, doi: 10.1080/01621459.1963.10500845.
A. P. Dempster, N. M. Laird, and D. B. Rubin, "Maximum likelihood from incomplete data via the EM algorithm," J. R. Stat. Soc. Series B, vol. 39, no. 1, pp. 1-38, 1977.
J. D. Banfield and A. E. Raftery, "Model-based Gaussian and non-Gaussian clustering," Biometrics, vol. 49, no. 3, pp. 803-821, 1993, doi: 10.2307/2532201.
M. Ester, H.-P. Kriegel, J. Sander, and X. Xu, "A density-based algorithm for discovering clusters in large spatial databases with noise," in Proc. 2nd Int. Conf. Knowledge Discovery and Data Mining, pp. 226-231, 1996.
Matdoan, M. Y., Fadhilah, R., Laamena, N. S., Safira, D. A., & Loklomin, S. B. (2024). Implementation of Centroid Clustering Method for Industrial Clusterization in Regencies and Cities in Maluku Province. Pattimura International Journal of Mathematics (PIJMath), 3(1), 09-14..
A. Y. Ng, M. I. Jordan, and Y. Weiss, "On spectral clustering: Analysis and an algorithm," in Advances in Neural Information Processing Systems, vol. 14, pp. 849-856, 2002.
U. von Luxburg, "A tutorial on spectral clustering," Stat. Comput., vol. 17, pp. 395-416, 2007, doi: 10.1007/s11222-007-9033-z.
Fadhilah, R., Matdoan, M. Y., Safira, D. A., & Tahalea, S. P. (2024). Clustering Shrimp Distribution in Indonesia Using the X-Means Clustering Algorithm. VARIANCE: Journal of Statistics and Its Applications, 6(1), 49-54.
Y. Ren et al., "Deep clustering: A comprehensive survey," IEEE Trans. Neural Netw. Learn. Syst., vol. 35, no. 9, pp. 12091-12110, 2024, doi: 10.1109/TNNLS.2023.3239379.
Y. Lu, H. Li, Y. Li, Y. Lin, and X. Peng, "A survey on deep clustering: From the prior perspective," AI and Ethics, 2024, doi: 10.1007/s44336-024-00001-w.
C. C. Aggarwal, A. Hinneburg, and D. A. Keim, "On the surprising behavior of distance metrics in high dimensional space," in Database Theory - ICDT 2001, pp. 420-434, 2001, doi: 10.1007/3-540-44503-X_27.
Matdoan, M. Y., Purnamasari, N. A., & Laamena, N. S. (2023). Application of the K-Means Algorithm for Clustering Production of Capture Fisheries in Maluku Province. Pattimura International Journal of Mathematics (PIJMath), 2(2), 63-70.
P. J. Rousseeuw, "Silhouettes: A graphical aid to the interpretation and validation of cluster analysis," J. Comput. Appl. Math., vol. 20, pp. 53-65, 1987, doi: 10.1016/0377-0427(87)90125-7.
D. L. Davies and D. W. Bouldin, "A cluster separation measure," IEEE Trans. Pattern Anal. Mach. Intell., vol. PAMI-1, no. 2, pp. 224-227, 1979, doi: 10.1109/TPAMI.1979.4766909.
T. Caliński and J. Harabasz, "A dendrite method for cluster analysis," Commun. Stat., vol. 3, no. 1, pp. 1-27, 1974, doi: 10.1080/03610927408827101.
L. Hubert and P. Arabie, "Comparing partitions," J. Classif., vol. 2, pp. 193-218, 1985, doi: 10.1007/BF01908075.
N. X. Vinh, J. Epps, and J. Bailey, "Information theoretic measures for clusterings comparison: Variants, properties, normalization and correction for chance," J. Mach. Learn. Res., vol. 11, pp. 2837-2854, 2010.
S. Aeberhard, D. Coomans, and O. de Vel, "Comparison of classifiers in high dimensional settings," Tech. Rep. no. 92-02, Dept. Computer Science and Dept. Mathematics and Statistics, James Cook University of North Queensland, 1992.
D. Dua and C. Graff, "UCI Machine Learning Repository," University of California, Irvine, School of Information and Computer Sciences, 2019. [Online]. Available: https://archive.ics.uci.edu/ml
F. Pedregosa et al., "Scikit-learn: Machine learning in Python," J. Mach. Learn. Res., vol. 12, pp. 2825-2830, 2011.
Implementasi metode Ward clustering untuk klasterisasi penduduk yang memiliki jaminan kesehatan pada kabupaten dan kota di Provinsi Maluku
P. Berkhin, "A survey of clustering data mining techniques," in Grouping Multidimensional Data, Springer, pp. 25-71, 2006, doi: 10.1007/3-540-28349-8_2.
L. McInnes, J. Healy, and J. Melville, "UMAP: Uniform manifold approximation and projection for dimension reduction," arXiv:1802.03426, 2018.
L. van der Maaten and G. Hinton, "Visualizing data using t-SNE," J. Mach. Learn. Res., vol. 9, pp. 2579-2605, 2008.
M. Ankerst, M. M. Breunig, H.-P. Kriegel, and J. Sander, "OPTICS: Ordering points to identify the clustering structure," in Proc. ACM SIGMOD Int. Conf. Management of Data, pp. 49-60, 1999, doi: 10.1145/304182.304187.
L. McInnes, J. Healy, and S. Astels, "hdbscan: Hierarchical density based clustering," J. Open Source Softw., vol. 2, no. 11, p. 205, 2017, doi: 10.21105/joss.00205.
T. Hastie, R. Tibshirani, and J. Friedman, The Elements of Statistical Learning: Data Mining, Inference, and Prediction, 2nd ed. New York, NY, USA: Springer, 2009.
G. James, D. Witten, T. Hastie, R. Tibshirani, and J. Taylor, An Introduction to Statistical Learning with Applications in Python. Cham, Switzerland: Springer, 2023, doi: 10.1007/978-3-031-38747-0.
D. Arthur and S. Vassilvitskii, "k-means++: The advantages of careful seeding," in Proc. 18th Annu. ACM-SIAM Symp. Discrete Algorithms, pp. 1027-1035, 2007.
G. Schwarz, "Estimating the dimension of a model," Ann. Statist., vol. 6, no. 2, pp. 461-464, 1978, doi: 10.1214/aos/1176344136.
C. Fraley and A. E. Raftery, "Model-based clustering, discriminant analysis, and density estimation," J. Am. Stat. Assoc., vol. 97, no. 458, pp. 611-631, 2002, doi: 10.1198/016214502760047131.
A. Strehl and J. Ghosh, "Cluster ensembles - A knowledge reuse framework for combining multiple partitions," J. Mach. Learn. Res., vol. 3, pp. 583-617, 2002.
Matdoan, M. Y., Ahsan, M., Wance, M., & Nukuhaly, N. A. (2023, January). Classification of provinces based on the Indonesian Democracy Index using the K-medoids clustering algorithm. In AIP Conference Proceedings (Vol. 2588, No. 1, p. 050024). AIP Publishing LLC.
M. Hahsler, M. Piekenbrock, and D. Doran, "dbscan: Fast density-based clustering with R," J. Stat. Softw., vol. 91, no. 1, pp. 1-30, 2019, doi: 10.18637/jss.v091.i01.
P. J. Rousseeuw and M. Hubert, "Anomaly detection by robust statistics," Wiley Interdiscip. Rev. Data Min. Knowl. Discov., vol. 8, no. 2, e1236, 2018, doi: 10.1002/widm.1236.
S. Zhou et al., "A comprehensive survey on deep clustering: Taxonomy, challenges, and future directions," arXiv:2206.07579, 2022.
M. Meila, "Comparing clusterings by the variation of information," in Learning Theory and Kernel Machines, pp. 173-187, 2003, doi: 10.1007/978-3-540-45167-9_14.
A. Saxena et al., "A review of clustering techniques and developments," Neurocomputing, vol. 267, pp. 664-681, 2017, doi: 10.1016/j.neucom.2017.06.053.