Imbalanced Data Clustering using Equilibrium K-Means

He, Yudong

Full-text links:

Download:

Current browse context:

cs.LG

< prev | next >

new | recent | 2402

Computer Science > Machine Learning

Title: Imbalanced Data Clustering using Equilibrium K-Means

Authors: Yudong He

(Submitted on 22 Feb 2024 (v1), last revised 28 Mar 2024 (this version, v2))

Abstract: Traditional centroid-based clustering algorithms, such as hard K-means (HKM, or Lloyd's algorithm) and fuzzy K-means (FKM, or Bezdek's algorithm), display degraded performance when true underlying groups of data have varying sizes (i.e., imbalanced data). This paper introduces equilibrium K-means (EKM), a novel fuzzy clustering algorithm that has the robustness to imbalanced data by preventing centroids from crowding together in the center of large clusters. EKM is simple, alternating between two steps; fast, with the same time and space complexity as FKM; and scalable to large datasets. We evaluate the performance of EKM on two synthetic and ten real datasets, comparing it to other centroid-based algorithms, including HKM, FKM, maximum-entropy fuzzy clustering (MEFC), two FKM variations designed for imbalanced data, and the Gaussian mixture model. The results show that EKM performs competitively on balanced data and significantly outperforms other algorithms on imbalanced data. Deep clustering experiments on the MNIST dataset demonstrate the significance of making representation have an EKM-friendly structure when dealing with imbalanced data; In comparison to deep clustering with HKM, deep clustering with EKM obtains a more discriminative representation and a 35% improvement in clustering accuracy. Additionally, we reformulate HKM, FKM, MEFC, and EKM in a general form of gradient descent, where fuzziness is introduced differently and more simply than in Bezdek's work, and demonstrate how the general form facilitates a uniform study of KM algorithms.

Subjects:	Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as:	arXiv:2402.14490 [cs.LG]
	(or arXiv:2402.14490v2 [cs.LG] for this version)

Submission history

From: Yudong He [view email]
[v1] Thu, 22 Feb 2024 12:27:38 GMT (3964kb,D)
[v2] Thu, 28 Mar 2024 08:36:27 GMT (3964kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2402.14490

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Machine Learning

Title: Imbalanced Data Clustering using Equilibrium K-Means

Submission history