MACHINE LEARNING / 7. CLUSTERING
Clustering
Finding structure in unlabeled data
EXPLANATION
Clustering groups similar data points together without any labels. Unsupervised learning — you don't tell the algorithm what the groups are. K-Means: • Pick K centroids randomly • Assign each point to nearest centroid • Move centroids to mean of assigned points • Repeat until convergence • Problem: you must choose K, sensitive to initialization, assumes spherical clusters DBSCAN (Density-Based): • Groups points that are close together (dense regions) • Points in sparse regions = noise/outliers • No need to specify K • Handles arbitrary cluster shapes and outliers Choosing K for K-Means: • Elbow method: plot inertia vs K, pick the elbow • Silhouette score: measures how similar a point is to its cluster vs others. Range (-1, 1), higher is better Use cases: customer segmentation, document grouping, anomaly detection, data compression.
DATA FLOW
K-Means (K=3): Iteration 0: random centroids ✕ ● ● ✕ ○ ○ ✕ ■ ■ ● ● ○ ✕ ○ ○ ■ ■ ■ Iteration N (converged): centroids at cluster centers ● ● ✕ ○ ○ ■ ■ ✕ ● ● ○ ○ ✕ ○ ■ ■ ■ DBSCAN: Core point: has min_samples within eps radius Border point: within eps of core but not core itself Noise point: isolated, not within eps of any core → outlier
CODE