diff --git a/Running Start/CSB320 - Machine Learning Concepts/Class 5-28 (Clustering).md b/Running Start/CSB320 - Machine Learning Concepts/Class 5-28 (Clustering).md index bbac9e7..54a723f 100644 --- a/Running Start/CSB320 - Machine Learning Concepts/Class 5-28 (Clustering).md +++ b/Running Start/CSB320 - Machine Learning Concepts/Class 5-28 (Clustering).md @@ -17,5 +17,19 @@ - centroid - mean position of a cluster's instances $$\overline X_{i}=\frac{\sum_{j \in C_{i}}X_{i}}{n_{i}}$$ - Inertia - - aerage squared distance of the instances from the centroid $$I_{i}=\frac{\sum_{j \in C_{i}}|\overline X_{j} - X_{i}|^2}{n_{i}}$$ - - \ No newline at end of file + - average squared distance of the instances from the centroid $$I_{i}=\frac{\sum_{j \in C_{i}}|\overline X_{j} - X_{i}|^2}{n_{i}}$$ + - partitioning approach + - create various partitions and evaluate based on metric (minimize sum of squared errors) + - K-means clustering + - assigns instances to the nearest centroid + - need to know k number of clusters beforehand + - algorithm + - k points are chosen randomly as initial centroids + - assign every data point to the closest centroid + - compute new centroids with assigned data + - if centroids don't change, stop. If they do repeat with new centroids + - Choosing the optimal number of clusters + - elbow method + - graph inertia against k and find point where elbow of data is (curve levels off) + - Silhouette method + - Use silhouette coeffic \ No newline at end of file