From db546b208d483b9c98f67f3d34ece9cdab31d400 Mon Sep 17 00:00:00 2001 From: ben-jaynes <1btjaynes@gmail.com> Date: Thu, 28 May 2026 18:55:41 -0700 Subject: [PATCH] vault backup: 2026-05-28 18:55:41 --- .../Class 5-28 (Clustering).md | 18 ++++++++++++++++-- 1 file changed, 16 insertions(+), 2 deletions(-) diff --git a/Running Start/CSB320 - Machine Learning Concepts/Class 5-28 (Clustering).md b/Running Start/CSB320 - Machine Learning Concepts/Class 5-28 (Clustering).md index bbac9e7..54a723f 100644 --- a/Running Start/CSB320 - Machine Learning Concepts/Class 5-28 (Clustering).md +++ b/Running Start/CSB320 - Machine Learning Concepts/Class 5-28 (Clustering).md @@ -17,5 +17,19 @@ - centroid - mean position of a cluster's instances $$\overline X_{i}=\frac{\sum_{j \in C_{i}}X_{i}}{n_{i}}$$ - Inertia - - aerage squared distance of the instances from the centroid $$I_{i}=\frac{\sum_{j \in C_{i}}|\overline X_{j} - X_{i}|^2}{n_{i}}$$ - - \ No newline at end of file + - average squared distance of the instances from the centroid $$I_{i}=\frac{\sum_{j \in C_{i}}|\overline X_{j} - X_{i}|^2}{n_{i}}$$ + - partitioning approach + - create various partitions and evaluate based on metric (minimize sum of squared errors) + - K-means clustering + - assigns instances to the nearest centroid + - need to know k number of clusters beforehand + - algorithm + - k points are chosen randomly as initial centroids + - assign every data point to the closest centroid + - compute new centroids with assigned data + - if centroids don't change, stop. If they do repeat with new centroids + - Choosing the optimal number of clusters + - elbow method + - graph inertia against k and find point where elbow of data is (curve levels off) + - Silhouette method + - Use silhouette coeffic \ No newline at end of file