vault backup: 2026-05-28 18:55:41

This commit is contained in:
ben committed 2026-05-28 18:55:41 -07:00
1 parent 5d5a107b0f
commit db546b208d
1 file changed
+16 -2
@@ -17,5 +17,19 @@
- centroid
- mean position of a cluster's instances $$\overline X_{i}=\frac{\sum_{j \in C_{i}}X_{i}}{n_{i}}$$
- Inertia
- aerage squared distance of the instances from the centroid $$I_{i}=\frac{\sum_{j \in C_{i}}|\overline X_{j} - X_{i}|^2}{n_{i}}$$
-
- average squared distance of the instances from the centroid $$I_{i}=\frac{\sum_{j \in C_{i}}|\overline X_{j} - X_{i}|^2}{n_{i}}$$
- partitioning approach
- create various partitions and evaluate based on metric (minimize sum of squared errors)
- K-means clustering
- assigns instances to the nearest centroid
- need to know k number of clusters beforehand
- algorithm
- k points are chosen randomly as initial centroids
- assign every data point to the closest centroid
- compute new centroids with assigned data
- if centroids don't change, stop. If they do repeat with new centroids
- Choosing the optimal number of clusters
- elbow method
- graph inertia against k and find point where elbow of data is (curve levels off)
- Silhouette method
- Use silhouette coeffic