Files
ObsidianVault/Running Start/CSB320 - Machine Learning Concepts/Class 5-28 (Clustering).md
T
2026-05-28 19:15:48 -07:00

1.9 KiB

#rs/class/csb320 #rs/notes


  • Everything so far has been supervised learning (mostly)
  • unsupervised learning (clustering)
    • do not know the labels of data.
  • clustering
    • groups instances based on feature similarity
    • results in group assignments, not target output
    • need to extrapolate meaning from clusters
    • !ClusteringBasic.excalidraw
    • applications
      • taxonomy of living things
      • clustering documents on topic
      • identify areas with similar land use
      • cluster groups of houses for city planning
    • good clustering: high intra-class similarity, low inter-class similarity
    • centroid
      • mean position of a cluster's instances \overline X_{i}=\frac{\sum_{j \in C_{i}}X_{i}}{n_{i}}
    • Inertia
      • average squared distance of the instances from the centroid I_{i}=\frac{\sum_{j \in C_{i}}|\overline X_{j} - X_{i}|^2}{n_{i}}
    • partitioning approach
      • create various partitions and evaluate based on metric (minimize sum of squared errors)
    • K-means clustering
      • assigns instances to the nearest centroid
      • need to know k number of clusters beforehand
      • algorithm
        • k points are chosen randomly as initial centroids
        • assign every data point to the closest centroid
        • compute new centroids with assigned data
        • if centroids don't change, stop. If they do repeat with new centroids
      • Choosing the optimal number of clusters
        • elbow method
          • graph inertia against k and find point where elbow of data is (curve levels off)
        • Silhouette method
          • Use silhouette coefficient to determine separation of clusters S(j) = \frac{{\overline{d_{out}(j)} - \overline{d_{in}(j)}}}{max(\overline{d_{out}(j)}, \overline{d_{in}(j)})}
          • \overline{d_{out}(j)} is the average distance of the instance i to the centroid of all other clusters
          • \overline{d_{in}(j)} is the average distance of the instance i to all other instances in its cluster
          • !SilhouetteMethod.excalidraw