Files
ObsidianVault/Running Start/CSB320 - Machine Learning Concepts/Class 5-28 (Clustering).md
T
2026-05-28 19:46:06 -07:00

2.9 KiB

#rs/class/csb320 #rs/notes


  • Everything so far has been supervised learning (mostly)
  • unsupervised learning (clustering)
    • do not know the labels of data.
  • clustering
    • groups instances based on feature similarity
    • results in group assignments, not target output
    • need to extrapolate meaning from clusters
    • !ClusteringBasic.excalidraw
    • applications
      • taxonomy of living things
      • clustering documents on topic
      • identify areas with similar land use
      • cluster groups of houses for city planning
    • good clustering: high intra-class similarity, low inter-class similarity
    • centroid
      • mean position of a cluster's instances \overline X_{i}=\frac{\sum_{j \in C_{i}}X_{i}}{n_{i}}
    • Inertia
      • average squared distance of the instances from the centroid I_{i}=\frac{\sum_{j \in C_{i}}|\overline X_{j} - X_{i}|^2}{n_{i}}
    • partitioning approach
      • create various partitions and evaluate based on metric (minimize sum of squared errors)
    • K-means clustering
      • assigns instances to the nearest centroid
      • need to know k number of clusters beforehand
      • algorithm
        • k points are chosen randomly as initial centroids
        • assign every data point to the closest centroid
        • compute new centroids with assigned data
        • if centroids don't change, stop. If they do repeat with new centroids
      • Choosing the optimal number of clusters
        • elbow method
          • graph inertia against k and find point where elbow of data is (curve levels off)
        • Silhouette method
          • Use silhouette coefficient to determine separation of clusters S(j) = \frac{{\overline{d_{out}(j)} - \overline{d_{in}(j)}}}{max(\overline{d_{out}(j)}, \overline{d_{in}(j)})}
          • \overline{d_{out}(j)} is the average distance of the instance i to the centroid of all other clusters
          • \overline{d_{in}(j)} is the average distance of the instance i to all other instances in its cluster
          • !SilhouetteMethod.excalidraw
      • effective for small/medium datasets
      • assumes spherical clusters
      • sensitive to initial conditions
      • equal-sized groups assumption
      • sensitive to outliers
        • k-medioids can be used
          • medioid is the most centrally located point in a cluster
          • PAM: partitioning around medioids
    • Hierarchical clustering
      • does not require number of clusters, requires stopping point
      • Agglomerative
        • starts with each instance as cluster and merges into larger clusters
        • measures distances between all clusters and merges closest two into new cluster
      • divisive
        • starts with one cluster and splits into smaller clusters
      • linkage methods
        • measure distance between two closes instances
        • measure distance between two farthest instances
        • distance between each cluster's centroid
      • dendrograms
        • used to visualize clustering hierarchy
        • basically just a cladogram
          • clades, links, leaves
      • does not assume spherical clusters
      • sensitive to outliers