#rs/class/csb320 #rs/notes - - - - Everything so far has been supervised learning (mostly) - unsupervised learning (clustering) - do not know the labels of data. - clustering - groups instances based on feature similarity - results in group assignments, not target output - need to extrapolate meaning from clusters - ![[ClusteringBasic.excalidraw]] - applications - taxonomy of living things - clustering documents on topic - identify areas with similar land use - cluster groups of houses for city planning - good clustering: high intra-class similarity, low inter-class similarity - centroid - mean position of a cluster's instances $$\overline X_{i}=\frac{\sum_{j \in C_{i}}X_{i}}{n_{i}}$$ - Inertia - average squared distance of the instances from the centroid $$I_{i}=\frac{\sum_{j \in C_{i}}|\overline X_{j} - X_{i}|^2}{n_{i}}$$ - partitioning approach - create various partitions and evaluate based on metric (minimize sum of squared errors) - K-means clustering - assigns instances to the nearest centroid - need to know k number of clusters beforehand - algorithm - k points are chosen randomly as initial centroids - assign every data point to the closest centroid - compute new centroids with assigned data - if centroids don't change, stop. If they do repeat with new centroids - Choosing the optimal number of clusters - elbow method - graph inertia against k and find point where elbow of data is (curve levels off) - Silhouette method - Use silhouette coeffic