#rs/class/csb320 #rs/notes - - - - Everything so far has been supervised learning (mostly) - unsupervised learning (clustering) - do not know the labels of data. - clustering - groups instances based on feature similarity - results in group assignments, not target output - need to extrapolate meaning from clusters - ![[ClusteringBasic.excalidraw]] - applications - taxonomy of living things - clustering documents on topic - identify areas with similar land use - cluster groups of houses for city planning - good clustering: high intra-class similarity, low inter-class similarity - centroid - mean position of a cluster's instances $$\overline X_{i}=\frac{\sum_{j \in C_{i}}X_{i}}{n_{i}}$$ - Inertia - average squared distance of the instances from the centroid $$I_{i}=\frac{\sum_{j \in C_{i}}|\overline X_{j} - X_{i}|^2}{n_{i}}$$ - partitioning approach - create various partitions and evaluate based on metric (minimize sum of squared errors) - K-means clustering - assigns instances to the nearest centroid - need to know k number of clusters beforehand - algorithm - k points are chosen randomly as initial centroids - assign every data point to the closest centroid - compute new centroids with assigned data - if centroids don't change, stop. If they do repeat with new centroids - Choosing the optimal number of clusters - elbow method - graph inertia against k and find point where elbow of data is (curve levels off) - Silhouette method - Use silhouette coefficient to determine separation of clusters $$S(j) = \frac{{\overline{d_{out}(j)} - \overline{d_{in}(j)}}}{max(\overline{d_{out}(j)}, \overline{d_{in}(j)})}$$ - $\overline{d_{out}(j)}$ is the average distance of the instance $i$ to the centroid of all other clusters - $\overline{d_{in}(j)}$ is the average distance of the instance $i$ to all other instances in its cluster - ![[SilhouetteMethod.excalidraw]] - effective for small/medium datasets - assumes spherical clusters - sensitive to initial conditions - equal-sized groups assumption - sensitive to outliers - k-medioids can be used - medioid is the most centrally located point in a cluster - PAM: partitioning around medioids - Hierarchical clustering - does not require number of clusters, requires stopping point - Agglomerative - starts with each instance as cluster and merges into larger clusters - measures distances between all clusters and merges closest two into new cluster - divisive - starts with one cluster and splits into smaller clusters - linkage methods - measure distance between two closes instances - measure distance between two farthest instances - distance between each cluster's centroid - dendrograms - used to visualize clustering hierarchy - basically just a cladogram - clades, links, leaves - does not assume spherical clusters - sensitive to outliers -