3.4 KiB
3.4 KiB
#rs/class/csb320 #rs/notes
- Everything so far has been supervised learning (mostly)
- unsupervised learning (clustering)
- do not know the labels of data.
- clustering
- groups instances based on feature similarity
- results in group assignments, not target output
- need to extrapolate meaning from clusters
- !ClusteringBasic.excalidraw
- applications
- taxonomy of living things
- clustering documents on topic
- identify areas with similar land use
- cluster groups of houses for city planning
- good clustering: high intra-class similarity, low inter-class similarity
- centroid
- mean position of a cluster's instances
\overline X_{i}=\frac{\sum_{j \in C_{i}}X_{i}}{n_{i}}
- mean position of a cluster's instances
- Inertia
- average squared distance of the instances from the centroid
I_{i}=\frac{\sum_{j \in C_{i}}|\overline X_{j} - X_{i}|^2}{n_{i}}
- average squared distance of the instances from the centroid
- partitioning approach
- create various partitions and evaluate based on metric (minimize sum of squared errors)
- K-means clustering
- assigns instances to the nearest centroid
- need to know k number of clusters beforehand
- algorithm
- k points are chosen randomly as initial centroids
- assign every data point to the closest centroid
- compute new centroids with assigned data
- if centroids don't change, stop. If they do repeat with new centroids
- Choosing the optimal number of clusters
- elbow method
- graph inertia against k and find point where elbow of data is (curve levels off)
- Silhouette method
- Use silhouette coefficient to determine separation of clusters
S(j) = \frac{{\overline{d_{out}(j)} - \overline{d_{in}(j)}}}{max(\overline{d_{out}(j)}, \overline{d_{in}(j)})} \overline{d_{out}(j)}is the average distance of the instanceito the centroid of all other clusters\overline{d_{in}(j)}is the average distance of the instanceito all other instances in its cluster- !SilhouetteMethod.excalidraw
- Use silhouette coefficient to determine separation of clusters
- elbow method
- effective for small/medium datasets
- assumes spherical clusters
- sensitive to initial conditions
- equal-sized groups assumption
- sensitive to outliers
- k-medioids can be used
- medioid is the most centrally located point in a cluster
- PAM: partitioning around medioids
- k-medioids can be used
- Hierarchical clustering
- does not require number of clusters, requires stopping point
- Agglomerative
- starts with each instance as cluster and merges into larger clusters
- measures distances between all clusters and merges closest two into new cluster
- divisive
- starts with one cluster and splits into smaller clusters
- linkage methods
- measure distance between two closes instances
- measure distance between two farthest instances
- distance between each cluster's centroid
- dendrograms
- used to visualize clustering hierarchy
- basically just a cladogram
- clades, links, leaves
- does not assume spherical clusters
- sensitive to outliers
- DBSCAN
- density-based clustering method
- clusters are defined as high-density regions separated by low-density regions
- core points, border points and outliers
- outliers are not assigned a cluster and thrown away
- doesn't work very well if clusters have varying densities (low-density cluster could be assigned entirely as an outlier and removed)
- mean-shift clustering
- iteratively shifts centroids to higher density regions
- affinity propagation
- similarity matrix to compute responsibilities and availability's