74 lines
3.4 KiB
Markdown
74 lines
3.4 KiB
Markdown
#rs/class/csb320 #rs/notes
|
|
- - -
|
|
- Everything so far has been supervised learning (mostly)
|
|
- unsupervised learning (clustering)
|
|
- do not know the labels of data.
|
|
- clustering
|
|
- groups instances based on feature similarity
|
|
- results in group assignments, not target output
|
|
- need to extrapolate meaning from clusters
|
|
- ![[ClusteringBasic.excalidraw]]
|
|
- applications
|
|
- taxonomy of living things
|
|
- clustering documents on topic
|
|
- identify areas with similar land use
|
|
- cluster groups of houses for city planning
|
|
- good clustering: high intra-class similarity, low inter-class similarity
|
|
- centroid
|
|
- mean position of a cluster's instances $$\overline X_{i}=\frac{\sum_{j \in C_{i}}X_{i}}{n_{i}}$$
|
|
- Inertia
|
|
- average squared distance of the instances from the centroid $$I_{i}=\frac{\sum_{j \in C_{i}}|\overline X_{j} - X_{i}|^2}{n_{i}}$$
|
|
- partitioning approach
|
|
- create various partitions and evaluate based on metric (minimize sum of squared errors)
|
|
- K-means clustering
|
|
- assigns instances to the nearest centroid
|
|
- need to know k number of clusters beforehand
|
|
- algorithm
|
|
- k points are chosen randomly as initial centroids
|
|
- assign every data point to the closest centroid
|
|
- compute new centroids with assigned data
|
|
- if centroids don't change, stop. If they do repeat with new centroids
|
|
- Choosing the optimal number of clusters
|
|
- elbow method
|
|
- graph inertia against k and find point where elbow of data is (curve levels off)
|
|
- Silhouette method
|
|
- Use silhouette coefficient to determine separation of clusters $$S(j) = \frac{{\overline{d_{out}(j)} - \overline{d_{in}(j)}}}{max(\overline{d_{out}(j)}, \overline{d_{in}(j)})}$$
|
|
- $\overline{d_{out}(j)}$ is the average distance of the instance $i$ to the centroid of all other clusters
|
|
- $\overline{d_{in}(j)}$ is the average distance of the instance $i$ to all other instances in its cluster
|
|
- ![[SilhouetteMethod.excalidraw]]
|
|
- effective for small/medium datasets
|
|
- assumes spherical clusters
|
|
- sensitive to initial conditions
|
|
- equal-sized groups assumption
|
|
- sensitive to outliers
|
|
- k-medioids can be used
|
|
- medioid is the most centrally located point in a cluster
|
|
- PAM: partitioning around medioids
|
|
- Hierarchical clustering
|
|
- does not require number of clusters, requires stopping point
|
|
- Agglomerative
|
|
- starts with each instance as cluster and merges into larger clusters
|
|
- measures distances between all clusters and merges closest two into new cluster
|
|
- divisive
|
|
- starts with one cluster and splits into smaller clusters
|
|
- linkage methods
|
|
- measure distance between two closes instances
|
|
- measure distance between two farthest instances
|
|
- distance between each cluster's centroid
|
|
- dendrograms
|
|
- used to visualize clustering hierarchy
|
|
- basically just a cladogram
|
|
- clades, links, leaves
|
|
- does not assume spherical clusters
|
|
- sensitive to outliers
|
|
- DBSCAN
|
|
- density-based clustering method
|
|
- clusters are defined as high-density regions separated by low-density regions
|
|
- core points, border points and outliers
|
|
- outliers are not assigned a cluster and thrown away
|
|
- doesn't work very well if clusters have varying densities (low-density cluster could be assigned entirely as an outlier and removed)
|
|
- mean-shift clustering
|
|
- iteratively shifts centroids to higher density regions
|
|
- affinity propagation
|
|
- similarity matrix to compute responsibilities and availability's
|
|
- |