1.9 KiB
1.9 KiB
#rs/class/csb320 #rs/notes
- Everything so far has been supervised learning (mostly)
- unsupervised learning (clustering)
- do not know the labels of data.
- clustering
- groups instances based on feature similarity
- results in group assignments, not target output
- need to extrapolate meaning from clusters
- !ClusteringBasic.excalidraw
- applications
- taxonomy of living things
- clustering documents on topic
- identify areas with similar land use
- cluster groups of houses for city planning
- good clustering: high intra-class similarity, low inter-class similarity
- centroid
- mean position of a cluster's instances
\overline X_{i}=\frac{\sum_{j \in C_{i}}X_{i}}{n_{i}}
- mean position of a cluster's instances
- Inertia
- average squared distance of the instances from the centroid
I_{i}=\frac{\sum_{j \in C_{i}}|\overline X_{j} - X_{i}|^2}{n_{i}}
- average squared distance of the instances from the centroid
- partitioning approach
- create various partitions and evaluate based on metric (minimize sum of squared errors)
- K-means clustering
- assigns instances to the nearest centroid
- need to know k number of clusters beforehand
- algorithm
- k points are chosen randomly as initial centroids
- assign every data point to the closest centroid
- compute new centroids with assigned data
- if centroids don't change, stop. If they do repeat with new centroids
- Choosing the optimal number of clusters
- elbow method
- graph inertia against k and find point where elbow of data is (curve levels off)
- Silhouette method
- Use silhouette coefficient to determine separation of clusters
S(j) = \frac{{\overline{d_{out}(j)} - \overline{d_{in}(j)}}}{max(\overline{d_{out}(j)}, \overline{d_{in}(j)})} \overline{d_{out}(j)}is the average distance of the instanceito the centroid of all other clusters\overline{d_{in}(j)}is the average distance of the instanceito all other instances in its cluster- !SilhouetteMethod.excalidraw
- Use silhouette coefficient to determine separation of clusters
- elbow method