vault backup: 2026-05-22 12:07:22

This commit is contained in:
ben committed 2026-05-22 12:07:22 -07:00
1 parent f5674d4591
commit 0972ee377a
2 files changed
+66 -63

No files matched your search

@@ -58,66 +58,3 @@
> - Top of equation is the distance between the averages of the data projected onto the new line
> - bottom is minimizing the scatter within each category
> - discriminant analysis determines decision boundary between classes
# Classification Models
Classification models are meant to categorize instances into new classes based on data and training. They can be supervised (using labeled data) and unsupervised (using unlabeled data). They can also either predict classes directly or the probabilities of belonging to certain classes.
### Dataset Terminology
| | Feature | Feature | Feature |
| -------- | ------- | ------- | ------- |
| Sample | | | |
| Tuple | | | |
| Instance | | | |
## Models
### K Nearest Neighbors
Attributes
- Lazy learner (memorizes data, doesn't create model during training)
- Gets very expensive as the dataset grows
- Requires normalization since it is based on distance measures
Prediction process
![[KNN.excalidraw]]
A KNN model selects the k closest points in the dataset and then predicts based on the most common class in this set. Ties are often broken by the class of the closest point.
Several different distance metrics can be used:
![[DistanceMetrics.excalidraw]]
- Euclidean Distance
- Good for when data is compact and continuous $$d(x,y) = (\sum_{i=1}^n {|x_{i} - y_{i}|}^2)^{ \frac{1}{2} }$$
- Manhattan Distance
- Good for when data is discrete or has large distances $$d(x,y) = \sum_{i=1}^n |x_{i} - y_{i}|$$
- Minkowski Distance
- Allows for interpolation between Euclidean and Manhattan distance $$d(x,y) = (\sum_{i=1}^n {|x_{i} - y_{i}|}^p)^{ \frac{1}{p} }$$
- if p = 1 it is the same as Manhattan distance
- if p = 2 it is the same as Euclidean distance
### Logistic Regression
A logistic regression is used to predict the probability of an instance belonging to a certain class. Linear regressions will often underfit data so a logistic regression can be a better choice.
![[LogisticVsLinear.excalidraw]]
The equation for a logistic regression is:
$$\hat{p}=\frac{\exp(w_0 + w_1x_i)}{1 + \exp(w_{0} + w_{1}x_{i})}$$
The output of this regression is compared to the threshold, T. This is usually set to 0.5 but can be changed to bias towards one class.
Logistic regressions can also be used with multiple input features rather than just one with the equation:
$$\hat{p}= \frac{\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}{1+\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}$$
### Gaussian Naive Bayes
A Gaussian NB model assumes that classes follow a Gaussian distribution and uses that to calculate the probability an instance will belong to each class.
$$P(x_{i}|y) = \frac{1}{\sigma \sqrt{ 2\pi }}e^{-\frac{(x-\mu)^2}{2\sigma^2}}$$
$x_{i}$ is the feature value
$\mu$ is the mean of the feature values for a given class $y_{k}$
$\sigma$ is the standard deviation of the feature values for the class
It also uses Bayes rule to calculate posterior probabilities:
$$P(A|B) = \frac{P(B|A) * P(A)}{P(B)}$$
The posterior probability ( $P(A|B)$ ) is the probability that $A$ happens given that $B$ has happened
The algorithm is referred to as "naive" because it makes a few assumptions:
- There is no correlation between the features in the dataset, they are all independent
- Each feature has an equal importance when predicting the output class
### Linear Discriminant Analysis
Linear Discriminant Analysis (LDA) is a supervised learning method that is used to reduce the dimensionality of a dataset. It attempts to maximize the distance between groups and minimize the variation within classes. Since it is supervised it knows the category labels and uses them to maximize the distance between classes.
![[LDA.excalidraw]]