initial commit
This commit is contained in:
commit
fbd12d4a8e
320 files changed
+124298
No files matched your search
+118
@@ -0,0 +1,118 @@
|
||||
#rs/notes #rs/class/csb320
|
||||
- - -
|
||||
|
||||
> [!NOTE]- Bullet Notes
|
||||
> - Classification models
|
||||
> - supervised
|
||||
> - classify instances into class (category)
|
||||
> - can predict class directly or probabilities
|
||||
> - Each column in data table is feature
|
||||
> - Each row is sample/tuple/instance
|
||||
> - K-Nearest neighbors
|
||||
> - lazy learner
|
||||
> - memorizes data, doesn't create model during training
|
||||
> - if k=6 it looks at 6 closes neighbors in dataset
|
||||
> - predicts based on what these are classified as
|
||||
> - there can be very different predictions based on k
|
||||
> - gets very expensive as the dataset grows
|
||||
> - selecting value of k
|
||||
> - k is a hyperparameter (user set value )
|
||||
> - typical values of 3-15
|
||||
> - lower value can be sensitive to noise
|
||||
> - higher value risks underfitting
|
||||
> - distance measures
|
||||
> - euclidean distance
|
||||
> - good when data is compact and continuous
|
||||
> - manhattan distance
|
||||
> - sum of absolute differences in coordinates
|
||||
> - good when data is discrete or with large distances
|
||||
> - Minkowski distance
|
||||
> - includes euclidean and manhattan
|
||||
> - parameter allows to interpolate between the two
|
||||
> - can be used for model tuning since distance function can be changed between euclidean and manhattan
|
||||
> - features should be standardized to ensure fair distance measures
|
||||
> - Logistic regression
|
||||
> - log-odds: natural log of the probability ratio
|
||||
> - uses a logistic regression to predict the chances of something being categorized in certain way
|
||||
> - logistic regression $$\hat{p}=\frac{\exp(w_0 + w_1x_i)}{1 + \exp(w_{0} + w_{1}x_{i})}$$
|
||||
> - output of logistic regression is compared to threshold T
|
||||
> - default T = 0.5
|
||||
> - linear regression will underfit for classifying data, logistic regression is better
|
||||
> - multiple input features: $$\hat{p}= \frac{\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}{1+\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}$$
|
||||
> - Gaussian Naive Bayes
|
||||
> - normal distributions for each outcome
|
||||
> - Baye's rule: $$P(A|B) = \frac{P(B|A) * P(A)}{P(B)}$$
|
||||
> - $P(A|B$): Posterior probability
|
||||
> - Assumptions
|
||||
> - all input features are independent
|
||||
> - all input features contribute equally to classification
|
||||
> - requires that each probability is 0
|
||||
> - if there is one option that is 0, can add 1 to each to ensure it works
|
||||
> - Linear Discriminant Analysis
|
||||
> - supervised
|
||||
> - dimensional reduction
|
||||
> - maximize the distance between groups
|
||||
> - within-class variance is minimized
|
||||
> - maximize the distances between the means of the two categories on the new axis
|
||||
> - trying to maximize: $$\frac{(\mu_{1}-\mu_{2})^2}{s_{1}^2-s_{2}^2}$$
|
||||
> - Top of equation is the distance between the averages of the data projected onto the new line
|
||||
> - bottom is minimizing the scatter within each category
|
||||
> - discriminant analysis determines decision boundary between classes
|
||||
# Classification Models
|
||||
Classification models are meant to categorize instances into new classes based on data and training. They can be supervised (using labeled data) and unsupervised (using unlabeled data). They can also either predict classes directly or the probabilities of belonging to certain classes.
|
||||
### Dataset Terminology
|
||||
|
||||
| | Feature | Feature | Feature |
|
||||
| -------- | ------- | ------- | ------- |
|
||||
| Sample | | | |
|
||||
| Tuple | | | |
|
||||
| Instance | | | |
|
||||
## Models
|
||||
### K Nearest Neighbors
|
||||
|
||||
Attributes
|
||||
- Lazy learner (memorizes data, doesn't create model during training)
|
||||
- Gets very expensive as the dataset grows
|
||||
- Requires normalization since it is based on distance measures
|
||||
|
||||
Prediction process
|
||||
![[KNN.excalidraw]]
|
||||
A KNN model selects the k closest points in the dataset and then predicts based on the most common class in this set. Ties are often broken by the class of the closest point.
|
||||
|
||||
Several different distance metrics can be used:
|
||||
![[DistanceMetrics.excalidraw]]
|
||||
- Euclidean Distance
|
||||
- Good for when data is compact and continuous $$d(x,y) = (\sum_{i=1}^n {|x_{i} - y_{i}|}^2)^{ \frac{1}{2} }$$
|
||||
- Manhattan Distance
|
||||
- Good for when data is discrete or has large distances $$d(x,y) = \sum_{i=1}^n |x_{i} - y_{i}|$$
|
||||
- Minkowski Distance
|
||||
- Allows for interpolation between Euclidean and Manhattan distance $$d(x,y) = (\sum_{i=1}^n {|x_{i} - y_{i}|}^p)^{ \frac{1}{p} }$$
|
||||
- if p = 1 it is the same as Manhattan distance
|
||||
- if p = 2 it is the same as Euclidean distance
|
||||
### Logistic Regression
|
||||
|
||||
A logistic regression is used to predict the probability of an instance belonging to a certain class. Linear regressions will often underfit data so a logistic regression can be a better choice.
|
||||
![[LogisticVsLinear.excalidraw]]
|
||||
|
||||
The equation for a logistic regression is:
|
||||
$$\hat{p}=\frac{\exp(w_0 + w_1x_i)}{1 + \exp(w_{0} + w_{1}x_{i})}$$
|
||||
|
||||
The output of this regression is compared to the threshold, T. This is usually set to 0.5 but can be changed to bias towards one class.
|
||||
|
||||
Logistic regressions can also be used with multiple input features rather than just one with the equation:
|
||||
$$\hat{p}= \frac{\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}{1+\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}$$
|
||||
### Gaussian Naive Bayes
|
||||
|
||||
A Gaussian NB model assumes that classes follow a Gaussian distribution and uses that to calculate the probability an instance will belong to each class.
|
||||
$$P(x_{i}|y) = \frac{1}{\sigma \sqrt{ 2\pi }}e^{-\frac{(x-\mu)^2}{2\sigma^2}}$$
|
||||
$x_{i}$ is the feature value
|
||||
$\mu$ is the mean of the feature values for a given class $y_{k}$
|
||||
$\sigma$ is the standard deviation of the feature values for the class
|
||||
|
||||
It also uses Bayes rule to calculate posterior probabilities:
|
||||
$$P(A|B) = \frac{P(B|A) * P(A)}{P(B)}$$
|
||||
The posterior probability ( $P(A|B)$ ) is the probability that $A$ happens given that $B$ has happened
|
||||
|
||||
The algorithm is referred to as "naive" because it makes a few assumptions:
|
||||
- There is no correlation between the features in the dataset, they are all independent
|
||||
- Each feature has an equal importance when predicting the output class
|
||||
Reference in new issue
Block a user