vault backup: 2026-05-22 12:17:26
This commit is contained in:
1 parent
0972ee377a
commit
d7103664a0
3 files changed
+142
-143
No files matched your search
+56
-57
@@ -1,60 +1,59 @@
|
||||
#rs/notes #rs/class/csb320
|
||||
- - -
|
||||
|
||||
> [!NOTE]- Bullet Notes
|
||||
> - Classification models
|
||||
> - supervised
|
||||
> - classify instances into class (category)
|
||||
> - can predict class directly or probabilities
|
||||
> - Each column in data table is feature
|
||||
> - Each row is sample/tuple/instance
|
||||
> - K-Nearest neighbors
|
||||
> - lazy learner
|
||||
> - memorizes data, doesn't create model during training
|
||||
> - if k=6 it looks at 6 closes neighbors in dataset
|
||||
> - predicts based on what these are classified as
|
||||
> - there can be very different predictions based on k
|
||||
> - gets very expensive as the dataset grows
|
||||
> - selecting value of k
|
||||
> - k is a hyperparameter (user set value )
|
||||
> - typical values of 3-15
|
||||
> - lower value can be sensitive to noise
|
||||
> - higher value risks underfitting
|
||||
> - distance measures
|
||||
> - euclidean distance
|
||||
> - good when data is compact and continuous
|
||||
> - manhattan distance
|
||||
> - sum of absolute differences in coordinates
|
||||
> - good when data is discrete or with large distances
|
||||
> - Minkowski distance
|
||||
> - includes euclidean and manhattan
|
||||
> - parameter allows to interpolate between the two
|
||||
> - can be used for model tuning since distance function can be changed between euclidean and manhattan
|
||||
> - features should be standardized to ensure fair distance measures
|
||||
> - Logistic regression
|
||||
> - log-odds: natural log of the probability ratio
|
||||
> - uses a logistic regression to predict the chances of something being categorized in certain way
|
||||
> - logistic regression $$\hat{p}=\frac{\exp(w_0 + w_1x_i)}{1 + \exp(w_{0} + w_{1}x_{i})}$$
|
||||
> - output of logistic regression is compared to threshold T
|
||||
> - default T = 0.5
|
||||
> - linear regression will underfit for classifying data, logistic regression is better
|
||||
> - multiple input features: $$\hat{p}= \frac{\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}{1+\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}$$
|
||||
> - Gaussian Naive Bayes
|
||||
> - normal distributions for each outcome
|
||||
> - Baye's rule: $$P(A|B) = \frac{P(B|A) * P(A)}{P(B)}$$
|
||||
> - $P(A|B$): Posterior probability
|
||||
> - Assumptions
|
||||
> - all input features are independent
|
||||
> - all input features contribute equally to classification
|
||||
> - requires that each probability is 0
|
||||
> - if there is one option that is 0, can add 1 to each to ensure it works
|
||||
> - Linear Discriminant Analysis
|
||||
> - supervised
|
||||
> - dimensional reduction
|
||||
> - maximize the distance between groups
|
||||
> - within-class variance is minimized
|
||||
> - maximize the distances between the means of the two categories on the new axis
|
||||
> - trying to maximize: $$\frac{(\mu_{1}-\mu_{2})^2}{s_{1}^2-s_{2}^2}$$
|
||||
> - Top of equation is the distance between the averages of the data projected onto the new line
|
||||
> - bottom is minimizing the scatter within each category
|
||||
> - discriminant analysis determines decision boundary between classes
|
||||
- Classification models
|
||||
- supervised
|
||||
- classify instances into class (category)
|
||||
- can predict class directly or probabilities
|
||||
- Each column in data table is feature
|
||||
- Each row is sample/tuple/instance
|
||||
- K-Nearest neighbors
|
||||
- lazy learner
|
||||
- memorizes data, doesn't create model during training
|
||||
- if k=6 it looks at 6 closes neighbors in dataset
|
||||
- predicts based on what these are classified as
|
||||
- there can be very different predictions based on k
|
||||
- gets very expensive as the dataset grows
|
||||
- selecting value of k
|
||||
- k is a hyperparameter (user set value )
|
||||
- typical values of 3-15
|
||||
- lower value can be sensitive to noise
|
||||
- higher value risks underfitting
|
||||
- distance measures
|
||||
- euclidean distance
|
||||
- good when data is compact and continuous
|
||||
- manhattan distance
|
||||
- sum of absolute differences in coordinates
|
||||
- good when data is discrete or with large distances
|
||||
- Minkowski distance
|
||||
- includes euclidean and manhattan
|
||||
- parameter allows to interpolate between the two
|
||||
- can be used for model tuning since distance function can be changed between euclidean and manhattan
|
||||
- features should be standardized to ensure fair distance measures
|
||||
- Logistic regression
|
||||
- log-odds: natural log of the probability ratio
|
||||
- uses a logistic regression to predict the chances of something being categorized in certain way
|
||||
- logistic regression $$\hat{p}=\frac{\exp(w_0 + w_1x_i)}{1 + \exp(w_{0} + w_{1}x_{i})}$$
|
||||
- output of logistic regression is compared to threshold T
|
||||
- default T = 0.5
|
||||
- linear regression will underfit for classifying data, logistic regression is better
|
||||
- multiple input features: $$\hat{p}= \frac{\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}{1+\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}$$
|
||||
- Gaussian Naive Bayes
|
||||
- normal distributions for each outcome
|
||||
- Baye's rule: $$P(A|B) = \frac{P(B|A) * P(A)}{P(B)}$$
|
||||
- $P(A|B$): Posterior probability
|
||||
- Assumptions
|
||||
- all input features are independent
|
||||
- all input features contribute equally to classification
|
||||
- requires that each probability is 0
|
||||
- if there is one option that is 0, can add 1 to each to ensure it works
|
||||
- Linear Discriminant Analysis
|
||||
- supervised
|
||||
- dimensional reduction
|
||||
- maximize the distance between groups
|
||||
- within-class variance is minimized
|
||||
- maximize the distances between the means of the two categories on the new axis
|
||||
- trying to maximize: $$\frac{(\mu_{1}-\mu_{2})^2}{s_{1}^2-s_{2}^2}$$
|
||||
- Top of equation is the distance between the averages of the data projected onto the new line
|
||||
- bottom is minimizing the scatter within each category
|
||||
- discriminant analysis determines decision boundary between classes
|
||||
Reference in new issue
Block a user