vault backup: 2026-05-22 12:17:26
This commit is contained in:
1 parent
0972ee377a
commit
d7103664a0
3 files changed
+142
-143
No files matched your search
+56
-57
@@ -1,60 +1,59 @@
|
||||
#rs/notes #rs/class/csb320
|
||||
- - -
|
||||
|
||||
> [!NOTE]- Bullet Notes
|
||||
> - Classification models
|
||||
> - supervised
|
||||
> - classify instances into class (category)
|
||||
> - can predict class directly or probabilities
|
||||
> - Each column in data table is feature
|
||||
> - Each row is sample/tuple/instance
|
||||
> - K-Nearest neighbors
|
||||
> - lazy learner
|
||||
> - memorizes data, doesn't create model during training
|
||||
> - if k=6 it looks at 6 closes neighbors in dataset
|
||||
> - predicts based on what these are classified as
|
||||
> - there can be very different predictions based on k
|
||||
> - gets very expensive as the dataset grows
|
||||
> - selecting value of k
|
||||
> - k is a hyperparameter (user set value )
|
||||
> - typical values of 3-15
|
||||
> - lower value can be sensitive to noise
|
||||
> - higher value risks underfitting
|
||||
> - distance measures
|
||||
> - euclidean distance
|
||||
> - good when data is compact and continuous
|
||||
> - manhattan distance
|
||||
> - sum of absolute differences in coordinates
|
||||
> - good when data is discrete or with large distances
|
||||
> - Minkowski distance
|
||||
> - includes euclidean and manhattan
|
||||
> - parameter allows to interpolate between the two
|
||||
> - can be used for model tuning since distance function can be changed between euclidean and manhattan
|
||||
> - features should be standardized to ensure fair distance measures
|
||||
> - Logistic regression
|
||||
> - log-odds: natural log of the probability ratio
|
||||
> - uses a logistic regression to predict the chances of something being categorized in certain way
|
||||
> - logistic regression $$\hat{p}=\frac{\exp(w_0 + w_1x_i)}{1 + \exp(w_{0} + w_{1}x_{i})}$$
|
||||
> - output of logistic regression is compared to threshold T
|
||||
> - default T = 0.5
|
||||
> - linear regression will underfit for classifying data, logistic regression is better
|
||||
> - multiple input features: $$\hat{p}= \frac{\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}{1+\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}$$
|
||||
> - Gaussian Naive Bayes
|
||||
> - normal distributions for each outcome
|
||||
> - Baye's rule: $$P(A|B) = \frac{P(B|A) * P(A)}{P(B)}$$
|
||||
> - $P(A|B$): Posterior probability
|
||||
> - Assumptions
|
||||
> - all input features are independent
|
||||
> - all input features contribute equally to classification
|
||||
> - requires that each probability is 0
|
||||
> - if there is one option that is 0, can add 1 to each to ensure it works
|
||||
> - Linear Discriminant Analysis
|
||||
> - supervised
|
||||
> - dimensional reduction
|
||||
> - maximize the distance between groups
|
||||
> - within-class variance is minimized
|
||||
> - maximize the distances between the means of the two categories on the new axis
|
||||
> - trying to maximize: $$\frac{(\mu_{1}-\mu_{2})^2}{s_{1}^2-s_{2}^2}$$
|
||||
> - Top of equation is the distance between the averages of the data projected onto the new line
|
||||
> - bottom is minimizing the scatter within each category
|
||||
> - discriminant analysis determines decision boundary between classes
|
||||
- Classification models
|
||||
- supervised
|
||||
- classify instances into class (category)
|
||||
- can predict class directly or probabilities
|
||||
- Each column in data table is feature
|
||||
- Each row is sample/tuple/instance
|
||||
- K-Nearest neighbors
|
||||
- lazy learner
|
||||
- memorizes data, doesn't create model during training
|
||||
- if k=6 it looks at 6 closes neighbors in dataset
|
||||
- predicts based on what these are classified as
|
||||
- there can be very different predictions based on k
|
||||
- gets very expensive as the dataset grows
|
||||
- selecting value of k
|
||||
- k is a hyperparameter (user set value )
|
||||
- typical values of 3-15
|
||||
- lower value can be sensitive to noise
|
||||
- higher value risks underfitting
|
||||
- distance measures
|
||||
- euclidean distance
|
||||
- good when data is compact and continuous
|
||||
- manhattan distance
|
||||
- sum of absolute differences in coordinates
|
||||
- good when data is discrete or with large distances
|
||||
- Minkowski distance
|
||||
- includes euclidean and manhattan
|
||||
- parameter allows to interpolate between the two
|
||||
- can be used for model tuning since distance function can be changed between euclidean and manhattan
|
||||
- features should be standardized to ensure fair distance measures
|
||||
- Logistic regression
|
||||
- log-odds: natural log of the probability ratio
|
||||
- uses a logistic regression to predict the chances of something being categorized in certain way
|
||||
- logistic regression $$\hat{p}=\frac{\exp(w_0 + w_1x_i)}{1 + \exp(w_{0} + w_{1}x_{i})}$$
|
||||
- output of logistic regression is compared to threshold T
|
||||
- default T = 0.5
|
||||
- linear regression will underfit for classifying data, logistic regression is better
|
||||
- multiple input features: $$\hat{p}= \frac{\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}{1+\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}$$
|
||||
- Gaussian Naive Bayes
|
||||
- normal distributions for each outcome
|
||||
- Baye's rule: $$P(A|B) = \frac{P(B|A) * P(A)}{P(B)}$$
|
||||
- $P(A|B$): Posterior probability
|
||||
- Assumptions
|
||||
- all input features are independent
|
||||
- all input features contribute equally to classification
|
||||
- requires that each probability is 0
|
||||
- if there is one option that is 0, can add 1 to each to ensure it works
|
||||
- Linear Discriminant Analysis
|
||||
- supervised
|
||||
- dimensional reduction
|
||||
- maximize the distance between groups
|
||||
- within-class variance is minimized
|
||||
- maximize the distances between the means of the two categories on the new axis
|
||||
- trying to maximize: $$\frac{(\mu_{1}-\mu_{2})^2}{s_{1}^2-s_{2}^2}$$
|
||||
- Top of equation is the distance between the averages of the data projected onto the new line
|
||||
- bottom is minimizing the scatter within each category
|
||||
- discriminant analysis determines decision boundary between classes
|
||||
+58
-86
@@ -1,89 +1,61 @@
|
||||
#rs/notes #rs/class/csb320
|
||||
- - -
|
||||
|
||||
> [!NOTE]- Bullet Notes
|
||||
> - loss functions
|
||||
> - quantifies differences between predictions and observed values
|
||||
> - should minimize loss
|
||||
> - absolute loss
|
||||
> - prediction of instance
|
||||
> - log loss
|
||||
> - penalizes wrong predictions more harshly when they are more confident $$L_{\log}(y_{i}, \hat{p}_{i}) = -(y_{i}\ln(\hat{p}_{i}) + (1-y_{i})\ln(1-\hat{p}_{i}))$$
|
||||
> - cross-entropy loss
|
||||
> - used for multi-class classification
|
||||
> - measures difference between observed and predicted distributions
|
||||
> - type I error: false positive
|
||||
> - type II error: false negative
|
||||
> - evaluation metrics
|
||||
> - accuracy $$\frac{\text{num correct predictions}}{\text{num incorrect predictions}}$$
|
||||
> - precision $$\frac{TP}{TP + FP}$$
|
||||
> - F-measures
|
||||
> - harmonic mean of precision and recall
|
||||
> - beta allows for emphasizing precision or recall in metric: $$F_{\beta} = (1 + \beta^2)\frac{{\text{precision} * \text{recall}}}{\beta^2 * \text{precision} + \text{recall}}$$
|
||||
> - $\beta>1$ emphasizes recall
|
||||
> - $\beta<1$ emphasizes precision
|
||||
> - kappa
|
||||
> - overall proportion of correct predictions
|
||||
> - evaluates performance when compared to random classifier
|
||||
> - 1: perfect classifier
|
||||
> - 0: same accuracy as random chance
|
||||
> - <0: worse than random chance
|
||||
> - accuracy fails for imbalanced datasets
|
||||
> - ex. when most people don't have cancer
|
||||
> - AUC-ROC curve
|
||||
> - Receiver operating curve
|
||||
> - AUC represents the degree of seperability
|
||||
> - ROC curves
|
||||
> - shows the trade off between true positive rate and false positive rate
|
||||
> - ![[rocCurve.excalidraw]]
|
||||
> - parametric graphs, false positive rate is on the x-axis and true positive rate is on the y-axis
|
||||
> - parameter is threshold for positive classification
|
||||
> - when threshold is raised there are less false positives but also less true positives
|
||||
> - vice versa
|
||||
> - area underneath the curve is a measure of the accuracy
|
||||
> - AUC (Area under the curve)
|
||||
> - When to use metrics
|
||||
> - accuracy is very bad when dataset is not balanced
|
||||
> - precision vs. recall
|
||||
> - depends on the situation (don't want to have false negatives for cancer)
|
||||
> - F1 can maximize precision and recall and gives balance
|
||||
> - when balanced classes, maximize accuracy
|
||||
> - unbalanced classes can prioritize F1 on only one class if one is more important
|
||||
> - Holdout method
|
||||
> - data is randomly partitioned into two independent sets
|
||||
> - validation set (often half of testing set) is used to decide optimal hyperparameter values
|
||||
> - Random subsampling
|
||||
> - holdout is repeated k times and accuracy is average
|
||||
> - Cross validation
|
||||
> - separate into sections, iterate through and use different section as the test set each time
|
||||
> - average the error
|
||||
> - if the average accuracy goes down with cross validation compared to holdout then model is likely overfitting
|
||||
> - stratified cross-validaiton
|
||||
> - ensure that the distribution of data is the same in each section as in the general data set
|
||||
|
||||
# Model Validation and Evaluation
|
||||
There are many ways to evaluate the performance of models. This is necessary to ensure that they are performing at an acceptable level and to help to tune them with different hyperparameters, preprocessing, data, and other pipeline changes.
|
||||
## Loss Functions
|
||||
Loss functions are meant to quantify the difference between observed values and the predictions of a model. They can be used to minimize loss and improve model performance.
|
||||
### Absolute Loss
|
||||
$$|Y_{i} - \hat{Y}_{i}|$$
|
||||
Used to measure the difference from the observed value and predicted value by the model.
|
||||
### Mean Absolute Error
|
||||
$$MAE=\frac{1}{n}\sum|\hat{Y}_{i} - Y_{i}|$$
|
||||
Calculates the average absolute loss of a model based on a training set of data. Is a metric that can be used to improve the accuracy of a model and train it. Should be minimized.
|
||||
|
||||
Mean squared error is also used to penalize predictions that are further from observed values more heavily.
|
||||
$$MSE = \frac{1}{n}\sum(\hat{Y}_{i} - Y_{i})^2$$
|
||||
Since MSE produces a convex curve in the error metric it also allows the use of algorithms like gradient descent optimization to tune the weights of a model.
|
||||
### Log Loss
|
||||
$$L_{\log}(y_{i}, \hat{p}_{i}) = -(y_{i}\ln(\hat{p}_{i}) + (1-y_{i})\ln(1-\hat{p}_{i}))$$
|
||||
Log loss is used to optimize the weights of a logistic regression. MSE cannot be used since a logistic regression is not linear and the error with respect to weights is not a convex curve. This makes it challenging to find the weights since algorithms like gradient descent cannot be used.
|
||||
|
||||
Logistic regressions also only predict between 0 - 1 (probabilities) so any error value will be between 0 - 1 using MSE which is not ideal.
|
||||
|
||||
The log loss function is used to penalize wrong predictions more harshly when they are more confident.
|
||||
|
||||
![[LogLoss.excalidraw]]
|
||||
Only one of the terms inside of the parentheses will be non-zero based on if the observed class is 0 or 1. The log loss function then penalizes the incorrect prediction much more heavily the closer it is to the incorrect value.
|
||||
### Cross Entropy Loss
|
||||
$$CE = -\sum{Y_{i} * \log(\hat{Y}_{i})}$$
|
||||
- loss functions
|
||||
- quantifies differences between predictions and observed values
|
||||
- should minimize loss
|
||||
- absolute loss
|
||||
- prediction of instance
|
||||
- log loss
|
||||
- penalizes wrong predictions more harshly when they are more confident $$L_{\log}(y_{i}, \hat{p}_{i}) = -(y_{i}\ln(\hat{p}_{i}) + (1-y_{i})\ln(1-\hat{p}_{i}))$$
|
||||
- cross-entropy loss
|
||||
- used for multi-class classification
|
||||
- measures difference between observed and predicted distributions
|
||||
- type I error: false positive
|
||||
- type II error: false negative
|
||||
- evaluation metrics
|
||||
- accuracy $$\frac{\text{num correct predictions}}{\text{num incorrect predictions}}$$
|
||||
- precision $$\frac{TP}{TP + FP}$$
|
||||
- F-measures
|
||||
- harmonic mean of precision and recall
|
||||
- beta allows for emphasizing precision or recall in metric: $$F_{\beta} = (1 + \beta^2)\frac{{\text{precision} * \text{recall}}}{\beta^2 * \text{precision} + \text{recall}}$$
|
||||
- $\beta>1$ emphasizes recall
|
||||
- $\beta<1$ emphasizes precision
|
||||
- kappa
|
||||
- overall proportion of correct predictions
|
||||
- evaluates performance when compared to random classifier
|
||||
- 1: perfect classifier
|
||||
- 0: same accuracy as random chance
|
||||
- <0: worse than random chance
|
||||
- accuracy fails for imbalanced datasets
|
||||
- ex. when most people don't have cancer
|
||||
- AUC-ROC curve
|
||||
- Receiver operating curve
|
||||
- AUC represents the degree of seperability
|
||||
- ROC curves
|
||||
- shows the trade off between true positive rate and false positive rate
|
||||
- ![[rocCurve.excalidraw]]
|
||||
- parametric graphs, false positive rate is on the x-axis and true positive rate is on the y-axis
|
||||
- parameter is threshold for positive classification
|
||||
- when threshold is raised there are less false positives but also less true positives
|
||||
- vice versa
|
||||
- area underneath the curve is a measure of the accuracy
|
||||
- AUC (Area under the curve)
|
||||
- When to use metrics
|
||||
- accuracy is very bad when dataset is not balanced
|
||||
- precision vs. recall
|
||||
- depends on the situation (don't want to have false negatives for cancer)
|
||||
- F1 can maximize precision and recall and gives balance
|
||||
- when balanced classes, maximize accuracy
|
||||
- unbalanced classes can prioritize F1 on only one class if one is more important
|
||||
- Holdout method
|
||||
- data is randomly partitioned into two independent sets
|
||||
- validation set (often half of testing set) is used to decide optimal hyperparameter values
|
||||
- Random subsampling
|
||||
- holdout is repeated k times and accuracy is average
|
||||
- Cross validation
|
||||
- separate into sections, iterate through and use different section as the test set each time
|
||||
- average the error
|
||||
- if the average accuracy goes down with cross validation compared to holdout then model is likely overfitting
|
||||
- stratified cross-validaiton
|
||||
- ensure that the distribution of data is the same in each section as in the general data set
|
||||
@@ -0,0 +1,28 @@
|
||||
#machine_learning
|
||||
- - -
|
||||
|
||||
There are many ways to evaluate the performance of models. This is necessary to ensure that they are performing at an acceptable level and to help to tune them with different hyperparameters, preprocessing, data, and other pipeline changes.
|
||||
## Loss Functions
|
||||
Loss functions are meant to quantify the difference between observed values and the predictions of a model. They can be used to minimize loss and improve model performance.
|
||||
### Absolute Loss
|
||||
$$|Y_{i} - \hat{Y}_{i}|$$
|
||||
Used to measure the difference from the observed value and predicted value by the model.
|
||||
### Mean Absolute Error
|
||||
$$MAE=\frac{1}{n}\sum|\hat{Y}_{i} - Y_{i}|$$
|
||||
Calculates the average absolute loss of a model based on a training set of data. Is a metric that can be used to improve the accuracy of a model and train it. Should be minimized.
|
||||
|
||||
Mean squared error is also used to penalize predictions that are further from observed values more heavily.
|
||||
$$MSE = \frac{1}{n}\sum(\hat{Y}_{i} - Y_{i})^2$$
|
||||
Since MSE produces a convex curve in the error metric it also allows the use of algorithms like gradient descent optimization to tune the weights of a model.
|
||||
### Log Loss
|
||||
$$L_{\log}(y_{i}, \hat{p}_{i}) = -(y_{i}\ln(\hat{p}_{i}) + (1-y_{i})\ln(1-\hat{p}_{i}))$$
|
||||
Log loss is used to optimize the weights of a logistic regression. MSE cannot be used since a logistic regression is not linear and the error with respect to weights is not a convex curve. This makes it challenging to find the weights since algorithms like gradient descent cannot be used.
|
||||
|
||||
Logistic regressions also only predict between 0 - 1 (probabilities) so any error value will be between 0 - 1 using MSE which is not ideal.
|
||||
|
||||
The log loss function is used to penalize wrong predictions more harshly when they are more confident.
|
||||
|
||||
![[LogLoss.excalidraw]]
|
||||
Only one of the terms inside of the parentheses will be non-zero based on if the observed class is 0 or 1. The log loss function then penalizes the incorrect prediction much more heavily the closer it is to the incorrect value.
|
||||
### Cross Entropy Loss
|
||||
$$CE = -\sum{Y_{i} * \log(\hat{Y}_{i})}$$
|
||||
Reference in new issue
Block a user