4.9 KiB
#rs/notes #rs/class/csb320
[!NOTE]- Bullet Notes
- loss functions
- quantifies differences between predictions and observed values
- should minimize loss
- absolute loss
- prediction of instance
- log loss
- penalizes wrong predictions more harshly when they are more confident
L_{\log}(y_{i}, \hat{p}_{i}) = -(y_{i}\ln(\hat{p}_{i}) + (1-y_{i})\ln(1-\hat{p}_{i}))- cross-entropy loss
- used for multi-class classification
- measures difference between observed and predicted distributions
- type I error: false positive
- type II error: false negative
- evaluation metrics
- accuracy
\frac{\text{num correct predictions}}{\text{num incorrect predictions}}- precision
\frac{TP}{TP + FP}- F-measures
- harmonic mean of precision and recall
- beta allows for emphasizing precision or recall in metric:
F_{\beta} = (1 + \beta^2)\frac{{\text{precision} * \text{recall}}}{\beta^2 * \text{precision} + \text{recall}}\beta>1emphasizes recall\beta<1emphasizes precision- kappa
- overall proportion of correct predictions
- evaluates performance when compared to random classifier
- 1: perfect classifier
- 0: same accuracy as random chance
- <0: worse than random chance
- accuracy fails for imbalanced datasets
- ex. when most people don't have cancer
- AUC-ROC curve
- Receiver operating curve
- AUC represents the degree of seperability
- ROC curves
- shows the trade off between true positive rate and false positive rate
- !rocCurve.excalidraw
- parametric graphs, false positive rate is on the x-axis and true positive rate is on the y-axis
- parameter is threshold for positive classification
- when threshold is raised there are less false positives but also less true positives
- vice versa
- area underneath the curve is a measure of the accuracy
- AUC (Area under the curve)
- When to use metrics
- accuracy is very bad when dataset is not balanced
- precision vs. recall
- depends on the situation (don't want to have false negatives for cancer)
- F1 can maximize precision and recall and gives balance
- when balanced classes, maximize accuracy
- unbalanced classes can prioritize F1 on only one class if one is more important
- Holdout method
- data is randomly partitioned into two independent sets
- validation set (often half of testing set) is used to decide optimal hyperparameter values
- Random subsampling
- holdout is repeated k times and accuracy is average
- Cross validation
- separate into sections, iterate through and use different section as the test set each time
- average the error
- if the average accuracy goes down with cross validation compared to holdout then model is likely overfitting
- stratified cross-validaiton
- ensure that the distribution of data is the same in each section as in the general data set
Model Validation and Evaluation
There are many ways to evaluate the performance of models. This is necessary to ensure that they are performing at an acceptable level and to help to tune them with different hyperparameters, preprocessing, data, and other pipeline changes.
Loss Functions
Loss functions are meant to quantify the difference between observed values and the predictions of a model. They can be used to minimize loss and improve model performance.
Absolute Loss
Y_{i} - \hat{Y}_{i}
Used to measure the difference from the observed value and predicted value by the model.
Mean Absolute Error
MAE=\frac{1}{n}\sum|\hat{Y}_{i} - Y_{i}|
Calculates the average absolute loss of a model based on a training set of data. Is a metric that can be used to improve the accuracy of a model and train it. Should be minimized.
Mean squared error is also used to penalize predictions that are further from observed values more heavily.
MSE = \frac{1}{n}\sum(\hat{Y}_{i} - Y_{i})^2
Since MSE produces a convex curve in the error metric it also allows the use of algorithms like gradient descent optimization to tune the weights of a model.
Log Loss
L_{\log}(y_{i}, \hat{p}_{i}) = -(y_{i}\ln(\hat{p}_{i}) + (1-y_{i})\ln(1-\hat{p}_{i}))
Log loss is used to optimize the weights of a logistic regression. MSE cannot be used since a logistic regression is not linear and the error with respect to weights is not a convex curve. This makes it challenging to find the weights since algorithms like gradient descent cannot be used.
Logistic regressions also only predict between 0 - 1 (probabilities) so any error value will be between 0 - 1 using MSE which is not ideal.
The log loss function is used to penalize wrong predictions more harshly when they are more confident.
!LogLoss.excalidraw Only one of the terms inside of the parentheses will be non-zero based on if the observed class is 0 or 1. The log loss function then penalizes the incorrect prediction much more heavily the closer it is to the incorrect value