85 lines
4.7 KiB
Markdown
85 lines
4.7 KiB
Markdown
#rs/notes #rs/class/csb320
|
|
- - -
|
|
|
|
> [!NOTE]- Bullet Notes
|
|
> - loss functions
|
|
> - quantifies differences between predictions and observed values
|
|
> - should minimize loss
|
|
> - absolute loss
|
|
> - prediction of instance
|
|
> - log loss
|
|
> - penalizes wrong predictions more harshly when they are more confident $$L_{\log}(y_{i}, \hat{p}_{i}) = -(y_{i}\ln(\hat{p}_{i}) + (1-y_{i})\ln(1-\hat{p}_{i}))$$
|
|
> - cross-entropy loss
|
|
> - used for multi-class classification
|
|
> - measures difference between observed and predicted distributions
|
|
> - type I error: false positive
|
|
> - type II error: false negative
|
|
> - evaluation metrics
|
|
> - accuracy $$\frac{\text{num correct predictions}}{\text{num incorrect predictions}}$$
|
|
> - precision $$\frac{TP}{TP + FP}$$
|
|
> - F-measures
|
|
> - harmonic mean of precision and recall
|
|
> - beta allows for emphasizing precision or recall in metric: $$F_{\beta} = (1 + \beta^2)\frac{{\text{precision} * \text{recall}}}{\beta^2 * \text{precision} + \text{recall}}$$
|
|
> - $\beta>1$ emphasizes recall
|
|
> - $\beta<1$ emphasizes precision
|
|
> - kappa
|
|
> - overall proportion of correct predictions
|
|
> - evaluates performance when compared to random classifier
|
|
> - 1: perfect classifier
|
|
> - 0: same accuracy as random chance
|
|
> - <0: worse than random chance
|
|
> - accuracy fails for imbalanced datasets
|
|
> - ex. when most people don't have cancer
|
|
> - AUC-ROC curve
|
|
> - Receiver operating curve
|
|
> - AUC represents the degree of seperability
|
|
> - ROC curves
|
|
> - shows the trade off between true positive rate and false positive rate
|
|
> - ![[rocCurve.excalidraw]]
|
|
> - parametric graphs, false positive rate is on the x-axis and true positive rate is on the y-axis
|
|
> - parameter is threshold for positive classification
|
|
> - when threshold is raised there are less false positives but also less true positives
|
|
> - vice versa
|
|
> - area underneath the curve is a measure of the accuracy
|
|
> - AUC (Area under the curve)
|
|
> - When to use metrics
|
|
> - accuracy is very bad when dataset is not balanced
|
|
> - precision vs. recall
|
|
> - depends on the situation (don't want to have false negatives for cancer)
|
|
> - F1 can maximize precision and recall and gives balance
|
|
> - when balanced classes, maximize accuracy
|
|
> - unbalanced classes can prioritize F1 on only one class if one is more important
|
|
> - Holdout method
|
|
> - data is randomly partitioned into two independent sets
|
|
> - validation set (often half of testing set) is used to decide optimal hyperparameter values
|
|
> - Random subsampling
|
|
> - holdout is repeated k times and accuracy is average
|
|
> - Cross validation
|
|
> - separate into sections, iterate through and use different section as the test set each time
|
|
> - average the error
|
|
> - if the average accuracy goes down with cross validation compared to holdout then model is likely overfitting
|
|
> - stratified cross-validaiton
|
|
> - ensure that the distribution of data is the same in each section as in the general data set
|
|
|
|
# Model Validation and Evaluation
|
|
There are many ways to evaluate the performance of models. This is necessary to ensure that they are performing at an acceptable level and to help to tune them with different hyperparameters, preprocessing, data, and other pipeline changes.
|
|
## Loss Functions
|
|
Loss functions are meant to quantify the difference between observed values and the predictions of a model. They can be used to minimize loss and improve model performance.
|
|
### Absolute Loss
|
|
$$Y_{i} - \hat{Y}_{i}$$
|
|
Used to measure the difference from the observed value and predicted value by the model.
|
|
### Mean Absolute Error
|
|
$$MAE=\frac{1}{n}\sum|\hat{Y}_{i} - Y_{i}|$$
|
|
Calculates the average absolute loss of a model based on a training set of data. Is a metric that can be used to improve the accuracy of a model and train it. Should be minimized.
|
|
|
|
Mean squared error is also used to penalize predictions that are further from observed values more heavily.
|
|
$$MSE = \frac{1}{n}\sum(\hat{Y}_{i} - Y_{i})^2$$
|
|
Since MSE produces a convex curve in the error metric it also allows the use of algorithms like gradient descent optimization to tune the weights of a model.
|
|
### Log Loss
|
|
$$L_{\log}(y_{i}, \hat{p}_{i}) = -(y_{i}\ln(\hat{p}_{i}) + (1-y_{i})\ln(1-\hat{p}_{i}))$$
|
|
Log loss is used to optimize the weights of a logistic regression. MSE cannot be used since a logistic regression is not linear and the error with respect to weights is not a convex curve. This makes it challenging to find the weights since algorithms like gradient descent cannot be used.
|
|
|
|
Logistic regressions also only predict between 0 - 1 (probabilities) so any error value will be between 0 - 1 using MSE which is not ideal.
|
|
|
|
The log loss function is used to penalize wrong predictions more harshly when they are more confident.
|