78 lines
5.8 KiB
Markdown
78 lines
5.8 KiB
Markdown
#machine_learning #rs/class/csb320
|
|
- - -
|
|
|
|
There are many ways to evaluate the performance of models. This is necessary to ensure that they are performing at an acceptable level and to help to tune them with different hyperparameters, preprocessing, data, and other pipeline changes.
|
|
## Loss Functions
|
|
Loss functions are meant to quantify the difference between observed values and the predictions of a model. They can be used to minimize loss and improve model performance.
|
|
### Absolute Loss
|
|
$$
|
|
|Y_{i} - \hat{Y}_{i}|
|
|
$$
|
|
Used to measure the difference from the observed value and predicted value by the model.
|
|
### Mean Absolute Error
|
|
$$
|
|
MAE=\frac{1}{n}\sum|\hat{Y}_{i} - Y_{i}|
|
|
$$
|
|
Calculates the average absolute loss of a model based on a training set of data. Is a metric that can be used to improve the accuracy of a model and train it. Should be minimized.
|
|
|
|
Mean squared error is also used to penalize predictions that are further from observed values more heavily.
|
|
$$
|
|
MSE = \frac{1}{n}\sum(\hat{Y}_{i} - Y_{i})^2
|
|
$$
|
|
Since MSE produces a convex curve in the error metric it also allows the use of algorithms like gradient descent optimization to tune the weights of a model.
|
|
### Log Loss
|
|
$$
|
|
L_{\log}(y_{i}, \hat{p}_{i}) = -(y_{i}\ln(\hat{p}_{i}) + (1-y_{i})\ln(1-\hat{p}_{i}))
|
|
$$
|
|
Log loss is used to optimize the weights of a [[Classification Models#Logistic Regression|logistic regression]]. MSE cannot be used since a logistic regression is not linear and the error with respect to weights is not a convex curve. This makes it challenging to find the weights since algorithms like gradient descent cannot be used.
|
|
|
|
Logistic regressions also only predict between 0 - 1 (probabilities) so any error value will be between 0 - 1 using MSE which is not ideal.
|
|
|
|
The log loss function is used to penalize wrong predictions more harshly when they are more confident.
|
|
|
|
![[LogLoss.excalidraw]]
|
|
Only one of the terms inside of the parentheses will be non-zero based on if the observed class is 0 or 1. The log loss function then penalizes the incorrect prediction much more heavily the closer it is to the incorrect value.
|
|
### Cross Entropy Loss
|
|
$$
|
|
CE = -\sum{Y_{i} * \log(\hat{Y}_{i})}
|
|
$$
|
|
Cross entropy loss is used for multi-class problems and due to the log that is used confident mistakes are heavily penalized. This function also incentivizes the model to decrease the number of uncertain predictions and predict one class. The loss is low when a model confidently predicts the correct class and it is high when it confidently predicts the wrong class.
|
|
|
|
## Evaluation Metrics
|
|
### Accuracy
|
|
Accuracy is defined as the percentage of predictions that the model gets correct.
|
|
$$
|
|
\frac{\text{num correct predictions}}{\text{num incorrect predictions}}
|
|
$$
|
|
This metric is simple but can often hide information about how a model performs on a dataset. One particular limitation is when the dataset is imbalanced since a model can have a high accuracy while predicting the minority class incorrectly most of the time. This is common in datasets relating to disease, especially when outputs such as whether someone has cancer need to be accurately predicted. If a model is biased towards the majority class there will likely be many false negatives that accuracy is not able to detect.
|
|
### Precision
|
|
$$
|
|
\frac{TP}{TP + FP}
|
|
$$
|
|
### Recall
|
|
|
|
### F-Measure
|
|
The F-measure is used to balance both the precision and recall metrics to better evaluate how a model performs. This helps to get a better metric on imbalanced datasets since it can help a model train for precision and recall.
|
|
$$
|
|
F_{\beta} = (1 + \beta^2)\frac{{\text{precision} * \text{recall}}}{\beta^2 * \text{precision} + \text{recall}}
|
|
$$
|
|
$\beta$ is used as a way to tune whether precision or recall is emphasized in the F-measure. When $\beta > 1$ recall is emphasized and when $\beta < 1$ precision is emphasized.
|
|
|
|
When $\beta = 1$ the F-measure is also called the F1-measure which is the most commonly used variant. This is where precision and recall are both weighted equally in the metric.
|
|
### Kappa
|
|
$$
|
|
\kappa = \frac{{p_{o} - p_{e}}}{1 - p_{e}}
|
|
$$
|
|
Where $p_{o}$ is the observed agreement and $p_{e}$ is the expected agreement. When using in the context of a binary classification problem $p_{o}$ is the observed accuracy of the model and $p_{e}$ is the expected number of times the model will agree with reality due to random chance (and the percentages that the model will pick each outcome)
|
|
|
|
When using the metric in a binary classification problem with a 2x2 confusion matrix the formula can be written as
|
|
$$
|
|
\kappa = \frac{{2 \times (TP \times TN - FN \times FP)}}{(TP + FP) \times (FP + TN) + (TP + FN) \times (FN + TN)}
|
|
$$
|
|
When using it in this context a value of 1 represents a perfect classifier (model and reality are in perfect agreement) and 0 represents the same accuracy as random chance. The value can also go down to -1 which represents complete disagreement and worse performance than random chance.
|
|
### ROC and AUC
|
|
ROC stands for the Receiver-operating characteristic curve. The graph is found by plotting the true positive rate on the y-axis and the false positive rate on the x-axis as the threshold changes. This is a parametric graph with the threshold as the parameter.
|
|
![[rocCurve.excalidraw]]
|
|
A perfect classifier is represented by a point at $(0, 1)$ which means that every prediction is correct (no false positives). A straight line from $(0, 0)$ to $(1, 1)$ represents a random classifier (same number of true and false positives). Curves will generally look like the green or blue ones with blue performing better than a random classifier and green performing worse.
|
|
|
|
The area under the ROC curve (AUC) represents the probability that given random positive and negative examples it will rank the positive example above the negative one. A perfect classifier has a AUC of 1.0 meaning that it will always put a positive instance above a negative one thus separating the two classes. |