vault backup: 2026-05-22 12:47:42

This commit is contained in:
ben committed 2026-05-22 12:47:42 -07:00
1 parent d7103664a0
commit d50f8c0665
1 file changed
+8 -1
@@ -16,7 +16,7 @@ $$MSE = \frac{1}{n}\sum(\hat{Y}_{i} - Y_{i})^2$$
Since MSE produces a convex curve in the error metric it also allows the use of algorithms like gradient descent optimization to tune the weights of a model. Since MSE produces a convex curve in the error metric it also allows the use of algorithms like gradient descent optimization to tune the weights of a model.
### Log Loss ### Log Loss
$$L_{\log}(y_{i}, \hat{p}_{i}) = -(y_{i}\ln(\hat{p}_{i}) + (1-y_{i})\ln(1-\hat{p}_{i}))$$ $$L_{\log}(y_{i}, \hat{p}_{i}) = -(y_{i}\ln(\hat{p}_{i}) + (1-y_{i})\ln(1-\hat{p}_{i}))$$
Log loss is used to optimize the weights of a logistic regression. MSE cannot be used since a logistic regression is not linear and the error with respect to weights is not a convex curve. This makes it challenging to find the weights since algorithms like gradient descent cannot be used. Log loss is used to optimize the weights of a [[Classification Models#Logistic Regression|logistic regression]]. MSE cannot be used since a logistic regression is not linear and the error with respect to weights is not a convex curve. This makes it challenging to find the weights since algorithms like gradient descent cannot be used.
Logistic regressions also only predict between 0 - 1 (probabilities) so any error value will be between 0 - 1 using MSE which is not ideal. Logistic regressions also only predict between 0 - 1 (probabilities) so any error value will be between 0 - 1 using MSE which is not ideal.
@@ -26,3 +26,10 @@ The log loss function is used to penalize wrong predictions more harshly when th
Only one of the terms inside of the parentheses will be non-zero based on if the observed class is 0 or 1. The log loss function then penalizes the incorrect prediction much more heavily the closer it is to the incorrect value. Only one of the terms inside of the parentheses will be non-zero based on if the observed class is 0 or 1. The log loss function then penalizes the incorrect prediction much more heavily the closer it is to the incorrect value.
### Cross Entropy Loss ### Cross Entropy Loss
$$CE = -\sum{Y_{i} * \log(\hat{Y}_{i})}$$ $$CE = -\sum{Y_{i} * \log(\hat{Y}_{i})}$$
Cross entropy loss is used for multi-class problems and due to the log that is used confident mistakes are heavily penalized. This function also incentivizes the model to decrease the number of uncertain predictions and predict one class. The loss is low when a model confidently predicts the correct class and it is high when it confidently predicts the wrong class.
## Evaluation Metrics
### Accuracy
Accuracy is defined as the percentage of predictions that the model gets correct.
$$\frac{\text{num correct predictions}}{\text{num incorrect predictions}}$$
This metric is simple but can often hide information about how a model performs on a dataset. One particular limitation is when the dataset is imbalanced since a model can have a high accuracy while predicting the minority class incorrectly most of the time. This is common in datasets relating to disease, especially when outputs such as whether someone has cancer