#machine_learning - - - There are many ways to evaluate the performance of models. This is necessary to ensure that they are performing at an acceptable level and to help to tune them with different hyperparameters, preprocessing, data, and other pipeline changes. ## Loss Functions Loss functions are meant to quantify the difference between observed values and the predictions of a model. They can be used to minimize loss and improve model performance. ### Absolute Loss $$|Y_{i} - \hat{Y}_{i}|$$ Used to measure the difference from the observed value and predicted value by the model. ### Mean Absolute Error $$MAE=\frac{1}{n}\sum|\hat{Y}_{i} - Y_{i}|$$ Calculates the average absolute loss of a model based on a training set of data. Is a metric that can be used to improve the accuracy of a model and train it. Should be minimized. Mean squared error is also used to penalize predictions that are further from observed values more heavily. $$MSE = \frac{1}{n}\sum(\hat{Y}_{i} - Y_{i})^2$$ Since MSE produces a convex curve in the error metric it also allows the use of algorithms like gradient descent optimization to tune the weights of a model. ### Log Loss $$L_{\log}(y_{i}, \hat{p}_{i}) = -(y_{i}\ln(\hat{p}_{i}) + (1-y_{i})\ln(1-\hat{p}_{i}))$$ Log loss is used to optimize the weights of a [[Classification Models#Logistic Regression|logistic regression]]. MSE cannot be used since a logistic regression is not linear and the error with respect to weights is not a convex curve. This makes it challenging to find the weights since algorithms like gradient descent cannot be used. Logistic regressions also only predict between 0 - 1 (probabilities) so any error value will be between 0 - 1 using MSE which is not ideal. The log loss function is used to penalize wrong predictions more harshly when they are more confident. ![[LogLoss.excalidraw]] Only one of the terms inside of the parentheses will be non-zero based on if the observed class is 0 or 1. The log loss function then penalizes the incorrect prediction much more heavily the closer it is to the incorrect value. ### Cross Entropy Loss $$CE = -\sum{Y_{i} * \log(\hat{Y}_{i})}$$ Cross entropy loss is used for multi-class problems and due to the log that is used confident mistakes are heavily penalized. This function also incentivizes the model to decrease the number of uncertain predictions and predict one class. The loss is low when a model confidently predicts the correct class and it is high when it confidently predicts the wrong class. ## Evaluation Metrics ### Accuracy Accuracy is defined as the percentage of predictions that the model gets correct. $$\frac{\text{num correct predictions}}{\text{num incorrect predictions}}$$ This metric is simple but can often hide information about how a model performs on a dataset. One particular limitation is when the dataset is imbalanced since a model can have a high accuracy while predicting the minority class incorrectly most of the time. This is common in datasets relating to disease, especially when outputs such as whether someone has cancer