#rs/notes #rs/class/csb320 - - - - loss functions - quantifies differences between predictions and observed values - should minimize loss - absolute loss - prediction of instance - log loss - penalizes wrong predictions more harshly when they are more confident $$L_{\log}(y_{i}, \hat{p}_{i}) = -(y_{i}\ln(\hat{p}_{i}) + (1-y_{i})\ln(1-\hat{p}_{i}))$$ - cross-entropy loss - used for multi-class classification - measures difference between observed and predicted distributions - type I error: false positive - type II error: false negative - evaluation metrics - accuracy $$\frac{\text{num correct predictions}}{\text{num incorrect predictions}}$$ - precision $$\frac{TP}{TP + FP}$$ - F-measures - harmonic mean of precision and recall - beta allows for emphasizing precision or recall in metric: $$F_{\beta} = (1 + \beta^2)\frac{{\text{precision} * \text{recall}}}{\beta^2 * \text{precision} + \text{recall}}$$ - $\beta>1$ emphasizes recall - $\beta<1$ emphasizes precision - kappa - overall proportion of correct predictions - evaluates performance when compared to random classifier - 1: perfect classifier - 0: same accuracy as random chance - <0: worse than random chance - accuracy fails for imbalanced datasets - ex. when most people don't have cancer - AUC-ROC curve - Receiver operating curve - AUC represents the degree of seperability - ROC curves - shows the trade off between true positive rate and false positive rate - ![[rocCurve.excalidraw]] - parametric graphs, false positive rate is on the x-axis and true positive rate is on the y-axis - parameter is threshold for positive classification - when threshold is raised there are less false positives but also less true positives - vice versa - area underneath the curve is a measure of the accuracy - AUC (Area under the curve) - When to use metrics - accuracy is very bad when dataset is not balanced - precision vs. recall - depends on the situation (don't want to have false negatives for cancer) - F1 can maximize precision and recall and gives balance - when balanced classes, maximize accuracy - unbalanced classes can prioritize F1 on only one class if one is more important - Holdout method - data is randomly partitioned into two independent sets - validation set (often half of testing set) is used to decide optimal hyperparameter values - Random subsampling - holdout is repeated k times and accuracy is average - Cross validation - separate into sections, iterate through and use different section as the test set each time - average the error - if the average accuracy goes down with cross validation compared to holdout then model is likely overfitting - stratified cross-validaiton - ensure that the distribution of data is the same in each section as in the general data set