vault backup: 2026-06-12 18:50:43

This commit is contained in:
ben committed 2026-06-12 18:50:43 -07:00
1 parent e0af2dc217
commit cbee7025f4
1 file changed
+4
@@ -80,3 +80,7 @@ The area under the ROC curve (AUC) represents the probability that given random
There are many different factors that need to be taken into account when choosing a metric to use to evaluate a model.
One of the biggest ones is whether the dataset is balanced or not. If the dataset is imbalanced accuracy is not a good metric as it can reward a model doing well on a majority class and poorly on a minority class. This is especially problematic in scenarios such as predicting a disease where false negatives are extremely detrimental.
When the dataset is imbalanced other metrics such as precision and recall can be used. These can minimize false negatives or positives and give a better idea of how a model is performing on an imbalanced dataset.
One of the best metrics to use with imbalanced data is F1 since it provides a balance between precision and recall. A F-score can also be tuned with $\beta$ to prioritize precision or recall while still giving a balance between the two.