diff --git a/Wiki/Machine Learning/Model Validation and Evaluation.md b/Wiki/Machine Learning/Model Validation and Evaluation.md index db89f5a..1d20a5c 100644 --- a/Wiki/Machine Learning/Model Validation and Evaluation.md +++ b/Wiki/Machine Learning/Model Validation and Evaluation.md @@ -79,4 +79,8 @@ The area under the ROC curve (AUC) represents the probability that given random ### Metric Usage There are many different factors that need to be taken into account when choosing a metric to use to evaluate a model. -One of the biggest ones is whether the dataset is balanced or not. If the dataset is imbalanced accuracy is not a good metric as it can reward a model doing well on a majority class and poorly on a minority class. This is especially problematic in scenarios such as predicting a disease where false negatives are extremely detrimental. \ No newline at end of file +One of the biggest ones is whether the dataset is balanced or not. If the dataset is imbalanced accuracy is not a good metric as it can reward a model doing well on a majority class and poorly on a minority class. This is especially problematic in scenarios such as predicting a disease where false negatives are extremely detrimental. + +When the dataset is imbalanced other metrics such as precision and recall can be used. These can minimize false negatives or positives and give a better idea of how a model is performing on an imbalanced dataset. + +One of the best metrics to use with imbalanced data is F1 since it provides a balance between precision and recall. A F-score can also be tuned with $\beta$ to prioritize precision or recall while still giving a balance between the two. \ No newline at end of file