diff --git a/Wiki/Machine Learning/Model Validation and Evaluation.md b/Wiki/Machine Learning/Model Validation and Evaluation.md index 02abedf..db89f5a 100644 --- a/Wiki/Machine Learning/Model Validation and Evaluation.md +++ b/Wiki/Machine Learning/Model Validation and Evaluation.md @@ -75,4 +75,8 @@ ROC stands for the Receiver-operating characteristic curve. The graph is found b ![[rocCurve.excalidraw]] A perfect classifier is represented by a point at $(0, 1)$ which means that every prediction is correct (no false positives). A straight line from $(0, 0)$ to $(1, 1)$ represents a random classifier (same number of true and false positives). Curves will generally look like the green or blue ones with blue performing better than a random classifier and green performing worse. -The area under the ROC curve (AUC) represents the probability that given random positive and negative examples it will rank the positive example above the negative one. A perfect classifier has a AUC of 1.0 meaning that it will always put a positive instance above a negative one thus separating the two classes. \ No newline at end of file +The area under the ROC curve (AUC) represents the probability that given random positive and negative examples it will rank the positive example above the negative one. A perfect classifier has a AUC of 1.0 meaning that it will always put a positive instance above a negative one thus separating the two classes. +### Metric Usage +There are many different factors that need to be taken into account when choosing a metric to use to evaluate a model. + +One of the biggest ones is whether the dataset is balanced or not. If the dataset is imbalanced accuracy is not a good metric as it can reward a model doing well on a majority class and poorly on a minority class. This is especially problematic in scenarios such as predicting a disease where false negatives are extremely detrimental. \ No newline at end of file