From dfb1c6be0fa2494f177c5ab052b7dd3e587273cd Mon Sep 17 00:00:00 2001 From: ben-jaynes <1btjaynes@gmail.com> Date: Mon, 8 Jun 2026 00:51:13 -0700 Subject: [PATCH] vault backup: 2026-06-08 00:51:13 --- .obsidian/community-plugins.json | 1 - Wiki/Machine Learning/Imbalanced Data.md | 5 ++++- 2 files changed, 4 insertions(+), 2 deletions(-) diff --git a/.obsidian/community-plugins.json b/.obsidian/community-plugins.json index 71d2ef8..7d05005 100644 --- a/.obsidian/community-plugins.json +++ b/.obsidian/community-plugins.json @@ -1,5 +1,4 @@ [ - "harper", "obsidian-advanced-uri", "obsidian-reading-time", "obsidian-vault-statistics-plugin", diff --git a/Wiki/Machine Learning/Imbalanced Data.md b/Wiki/Machine Learning/Imbalanced Data.md index 1dd6c58..324e495 100644 --- a/Wiki/Machine Learning/Imbalanced Data.md +++ b/Wiki/Machine Learning/Imbalanced Data.md @@ -3,5 +3,8 @@ ## Methods ### Undersampling -Undersampling is one of the easiest ways to deal with imbalanced data. The idea is that you remove samples from the majority class at random until the two classes have an equal number of instances. This approach is simple and comes with the advantage that no synthetic data ne +Undersampling is one of the easiest ways to deal with imbalanced data. The idea is that you remove samples from the majority class at random until the two classes have an equal number of instances. This approach is simple and comes with the advantage that no synthetic data needs to be created and all resulting instances are from the original dataset. + +The largest problem with this approach comes when the dataset is very imbalanced or there are few instances in it. Since data is removed the total number of instances shrinks. When data is very imbalanced (such as in datasets on a rare disease) this massively reduces the amount of usable data since the most you can have in each class is the size of the smallest class. ### Oversampling + \ No newline at end of file