vault backup: 2026-06-08 00:51:13

This commit is contained in:
ben committed 2026-06-08 00:51:13 -07:00
1 parent 2fea9cb507
commit dfb1c6be0f
2 files changed
+4 -2

No files matched your search

-1
View File
@@ -1,5 +1,4 @@
[ [
"harper",
"obsidian-advanced-uri", "obsidian-advanced-uri",
"obsidian-reading-time", "obsidian-reading-time",
"obsidian-vault-statistics-plugin", "obsidian-vault-statistics-plugin",
+4 -1
View File
@@ -3,5 +3,8 @@
## Methods ## Methods
### Undersampling ### Undersampling
Undersampling is one of the easiest ways to deal with imbalanced data. The idea is that you remove samples from the majority class at random until the two classes have an equal number of instances. This approach is simple and comes with the advantage that no synthetic data ne Undersampling is one of the easiest ways to deal with imbalanced data. The idea is that you remove samples from the majority class at random until the two classes have an equal number of instances. This approach is simple and comes with the advantage that no synthetic data needs to be created and all resulting instances are from the original dataset.
The largest problem with this approach comes when the dataset is very imbalanced or there are few instances in it. Since data is removed the total number of instances shrinks. When data is very imbalanced (such as in datasets on a rare disease) this massively reduces the amount of usable data since the most you can have in each class is the size of the smallest class.
### Oversampling ### Oversampling