vault backup: 2026-06-08 09:36:02

This commit is contained in:
ben committed 2026-06-08 09:36:02 -07:00
1 parent 32cb6011b7
commit 70ab6935f2
1 file changed
+6 -1
+6 -1
View File
@@ -7,4 +7,9 @@ Undersampling is one of the easiest ways to deal with imbalanced data. The idea
The largest problem with this approach comes when the dataset is very imbalanced or there are few instances in it. Since data is removed the total number of instances shrinks. When data is very imbalanced (such as in datasets on a rare disease) this massively reduces the amount of usable data since the most you can have in each class is the size of the smallest class.
### Oversampling
Oversampling is when synthetic data is created for the minority class to equalize the class distributions. One of the most common methods for oversampling is SMOTE.
#### SMOTE
The synthetic minority oversampling technique (SMOTE) is used to generate synthetic data to increase the size of a minority class. To create the synthetic data you:
1. Choose a point in the minority class
2.
![[SMOTE.excalidraw]]