vault backup: 2026-06-08 09:36:02
This commit is contained in:
1 parent
32cb6011b7
commit
70ab6935f2
1 file changed
+6
-1
@@ -7,4 +7,9 @@ Undersampling is one of the easiest ways to deal with imbalanced data. The idea
|
||||
|
||||
The largest problem with this approach comes when the dataset is very imbalanced or there are few instances in it. Since data is removed the total number of instances shrinks. When data is very imbalanced (such as in datasets on a rare disease) this massively reduces the amount of usable data since the most you can have in each class is the size of the smallest class.
|
||||
### Oversampling
|
||||
|
||||
Oversampling is when synthetic data is created for the minority class to equalize the class distributions. One of the most common methods for oversampling is SMOTE.
|
||||
#### SMOTE
|
||||
The synthetic minority oversampling technique (SMOTE) is used to generate synthetic data to increase the size of a minority class. To create the synthetic data you:
|
||||
1. Choose a point in the minority class
|
||||
2.
|
||||
![[SMOTE.excalidraw]]
|
||||
Reference in new issue
Block a user