Files
ObsidianVault/Wiki/Machine Learning/Imbalanced Data.md
T
2026-06-08 13:16:43 -07:00

1.8 KiB

#machine_learning #rs/class/csb320


Methods

Undersampling

Undersampling is one of the easiest ways to deal with imbalanced data. The idea is that you remove samples from the majority class at random until the two classes have an equal number of instances. This approach is simple and comes with the advantage that no synthetic data needs to be created and all resulting instances are from the original dataset.

The largest problem with this approach comes when the dataset is very imbalanced or there are few instances in it. Since data is removed the total number of instances shrinks. When data is very imbalanced (such as in datasets on a rare disease) this massively reduces the amount of usable data since the most you can have in each class is the size of the smallest class.

Oversampling

Oversampling is when synthetic data is created for the minority class to equalize the class distributions. One of the most common methods for oversampling is SMOTE.

SMOTE

The synthetic minority oversampling technique (SMOTE) is used to generate synthetic data to increase the size of a minority class. To create the synthetic data you:

  1. Choose a point in the minority class
  2. Find k nearest minority neighbors
  3. Select j of these neighbors
  4. Create new synthetic data point along line between first point and selected neighbors
  5. Repeat !SMOTE.excalidraw SMOTE can be helpful to increase the size of a minority dataset but has a few limitations. If the minority class is too small SMOTE can overgeneralize. It also does not work very well with categorical data and can create some data points with values that don't make sense for the feature. There is also no way for it to predict/generate outliers which could cause a model to overfit the data.

Algorithm-Level