Files
ObsidianVault/Running Start/CSB320 - Machine Learning Concepts/Class 5-21 (Ensemble Models).md
T
2026-05-21 18:43:20 -07:00

1.6 KiB

  • The average of guesses of a large number of people will typically be more accurate than an individual can be
    • (wisdom of crowds)
  • Bootstrapping
    • generates simulated samples by sampling with replacement from existing sample
    • sampling is random
    • Out of bag sample
      • data that was never picked in bootstrapping
      • is often used for testing data (OOB samples were not in the training samples so the model has never seen them)
      • with common techniques often ~30% of data is not picked (and forms OOB set)
    • distribution of classes with bootstrapping will form a Gaussian distribution
    • when enough samples are used bootstrap distributions will approximate population statistics
      • Assumption is made with bootstrapping that the sample approximates the original population
    • bootstrap samples are used as training sets, OOB samples serve as testing sets
bootstrap_samples = [
	resample(df, replace=True, n_samples=n random_state=i) for i in range(5)
]
  • gets a list of dataframes with sampled data
  • ensemble models can often greatly improve performance
    • decision trees are often used as bases
    • can increase computational complexity
  • parallel ensembles
    • base models are trained independently
    • training is faster
  • sequential ensembles
    • models trained iteratively, adjusting for previous errors
  • bagging
    • (bootstrap aggregation)
    • create bootstrap sets, train weak models and then aggregate predictions to get more accurate prediction
    • decision trees in bagging have low bias but high variance
    • bagging reduces model variance, not data variance
      • (how much does model change if data changes slightly)