1.6 KiB
1.6 KiB
- The average of guesses of a large number of people will typically be more accurate than an individual can be
- (wisdom of crowds)
- Bootstrapping
- generates simulated samples by sampling with replacement from existing sample
- sampling is random
- Out of bag sample
- data that was never picked in bootstrapping
- is often used for testing data (OOB samples were not in the training samples so the model has never seen them)
- with common techniques often ~30% of data is not picked (and forms OOB set)
- distribution of classes with bootstrapping will form a Gaussian distribution
- when enough samples are used bootstrap distributions will approximate population statistics
- Assumption is made with bootstrapping that the sample approximates the original population
- bootstrap samples are used as training sets, OOB samples serve as testing sets
bootstrap_samples = [
resample(df, replace=True, n_samples=n random_state=i) for i in range(5)
]
- gets a list of dataframes with sampled data
- ensemble models can often greatly improve performance
- decision trees are often used as bases
- can increase computational complexity
- parallel ensembles
- base models are trained independently
- training is faster
- sequential ensembles
- models trained iteratively, adjusting for previous errors
- bagging
- (bootstrap aggregation)
- create bootstrap sets, train weak models and then aggregate predictions to get more accurate prediction
- decision trees in bagging have low bias but high variance
- bagging reduces model variance, not data variance
- (how much does model change if data changes slightly)