2.1 KiB
2.1 KiB
- The average of guesses of a large number of people will typically be more accurate than an individual can be
- (wisdom of crowds)
- Bootstrapping
- generates simulated samples by sampling with replacement from existing sample
- sampling is random
- Out of bag sample
- data that was never picked in bootstrapping
- is often used for testing data (OOB samples were not in the training samples so the model has never seen them)
- with common techniques often ~30% of data is not picked (and forms OOB set)
- distribution of classes with bootstrapping will form a Gaussian distribution
- when enough samples are used bootstrap distributions will approximate population statistics
- Assumption is made with bootstrapping that the sample approximates the original population
- bootstrap samples are used as training sets, OOB samples serve as testing sets
bootstrap_samples = [
resample(df, replace=True, n_samples=n random_state=i) for i in range(5)
]
- gets a list of dataframes with sampled data
- ensemble models can often greatly improve performance
- decision trees are often used as bases
- can increase computational complexity
- parallel ensembles
- base models are trained independently
- training is faster
- sequential ensembles
- models trained iteratively, adjusting for previous errors
- bagging
- (bootstrap aggregation)
- create bootstrap sets, train weak models and then aggregate predictions to get more accurate prediction
- decision trees in bagging have low bias but high variance
- bagging reduces model variance, not data variance
- (how much does model change if data changes slightly)
- bagging helps to make predictions more stable
- model variance should be low to increase prediction stability, data variance should be high so patterns are easier to find
- 10-25 decision trees are typically sufficient
- too many results in diminishing returns
- bagging gives high accuracy and better generalization but can be hard to interpret and have large computational costs
- random forests
- uses decision trees trained on feature subsets
- predictions are aggregated using averaging (regression) or majority vote (classification)