- The average of guesses of a large number of people will typically be more accurate than an individual can be - (wisdom of crowds) - Bootstrapping - generates simulated samples by sampling with replacement from existing sample - sampling is random - Out of bag sample - data that was never picked in bootstrapping - is often used for testing data (OOB samples were not in the training samples so the model has never seen them) - with common techniques often ~30% of data is not picked (and forms OOB set) - distribution of classes with bootstrapping will form a Gaussian distribution - when enough samples are used bootstrap distributions will approximate population statistics - Assumption is made with bootstrapping that the sample approximates the original population - bootstrap samples are used as training sets, OOB samples serve as testing sets ```python bootstrap_samples = [ resample(df, replace=True, n_samples=n random_state=i) for i in range(5) ] ``` - gets a list of dataframes with sampled data - ensemble models can often greatly improve performance - decision trees are often used as bases - can increase computational complexity - parallel ensembles - base models are trained independently - training is faster - sequential ensembles - models trained iteratively, adjusting for previous errors - bagging - (bootstrap aggregation) - create bootstrap sets, train weak models and then aggregate predictions to get more accurate prediction - decision trees in bagging have low bias but high variance - bagging reduces model variance, not data variance - (how much does model change if data changes slightly) - bagging helps to make predictions more stable - model variance should be low to increase prediction stability, data variance should be high so patterns are easier to find - 10-25 decision trees are typically sufficient - too many results in diminishing returns - bagging gives high accuracy and better generalization but can be hard to interpret and have large computational costs - random forests - uses decision trees trained on feature subsets - predictions are aggregated using averaging (regression) or majority vote (classification) - features can also be selected at random as well as instances to decrease correlation between trees - this can also increase computational efficiency - prevents single features from dominating - boosting - base models fit iteratively, adjusting for previous error - uses same training data without bootstrap sampling - assigns higher weight to misclassified instances - gradient boosting - tries to minimize the gradient (derivative of the loss function) in each iteration - learning rate determines model update strength (step sizes) - ![[GradientDescent.excalidraw]] ## Zybooks notes -