diff --git a/.obsidian/graph.json b/.obsidian/graph.json index fe44a4b..1029bae 100644 --- a/.obsidian/graph.json +++ b/.obsidian/graph.json @@ -17,6 +17,6 @@ "repelStrength": 10, "linkStrength": 1, "linkDistance": 250, - "scale": 0.1719770322750558, + "scale": 0.18153036664097952, "close": true } \ No newline at end of file diff --git a/Running Start/CSB320 - Machine Learning Concepts/Class 5-21 (Ensemble Models).md b/Running Start/CSB320 - Machine Learning Concepts/Class 5-21 (Ensemble Models).md index a9daa0a..29da032 100644 --- a/Running Start/CSB320 - Machine Learning Concepts/Class 5-21 (Ensemble Models).md +++ b/Running Start/CSB320 - Machine Learning Concepts/Class 5-21 (Ensemble Models).md @@ -12,5 +12,23 @@ - Assumption is made with bootstrapping that the sample approximates the original population - bootstrap samples are used as training sets, OOB samples serve as testing sets ```python -bootstrap_samples = [resample(df, replace=True, random_state=i) for i in range(5)] +bootstrap_samples = [ + resample(df, replace=True, n_samples=n random_state=i) for i in range(5) +] ``` +- gets a list of dataframes with sampled data +- ensemble models can often greatly improve performance + - decision trees are often used as bases + - can increase computational complexity +- parallel ensembles + - base models are trained independently + - training is faster +- sequential ensembles + - models trained iteratively, adjusting for previous errors +- bagging + - (bootstrap aggregation) + - create bootstrap sets, train weak models and then aggregate predictions to get more accurate prediction + - decision trees in bagging have low bias but high variance + - bagging reduces model variance, not data variance + - (how much does model change if data changes slightly) + - \ No newline at end of file