From cda5d68cb5080e37ecd6383957f165e2a428bac2 Mon Sep 17 00:00:00 2001 From: ben-jaynes <1btjaynes@gmail.com> Date: Thu, 4 Jun 2026 18:28:57 -0700 Subject: [PATCH] vault backup: 2026-06-04 18:28:57 --- .../Class 6-4 (Reinforcement Learning).md | 13 ++++++++++++- 1 file changed, 12 insertions(+), 1 deletion(-) diff --git a/Running Start/CSB320 - Machine Learning Concepts/Class 6-4 (Reinforcement Learning).md b/Running Start/CSB320 - Machine Learning Concepts/Class 6-4 (Reinforcement Learning).md index 56d8f1d..1c5028e 100644 --- a/Running Start/CSB320 - Machine Learning Concepts/Class 6-4 (Reinforcement Learning).md +++ b/Running Start/CSB320 - Machine Learning Concepts/Class 6-4 (Reinforcement Learning).md @@ -1,4 +1,15 @@ #rs/class/csb320 #rs/notes - - - - Models are trained by learning from their mistakes -- \ No newline at end of file +- multi-armed bandits + - choose action from k possibilities + - receive a reward + - reward is dependent on the action taken + - explore vs. exploit + - exploit is taking the greedy action, the on that is known to produce the greatest reward + - explore is taking other random actions to learn values + - a combination of the two produce the best results + - $\varepsilon$-greedy strategy + - with probability $\varepsilon$ take a random action, with probability $1-\varepsilon$ take the best known actions + - $\varepsilon$ often starts high and decreases over time + - \ No newline at end of file