From c1595eb37f1bc7f6a77da98252b0bbacfca0364f Mon Sep 17 00:00:00 2001 From: ben-jaynes <1btjaynes@gmail.com> Date: Thu, 4 Jun 2026 18:39:00 -0700 Subject: [PATCH] vault backup: 2026-06-04 18:39:00 --- .../Class 6-4 (Reinforcement Learning).md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/Running Start/CSB320 - Machine Learning Concepts/Class 6-4 (Reinforcement Learning).md b/Running Start/CSB320 - Machine Learning Concepts/Class 6-4 (Reinforcement Learning).md index 1c5028e..81c4a3f 100644 --- a/Running Start/CSB320 - Machine Learning Concepts/Class 6-4 (Reinforcement Learning).md +++ b/Running Start/CSB320 - Machine Learning Concepts/Class 6-4 (Reinforcement Learning).md @@ -12,4 +12,6 @@ - $\varepsilon$-greedy strategy - with probability $\varepsilon$ take a random action, with probability $1-\varepsilon$ take the best known actions - $\varepsilon$ often starts high and decreases over time + - $\rho$ (regret) can be used to find how good a strategy is + - this is the difference between how much reward was gotten and the maximum reward possible if distributions are known ahead of time - \ No newline at end of file