From b46d6795a32676e759328d7aa1b548e6e3c7dc44 Mon Sep 17 00:00:00 2001 From: ben-jaynes <1btjaynes@gmail.com> Date: Thu, 4 Jun 2026 18:49:04 -0700 Subject: [PATCH] vault backup: 2026-06-04 18:49:04 --- .../Class 6-4 (Reinforcement Learning).md | 3 +++ 1 file changed, 3 insertions(+) diff --git a/Running Start/CSB320 - Machine Learning Concepts/Class 6-4 (Reinforcement Learning).md b/Running Start/CSB320 - Machine Learning Concepts/Class 6-4 (Reinforcement Learning).md index 81c4a3f..4dce4b2 100644 --- a/Running Start/CSB320 - Machine Learning Concepts/Class 6-4 (Reinforcement Learning).md +++ b/Running Start/CSB320 - Machine Learning Concepts/Class 6-4 (Reinforcement Learning).md @@ -14,4 +14,7 @@ - $\varepsilon$ often starts high and decreases over time - $\rho$ (regret) can be used to find how good a strategy is - this is the difference between how much reward was gotten and the maximum reward possible if distributions are known ahead of time +- markov decision process + - future states only depend on the present state, not what came before + - trying to maximize reward without knowing entire history - \ No newline at end of file