vault backup: 2026-06-04 19:19:24

This commit is contained in:
ben committed 2026-06-04 19:19:24 -07:00
1 parent c86a2d02b9
commit 4d1ea0de46
1 file changed
+5
@@ -21,3 +21,8 @@
- updates model after every episode
- temporal difference
- learns after every step, not every episode
- q-learning
- assembles all possible q values on the way to end reward
- updates q values to find best path
- $$Q^{new}(s_{t}, a_{t}) \leftarrow (1 - a) * Q(s_{t}, a_{t}) + a * (r_{t} + \gamma * max_{a} Q(s_{t+1}, a))$$
-