vault backup: 2026-06-04 19:19:24
This commit is contained in:
1 parent
c86a2d02b9
commit
4d1ea0de46
1 file changed
+5
@@ -21,3 +21,8 @@
|
||||
- updates model after every episode
|
||||
- temporal difference
|
||||
- learns after every step, not every episode
|
||||
- q-learning
|
||||
- assembles all possible q values on the way to end reward
|
||||
- updates q values to find best path
|
||||
- $$Q^{new}(s_{t}, a_{t}) \leftarrow (1 - a) * Q(s_{t}, a_{t}) + a * (r_{t} + \gamma * max_{a} Q(s_{t+1}, a))$$
|
||||
-
|
||||
Reference in new issue
Block a user