vault backup: 2026-06-04 19:49:35
This commit is contained in:
1 parent
c8c09c4d8b
commit
3bccadc873
2 files changed
+6
-4
No files matched your search
+3
-1
@@ -36,4 +36,6 @@
|
||||
- This is what prioritizes future rewards vs immediate rewards
|
||||
- future rewards (the $max$ part) are deprioritized in relation to immediate rewards
|
||||
- $maxQ(s_{t+1}, a)$ is the best q value from the next state (future reward)
|
||||
- this is what backpropagates future rewards
|
||||
- this is what backpropagates future rewards
|
||||
- during exploration phase paths are usually explored randomly to attempt to find the q values for each path
|
||||
-
|
||||
Reference in new issue
Block a user