vault backup: 2026-06-04 19:39:32
This commit is contained in:
1 parent
a875760823
commit
c8c09c4d8b
2 files changed
+139
-186
No files matched your search
+11
-2
@@ -24,7 +24,16 @@
|
||||
- q-learning
|
||||
- assembles all possible q values on the way to end reward
|
||||
- updates q values to find best path
|
||||
- $$Q^{new}(s_{t}, a_{t}) \leftarrow (1 - a) * Q(s_{t}, a_{t}) + a * (r_{t} + \gamma * max_{a} Q(s_{t+1}, a))$$
|
||||
- $$Q^{new}(s_{t}, a_{t}) \leftarrow (1 - a) * Q(s_{t}, a_{t}) + a * (r_{t} + \gamma * max Q(s_{t+1}, a))$$
|
||||
- ![[BellmanEquation.excalidraw]]
|
||||
- This equation can be used to update the q values of the grid
|
||||
- it takes the old value and adds
|
||||
- it takes the old value and adds a learned value to it
|
||||
- this allows the algorithm to slowly "learn" the best path
|
||||
- the learned value takes the reward from the max step from the *next* square
|
||||
- $Q(s, a)$ is the quality of taking action $a$ from state $s$
|
||||
- $r$ is the immediate reward after taking the action
|
||||
- $\gamma$ is the discount factor (0-1)
|
||||
- This is what prioritizes future rewards vs immediate rewards
|
||||
- future rewards (the $max$ part) are deprioritized in relation to immediate rewards
|
||||
- $maxQ(s_{t+1}, a)$ is the best q value from the next state (future reward)
|
||||
- this is what backpropagates future rewards
|
||||
Reference in new issue
Block a user