39 lines
2.0 KiB
Markdown
39 lines
2.0 KiB
Markdown
#rs/class/csb320 #rs/notes
|
|
- - -
|
|
- Models are trained by learning from their mistakes
|
|
- multi-armed bandits
|
|
- choose action from k possibilities
|
|
- receive a reward
|
|
- reward is dependent on the action taken
|
|
- explore vs. exploit
|
|
- exploit is taking the greedy action, the on that is known to produce the greatest reward
|
|
- explore is taking other random actions to learn values
|
|
- a combination of the two produce the best results
|
|
- $\varepsilon$-greedy strategy
|
|
- with probability $\varepsilon$ take a random action, with probability $1-\varepsilon$ take the best known actions
|
|
- $\varepsilon$ often starts high and decreases over time
|
|
- $\rho$ (regret) can be used to find how good a strategy is
|
|
- this is the difference between how much reward was gotten and the maximum reward possible if distributions are known ahead of time
|
|
- markov decision process
|
|
- future states only depend on the present state, not what came before
|
|
- trying to maximize reward without knowing entire history
|
|
- monte-carlo
|
|
- updates model after every episode
|
|
- temporal difference
|
|
- learns after every step, not every episode
|
|
- q-learning
|
|
- assembles all possible q values on the way to end reward
|
|
- updates q values to find best path
|
|
- $$Q^{new}(s_{t}, a_{t}) \leftarrow (1 - a) * Q(s_{t}, a_{t}) + a * (r_{t} + \gamma * max Q(s_{t+1}, a))$$
|
|
- ![[BellmanEquation.excalidraw]]
|
|
- This equation can be used to update the q values of the grid
|
|
- it takes the old value and adds a learned value to it
|
|
- this allows the algorithm to slowly "learn" the best path
|
|
- the learned value takes the reward from the max step from the *next* square
|
|
- $Q(s, a)$ is the quality of taking action $a$ from state $s$
|
|
- $r$ is the immediate reward after taking the action
|
|
- $\gamma$ is the discount factor (0-1)
|
|
- This is what prioritizes future rewards vs immediate rewards
|
|
- future rewards (the $max$ part) are deprioritized in relation to immediate rewards
|
|
- $maxQ(s_{t+1}, a)$ is the best q value from the next state (future reward)
|
|
- this is what backpropagates future rewards |