2.0 KiB
2.0 KiB
#rs/class/csb320 #rs/notes
- Models are trained by learning from their mistakes
- multi-armed bandits
- choose action from k possibilities
- receive a reward
- reward is dependent on the action taken
- explore vs. exploit
- exploit is taking the greedy action, the on that is known to produce the greatest reward
- explore is taking other random actions to learn values
- a combination of the two produce the best results
- $\varepsilon$-greedy strategy
- with probability
\varepsilontake a random action, with probability1-\varepsilontake the best known actions \varepsilonoften starts high and decreases over time
- with probability
\rho(regret) can be used to find how good a strategy is- this is the difference between how much reward was gotten and the maximum reward possible if distributions are known ahead of time
- choose action from k possibilities
- markov decision process
- future states only depend on the present state, not what came before
- trying to maximize reward without knowing entire history
- monte-carlo
- updates model after every episode
- temporal difference
- learns after every step, not every episode
- q-learning
- assembles all possible q values on the way to end reward
- updates q values to find best path
-
Q^{new}(s_{t}, a_{t}) \leftarrow (1 - a) * Q(s_{t}, a_{t}) + a * (r_{t} + \gamma * max Q(s_{t+1}, a)) - !BellmanEquation.excalidraw
- This equation can be used to update the q values of the grid
- it takes the old value and adds a learned value to it
- this allows the algorithm to slowly "learn" the best path
- the learned value takes the reward from the max step from the next square
Q(s, a)is the quality of taking actionafrom statesris the immediate reward after taking the action\gammais the discount factor (0-1)- This is what prioritizes future rewards vs immediate rewards
- future rewards (the
maxpart) are deprioritized in relation to immediate rewards
maxQ(s_{t+1}, a)is the best q value from the next state (future reward)- this is what backpropagates future rewards