#rs/class/csb320 #rs/notes - - - - Models are trained by learning from their mistakes - multi-armed bandits - choose action from k possibilities - receive a reward - reward is dependent on the action taken - explore vs. exploit - exploit is taking the greedy action, the on that is known to produce the greatest reward - explore is taking other random actions to learn values - a combination of the two produce the best results - $\varepsilon$-greedy strategy - with probability $\varepsilon$ take a random action, with probability $1-\varepsilon$ take the best known actions - $\varepsilon$ often starts high and decreases over time - $\rho$ (regret) can be used to find how good a strategy is - this is the difference between how much reward was gotten and the maximum reward possible if distributions are known ahead of time - markov decision process - future states only depend on the present state, not what came before - trying to maximize reward without knowing entire history - monte-carlo - updates model after every episode - temporal difference - learns after every step, not every episode - q-learning - assembles all possible q values on the way to end reward - updates q values to find best path - $$Q^{new}(s_{t}, a_{t}) \leftarrow (1 - a) * Q(s_{t}, a_{t}) + a * (r_{t} + \gamma * max Q(s_{t+1}, a))$$ - ![[BellmanEquation.excalidraw]] - This equation can be used to update the q values of the grid - it takes the old value and adds a learned value to it - this allows the algorithm to slowly "learn" the best path - the learned value takes the reward from the max step from the *next* square - $Q(s, a)$ is the quality of taking action $a$ from state $s$ - $r$ is the immediate reward after taking the action - $\gamma$ is the discount factor (0-1) - This is what prioritizes future rewards vs immediate rewards - future rewards (the $max$ part) are deprioritized in relation to immediate rewards - $maxQ(s_{t+1}, a)$ is the best q value from the next state (future reward) - this is what backpropagates future rewards - during exploration phase paths are usually explored randomly to attempt to find the q values for each path -