Files
ObsidianVault/Running Start/CSB320 - Machine Learning Concepts/Class 6-4 (Reinforcement Learning).md
T
2026-06-04 19:39:32 -07:00

2.0 KiB

#rs/class/csb320 #rs/notes


  • Models are trained by learning from their mistakes
  • multi-armed bandits
    • choose action from k possibilities
      • receive a reward
    • reward is dependent on the action taken
    • explore vs. exploit
      • exploit is taking the greedy action, the on that is known to produce the greatest reward
      • explore is taking other random actions to learn values
      • a combination of the two produce the best results
    • $\varepsilon$-greedy strategy
      • with probability \varepsilon take a random action, with probability 1-\varepsilon take the best known actions
      • \varepsilon often starts high and decreases over time
    • \rho (regret) can be used to find how good a strategy is
      • this is the difference between how much reward was gotten and the maximum reward possible if distributions are known ahead of time
  • markov decision process
    • future states only depend on the present state, not what came before
    • trying to maximize reward without knowing entire history
  • monte-carlo
    • updates model after every episode
  • temporal difference
    • learns after every step, not every episode
  • q-learning
    • assembles all possible q values on the way to end reward
    • updates q values to find best path
    • Q^{new}(s_{t}, a_{t}) \leftarrow (1 - a) * Q(s_{t}, a_{t}) + a * (r_{t} + \gamma * max Q(s_{t+1}, a))
    • !BellmanEquation.excalidraw
    • This equation can be used to update the q values of the grid
    • it takes the old value and adds a learned value to it
      • this allows the algorithm to slowly "learn" the best path
    • the learned value takes the reward from the max step from the next square
    • Q(s, a) is the quality of taking action a from state s
    • r is the immediate reward after taking the action
    • \gamma is the discount factor (0-1)
      • This is what prioritizes future rewards vs immediate rewards
      • future rewards (the max part) are deprioritized in relation to immediate rewards
    • maxQ(s_{t+1}, a) is the best q value from the next state (future reward)
      • this is what backpropagates future rewards