Files
ObsidianVault/Running Start/CSB320 - Machine Learning Concepts/Class 6-4 (Reinforcement Learning).md
T
2026-06-04 19:19:24 -07:00

1.3 KiB

#rs/class/csb320 #rs/notes


  • Models are trained by learning from their mistakes
  • multi-armed bandits
    • choose action from k possibilities
      • receive a reward
    • reward is dependent on the action taken
    • explore vs. exploit
      • exploit is taking the greedy action, the on that is known to produce the greatest reward
      • explore is taking other random actions to learn values
      • a combination of the two produce the best results
    • $\varepsilon$-greedy strategy
      • with probability \varepsilon take a random action, with probability 1-\varepsilon take the best known actions
      • \varepsilon often starts high and decreases over time
    • \rho (regret) can be used to find how good a strategy is
      • this is the difference between how much reward was gotten and the maximum reward possible if distributions are known ahead of time
  • markov decision process
    • future states only depend on the present state, not what came before
    • trying to maximize reward without knowing entire history
  • monte-carlo
    • updates model after every episode
  • temporal difference
    • learns after every step, not every episode
  • q-learning
    • assembles all possible q values on the way to end reward
    • updates q values to find best path
    • Q^{new}(s_{t}, a_{t}) \leftarrow (1 - a) * Q(s_{t}, a_{t}) + a * (r_{t} + \gamma * max_{a} Q(s_{t+1}, a))