Files
ObsidianVault/Running Start/CSB320 - Machine Learning Concepts/Class 6-4 (Reinforcement Learning).md
T
2026-06-04 18:28:57 -07:00

15 lines
654 B
Markdown

#rs/class/csb320 #rs/notes
- - -
- Models are trained by learning from their mistakes
- multi-armed bandits
- choose action from k possibilities
- receive a reward
- reward is dependent on the action taken
- explore vs. exploit
- exploit is taking the greedy action, the on that is known to produce the greatest reward
- explore is taking other random actions to learn values
- a combination of the two produce the best results
- $\varepsilon$-greedy strategy
- with probability $\varepsilon$ take a random action, with probability $1-\varepsilon$ take the best known actions
- $\varepsilon$ often starts high and decreases over time
-