17 lines
851 B
Markdown
17 lines
851 B
Markdown
#rs/class/csb320 #rs/notes
|
|
- - -
|
|
- Models are trained by learning from their mistakes
|
|
- multi-armed bandits
|
|
- choose action from k possibilities
|
|
- receive a reward
|
|
- reward is dependent on the action taken
|
|
- explore vs. exploit
|
|
- exploit is taking the greedy action, the on that is known to produce the greatest reward
|
|
- explore is taking other random actions to learn values
|
|
- a combination of the two produce the best results
|
|
- $\varepsilon$-greedy strategy
|
|
- with probability $\varepsilon$ take a random action, with probability $1-\varepsilon$ take the best known actions
|
|
- $\varepsilon$ often starts high and decreases over time
|
|
- $\rho$ (regret) can be used to find how good a strategy is
|
|
- this is the difference between how much reward was gotten and the maximum reward possible if distributions are known ahead of time
|
|
- |