1009 B
1009 B
#rs/class/csb320 #rs/notes
- Models are trained by learning from their mistakes
- multi-armed bandits
- choose action from k possibilities
- receive a reward
- reward is dependent on the action taken
- explore vs. exploit
- exploit is taking the greedy action, the on that is known to produce the greatest reward
- explore is taking other random actions to learn values
- a combination of the two produce the best results
- $\varepsilon$-greedy strategy
- with probability
\varepsilontake a random action, with probability1-\varepsilontake the best known actions \varepsilonoften starts high and decreases over time
- with probability
\rho(regret) can be used to find how good a strategy is- this is the difference between how much reward was gotten and the maximum reward possible if distributions are known ahead of time
- choose action from k possibilities
- markov decision process
- future states only depend on the present state, not what came before
- trying to maximize reward without knowing entire history