vault backup: 2026-06-04 18:28:57
This commit is contained in:
1 parent
99cf851191
commit
cda5d68cb5
1 file changed
+12
-1
+12
-1
@@ -1,4 +1,15 @@
|
||||
#rs/class/csb320 #rs/notes
|
||||
- - -
|
||||
- Models are trained by learning from their mistakes
|
||||
-
|
||||
- multi-armed bandits
|
||||
- choose action from k possibilities
|
||||
- receive a reward
|
||||
- reward is dependent on the action taken
|
||||
- explore vs. exploit
|
||||
- exploit is taking the greedy action, the on that is known to produce the greatest reward
|
||||
- explore is taking other random actions to learn values
|
||||
- a combination of the two produce the best results
|
||||
- $\varepsilon$-greedy strategy
|
||||
- with probability $\varepsilon$ take a random action, with probability $1-\varepsilon$ take the best known actions
|
||||
- $\varepsilon$ often starts high and decreases over time
|
||||
-
|
||||
Reference in new issue
Block a user