vault backup: 2026-06-04 18:39:00
This commit is contained in:
1 parent
cda5d68cb5
commit
c1595eb37f
1 file changed
+2
@@ -12,4 +12,6 @@
|
||||
- $\varepsilon$-greedy strategy
|
||||
- with probability $\varepsilon$ take a random action, with probability $1-\varepsilon$ take the best known actions
|
||||
- $\varepsilon$ often starts high and decreases over time
|
||||
- $\rho$ (regret) can be used to find how good a strategy is
|
||||
- this is the difference between how much reward was gotten and the maximum reward possible if distributions are known ahead of time
|
||||
-
|
||||
Reference in new issue
Block a user