vault backup: 2026-06-04 18:39:00
This commit is contained in:
1 parent
cda5d68cb5
commit
c1595eb37f
1 file changed
+2
@@ -12,4 +12,6 @@
|
|||||||
- $\varepsilon$-greedy strategy
|
- $\varepsilon$-greedy strategy
|
||||||
- with probability $\varepsilon$ take a random action, with probability $1-\varepsilon$ take the best known actions
|
- with probability $\varepsilon$ take a random action, with probability $1-\varepsilon$ take the best known actions
|
||||||
- $\varepsilon$ often starts high and decreases over time
|
- $\varepsilon$ often starts high and decreases over time
|
||||||
|
- $\rho$ (regret) can be used to find how good a strategy is
|
||||||
|
- this is the difference between how much reward was gotten and the maximum reward possible if distributions are known ahead of time
|
||||||
-
|
-
|
||||||
Reference in new issue
Block a user