vault backup: 2026-06-04 18:49:04
This commit is contained in:
1 parent
c1595eb37f
commit
b46d6795a3
1 file changed
+3
@@ -14,4 +14,7 @@
|
|||||||
- $\varepsilon$ often starts high and decreases over time
|
- $\varepsilon$ often starts high and decreases over time
|
||||||
- $\rho$ (regret) can be used to find how good a strategy is
|
- $\rho$ (regret) can be used to find how good a strategy is
|
||||||
- this is the difference between how much reward was gotten and the maximum reward possible if distributions are known ahead of time
|
- this is the difference between how much reward was gotten and the maximum reward possible if distributions are known ahead of time
|
||||||
|
- markov decision process
|
||||||
|
- future states only depend on the present state, not what came before
|
||||||
|
- trying to maximize reward without knowing entire history
|
||||||
-
|
-
|
||||||
Reference in new issue
Block a user