Exercise 5.13 Show the steps to derive (5.14) from (5.12).
ρt:T−1Rt+1=π(At∣St)b(At∣St)π(At+1∣St+1)b(At+1∣St+1)π(At+2∣St+2)b(At+2∣St+2)⋯π(AT−1∣ST−1)b(AT−1∣ST−1)Rt+1(5.12)
\rho_{t:T-1}R_{t+1} = \frac{\pi(A_t \mid S_t)}{b(A_t \mid S_t)}\frac{\pi(A_{t+1} \mid S_{t + 1})}{b(A_{t+1} \mid S_{t + 1})}\frac{\pi(A_{t+2} \mid S_{t + 2})}{b(A_{t+2} \mid S_{t + 2})}\cdots\frac{\pi(A_{T-1} \mid S_{T - 1})}{b(A_{T-1} \mid S_{T - 1})}R_{t+1} \qquad{(5.12)}
ρt:T−1Rt+1=b(At∣St)π(At∣St)b(At+1∣St+1)π(At+1∣St+1)b(At+2∣St+2)π(At+2∣St+2)⋯b(AT−1∣ST−1)π(AT−1∣ST−1)Rt+1(5.12)
According to the property of Markov Procedure, each action step is independent, so we have
E(ρt:T−1Rt+1)=E(π(At∣St)b(At∣St)Rt+1)E(π(At+1∣St+1)b(At+1∣St+1))E(π(At+2∣St+2)b(At+2∣St+2))⋯E(π(AT−1∣ST−1)b(AT−1∣ST−1))
\mathbb E(\rho_{t:T-1}R_{t+1})=\mathbb E( \frac{\pi(A_t \mid S_t)}{b(A_t \mid S_t)}R_{t+1})\mathbb E(\frac{\pi(A_{t+1} \mid S_{t + 1})}{b(A_{t+1} \mid S_{t + 1})})\mathbb E(\frac{\pi(A_{t+2} \mid S_{t + 2})}{b(A_{t+2} \mid S_{t + 2})})\cdots\mathbb E(\frac{\pi(A_{T-1} \mid S_{T - 1})}{b(A_{T-1} \mid S_{T - 1})})
E(ρt:T−1Rt+1)=E(b(At∣St)π(At∣St)Rt+1)E(b(At+1∣St+1)π(At+1∣St+1))E(b(At+2∣St+2)π(At+2∣St+2))⋯E(b(AT−1∣ST−1)π(AT−1∣ST−1))
And according to equation (5.13), the expectation of each step is 1, except the first.
E[π(Ak∣Sk)b(Ak∣Sk)]≐∑ab(a∣Sk)π(a∣Sk)b(a∣Sk)=∑aπ(a∣Sk)=1(5.13)
\mathbb E\Biggl[ \frac{\pi(A_k \mid S_k)}{b(A_k \mid S_k)}\Biggr] \doteq \sum_a b(a\mid S_k) \frac{\pi(a \mid S_k)}{b(a \mid S_k)} = \sum_a \pi(a \mid S_k) = 1 \qquad{(5.13)}
E[b(Ak∣Sk)π(Ak∣Sk)]≐a∑b(a∣Sk)b(a∣Sk)π(a∣Sk)=a∑π(a∣Sk)=1(5.13)
So,
E(ρt:T−1Rt+1)=E(π(At∣St)b(At∣St)Rt+1)=E(ρt:tRt+1)(5.14)
\begin{aligned}
\mathbb E(\rho_{t:T-1}R_{t+1})&=\mathbb E( \frac{\pi(A_t \mid S_t)}{b(A_t \mid S_t)}R_{t+1})\\
&=\mathbb E(\rho_{t:t}R_{t+1}) \qquad \qquad \qquad \qquad \qquad \qquad \qquad \qquad{(5.14)}
\end{aligned}
E(ρt:T−1Rt+1)=E(b(At∣St)π(At∣St)Rt+1)=E(ρt:tRt+1)(5.14)
Reinforcement Learning Exercise 5.13
最新推荐文章于 2026-04-18 09:20:44 发布
本文详细解析了强化学习中从公式(5.12)到(5.14)的推导过程,展示了如何利用马尔科夫决策过程的独立性属性,将复杂策略评估表达式简化为更简单的形式。

1237

被折叠的 条评论
为什么被折叠?



