尧图网络科技YAOTU DIGITAL 获取报价
获取报价
首页 / 资讯中心 / 文章详情

RL-赵-(八)-ValueBased03-ActionValue估算:Q-learning函数逼近算法【目标:计算出最优“值函数”参数,通过该“值函数”计算出的Action Value最优】

发布时间:2026/9/26 6:35:06

资讯中心
01
ARTICLE

RL-赵-(八)-ValueBased03-ActionValue估算:Q-learning函数逼近算法【目标:计算出最优“值函数”参数,通过该“值函数”计算出的Action Value最优】

RL-赵-(八)-ValueBased03-ActionValue估算:Q-learning函数逼近算法【目标:计算出最优“值函数”参数,通过该“值函数”计算出的Action Value最优】
我们知道:“TD learning”with“value function approximate”:w t + 1 = w t + α t [ r t + 1 + γ v ^ ( s t + 1 , w t ) − v ^ ( s t , w t ) ] ∇ w v ^ ( s t , w t ) \color{red}{w_{t+1}=w_t+\alpha_t\left[r_{t+1}+\gamma\hat{v}(s_{t+1},w_t)-\hat{v}(s_t,w_t)\right]\nabla_w\hat{v}(s_t,w_t)}wt+1​=wt​+αt​[rt+1​+γv^(st+1​,wt​)−v^(st​,wt​)]∇w​v^(st​,wt​)“Sarsa算法”with“value function approximate”:w t + 1 = w t + α t [ r t + 1 + γ q ^ ( s t + 1 , a t + 1 , w t ) − q ^ ( s t , a t , w t ) ] ∇ w q ^ ( s t , a t , w t ) \color{red}{w_{t+1}=w_t+\alpha_t\left[r_{t+1}+\gamma\hat{q}(s_{t+1},a_{t+1},w_t)-\hat{q}(s_t,a_t,w_t)\right]\nabla_w\hat{q}(s_t,a_t,w_t)}wt+1​=wt​+αt​[rt+1​
02
RELATED NEWS

相关资讯

更多网站建设与数字化升级内容

03
WHY YAOTU

想打造同款高转化官网?

懂行业、懂生意,从建站到增长一站式陪跑

◈

场景化定制

不做模板站,围绕你的业务场景量身设计,小众不撞款。

◐

营销型架构

以转化目标组织内容与路径,让官网真正带来询盘。

▲

全周期服务

设计、开发、运营、运维一体,上线只是开始。

免费获取你的建站方案

留下需求,专属顾问 24 小时内为你输出方案建议。