arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06667cs.LG

行为克隆优于熵正则化强化学习:批评者驱动的演员-评论家方法在自适应肿瘤治疗中的失败

Behavioral Cloning Outperforms Entropy-Regularized RL: Critic-Driven Failure of Actor-Critic Methods on Adaptive Tumor Treatment

  • Tilburg University(蒂尔堡大学)

机构由 AI 辅助整理,请以论文原文为准。

Aleksandar Dimitrov, Giacomo Spigler

中文总结 AI 辅助

本研究通过最优控制参考证明,行为克隆优于熵正则化强化学习,而演员-评论家方法因批评者误导在自适应肿瘤治疗中失败。

中文摘要 AI 辅助

自适应给药需要能够在不过度产生毒性的情况下减少肿瘤负担的策略。学习到的给药策略通常与历史或启发式比较器进行对比评估,这无法表明一个策略是否找到了可用的最佳行为。我们转而研究一个三群体肿瘤控制常微分方程,其中最优控制分析确定了良好调度形式——由奇异弧分隔的开关式给药——并构建该形式的数值控制器作为近似最优行为的代理。在持续治愈标准(连续200天低于5%承载能力)下与此参考对比评估时,从头训练的软演员-评论家(SAC)从未达到治愈。对参考的行为克隆(BC)能够复现它(100%持续治愈,30/30个种子),但对克隆策略的SAC微调在五个熵系数下均破坏了它,且TD3和BC正则化SAC同样失败;该模式在乘法药代动力学作用噪声下持续存在。沿治愈轨迹,崩溃后的批评者在96%的状态中将崩溃策略的动作排在参考动作之上,集中在维持阶段,且策略陷入非治愈的自适应治疗均衡。参考使得这一点清晰可见:与启发式比较器相比,微调后的策略会被视为称职的控制器而非失败。

英文摘要

Adaptive dosing requires policies that reduce tumor burden without excessive toxicity. Learned dosing policies are typically judged against historical or heuristic comparators, which cannot show whether a policy has found the best behavior available. We instead study a three-population tumor-control ODE in which optimal-control analysis fixes the form of a good schedule -- bang-bang dosing punctuated by a singular arc -- and construct a numerical controller of that form as a proxy for near-optimal behavior. Judged against this reference under a sustained-cure criterion -- 200 consecutive days below 5% carrying capacity -- Soft Actor-Critic (SAC) trained from scratch never reaches cure. Behavioral cloning (BC) of the reference reproduces it (100% sustained cure, 30/30 seeds), but SAC fine-tuning of the cloned policy destroys it across five entropy coefficients, and TD3 and BC-regularized SAC fail identically; the pattern persists under multiplicative pharmacokinetic action noise. Along curative trajectories the post-collapse critic ranks the collapsed-policy action above the reference action in 96% of states, concentrated in the maintenance phase, and the policy settles into a non-curative adaptive-therapy equilibrium. The reference is what makes this legible: against a heuristic comparator the fine-tuned policy would read as a competent controller rather than a failure.

补充信息

↑