arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12716stat.ME

递归Q学习的反馈感知调优

Feedback-Aware Tuning of Recursive Q-Learning

Masahiro Kojima

首次发表
浏览论文内容

中文总结 AI 辅助

针对反向Q学习模型选择中的递归反馈问题,提出反馈感知软调优方法,以整个反向拟合规则为比较单元,建立风险恒等式与oracle不等式,并在模拟试验中验证其性能。

中文摘要 AI 辅助

在反向Q学习中进行模型选择是递归的,因为后期阶段的选择会改变提供给早期回归的响应,并可能改变其模型比较统计量。分阶段的独立准则不能直接评估完整Q学习拟合的目标阶段预测风险。我们通过将整个反向拟合规则视为比较单元来解决这一不匹配问题,并为序贯多重分配随机试验提出了反馈感知的软调优方法。每次反向拟合使用其自身生成的响应,并在共同的预测目标上进行评估。风险准则保留了下游效应对上游比较的影响,而单独的校正则考虑了从相同观测中估计最终指数权重的问题。对于固定的有限光滑递归映射库,我们在高斯移位模型下建立了精确的风险恒等式,并给出了带有显式自适应余项的oracle不等式。在坐标表示和矩条件下,这些保证可转移到任何预定阶段的预测风险,其中阶段数随样本量增加而固定。一个两阶段构造提供了显式的可观测实现。数值研究考察了风险估计和有限样本性能,一项模拟的注意力缺陷/多动障碍试验说明了比较反馈与治疗推荐之间的关系。

英文摘要

Model choice in backward Q-learning is recursive because a later-stage choice changes the response supplied to an earlier regression and can alter its model-comparison statistic. Separate stagewise criteria do not directly assess the target-stage prediction risk of a completed Q-learning fit. We address this mismatch by treating the entire backward-fitting rule as the unit of comparison and propose feedback-aware soft tuning for sequential multiple assignment randomised trials. Each backward fit uses its own generated responses and is assessed at a common prediction target. The risk criterion retains downstream effects on upstream comparisons, while a separate correction accounts for estimating the final exponential weights from the same observations. For a fixed finite library of smooth recursive maps, we establish an exact risk identity under a Gaussian shift model and an oracle inequality with an explicit adaptation remainder. Under coordinate-representation and moment conditions, these guarantees transfer to prediction risk at any prespecified stage, with the number of stages fixed as sample size increases. A two-stage construction supplies an explicit observable implementation. Numerical studies examine risk estimation and finite-sample performance, and a simulated attention-deficit/hyperactivity-disorder trial illustrates the relation between comparison feedback and treatment recommendations.

发表机构

  • Chuo University(中央大学)

机构由 AI 辅助整理,请以论文原文为准。

↑