打破反馈盲目性:用于序列决策的效用增强Transformer
Breaking Feedback-Blindness: Utility-Augmented Transformer for Sequential Decision Making
浏览论文内容
中文总结 AI 辅助
研究非平稳和部分可观测环境下序列决策问题,提出效用增强Transformer(UAT),通过紧凑效用状态调制注意力检索,解决现有模型反馈盲目性问题,在四个非平稳基准测试中性能优于其他基线。
中文摘要 AI 辅助
在非平稳和部分可观测环境中的序列决策需要快速适应潜在的状态变化。然而,现有的Transformer决策模型在检索机制上面临结构瓶颈:即使奖励用于训练或作为输入令牌暴露,注意力检索仍主要由观察得出的相似性驱动。我们将此限制形式化为反馈盲目检索,并证明在反馈信息任务中,具有不同行动-奖励结果的观察等效历史不能通过任何仅观察的注意力来区分,导致次优选择。为解决此不匹配问题,我们提出了效用增强Transformer(UAT),一种新的反馈条件检索注意力架构,其中紧凑的效用状态调制查询、键和值投影,允许行动-奖励历史在正向传播期间直接改变上下文检索。UAT还具有精确的零门退化特性,当反馈无信息时恢复为普通Transformer。在有限视界紧凑性和Lipschitz假设下,我们证明UAT严格扩大了仅观察的Transformer类,并能均匀逼近依赖反馈的决策映射。在四个非平稳基准测试中,UAT始终优于仅观察、测试时适应和输入级反馈基线,在需要更强适应的噪声更大的情况下有特别大的提升。
英文摘要
Sequential decision making in non-stationary and partially observable environments requires rapid adaptation to latent regime changes. However, existing Transformer decision models face a structural bottleneck in the retrieval mechanism: even when reward is used for training or exposed as an input token, attention retrieval remains primarily driven by observation-derived similarity. We formalize this limitation as feedback-blind retrieval, and formally show that, on feedback-informative tasks, observation-equivalent histories with different action-reward outcomes cannot be distinguished by any observation-only attention, resulting in suboptimal choice. To address this mismatch, we propose the Utility-Augmented Transformer (UAT), a new feedback-conditioned retrieval attention architecture in which a compact utility state modulates the query, key, and value projections, allowing action-reward history to directly alter context retrieval during the forward pass. UAT also enjoys an exact zero-gate degradation property that recovers the Vanilla Transformer when feedback is uninformative. Under finite-horizon compactness and Lipschitz assumptions, we prove that UAT strictly enlarges the observation-only Transformer class and can uniformly approximate feedback-dependent decision maps. Across four non-stationary benchmarks: synthetic navigation with hidden goal shifts, non-stationary sepsis treatment, cross-market portfolio allocation, and delayed-feedback recommendation, UAT consistently improves performance over observation-only, test-time adaptation, and input-level feedback baselines, with particularly large gains in noisier regimes that require stronger adaptation.
发表机构
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
- Shenzhen Research Institute of Big Data(深圳大数据研究院)
- University of International Business and Economics(对外经济贸易大学)
机构由 AI 辅助整理,请以论文原文为准。