arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24815cs.IR

将回报条件作为控制旋钮进行审计:用于决策转换器推荐的离线诊断方法

Auditing Return Conditioning as a Control Knob: An Offline Diagnostic for Decision Transformer Recommendation

Jingyu Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对Decision Transformer推荐系统,提出将回报条件作为控制旋钮的离线诊断方法,通过在MovieLens 25M与MAL数据集上开展实验,验证了RTG局部性等四项检查对回报控制的审计作用。

中文摘要 AI 辅助

离线逐次回报(RTG)扫描可测试基于回报条件的推荐系统是否可控,但该干预措施很少被审计。重写每个历史RTG token会产生越来越多的合成上下文,而仅重写当前token则更具局部性。我们在固定窗口的离线设置中测试这种差异,在MovieLens 25M和MyAnimeList 2020(MAL)数据集上,评估采用RTG局部性阶梯、无RTG控制、记录匹配与得分奖励检查以及轨迹内打乱RTG消融的Decision Transformer。在MovieLens上,仅应用于真实上下文位置、覆盖完整上下文的K=20干预,使犯罪类型预测的占比从验证集第5百分位到第95百分位提升了+23.61±2.96个百分点,而仅改变当前位置则提升+1.77±1.17个百分点;打乱RTG模型大幅消除了该响应(K=20时为+2.08±1.20个百分点)。在MAL上,相同协议未产生戏剧类型响应:K=20使戏剧类型变化-0.03±0.07个百分点,K=1时为-0.01±0.01个百分点。真实RTG、无RTG与打乱RTG的类型预测准确率数值相近,K=1时记录匹配率与匹配评分变化很小。由于数据集和重点类型选择是探索性的,这些幅度仅具描述性;局部性、打乱RTG及MAL上空结果的跨诊断模式未确立回报控制,我们提出四项检查:干预局部性、无RTG基线、奖励检查和RTG内容消融。

英文摘要

Offline return-to-go (RTG) sweeps can test whether a recommender conditioned on return is controllable, but the intervention is rarely audited. Rewriting every historical RTG token creates an increasingly synthetic context, while rewriting only the current token is more local. We test this distinction in an offline setting with a fixed window. On MovieLens 25M and MyAnimeList 2020 (MAL), we evaluate a Decision Transformer using an RTG locality ladder, a control without RTG, a logged match and score reward check, and a within-trajectory shuffled RTG ablation. On MovieLens, a $K=20$ intervention that covers the full context, applied only to real context positions, shifts the share of Crime predictions by $+23.61 \pm 2.96$ percentage points from the validation 5th to 95th percentile, whereas changing only the current slot shifts it by $+1.77 \pm 1.17$ points. The shuffled RTG model largely removes this response ($+2.08 \pm 1.20$ points at $K=20$). On MAL, the same protocol does not produce a Drama response: $K=20$ changes Drama by $-0.03 \pm 0.07$ points, and $K=1$ by $-0.01 \pm 0.01$. Genre prediction accuracy is numerically close across real RTG, no RTG, and shuffled RTG, and at $K=1$ logged match rates and matched ratings change little. Because dataset and focus-genre selection were exploratory, these magnitudes are descriptive; the cross-diagnostic pattern across locality, shuffled RTG, and the null result on MAL does not establish reward control. We propose four checks: intervention locality, a no-RTG baseline, a reward check, and RTG-content ablation.

补充信息

↑