arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在蒸馏决策器的实时间隙中分离决策时间与决策质量:来自一个游戏和一个传送带模拟器的证据

Separating Decision Time from Decision Quality in the Real-Time Gap of Distilled Deciders: Evidence from a Game and a Conveyor Simulator

Chihoon Shin, Junyeong Lee, Kihyeok Jeong, Wonok Kwon

arXiv 2610.04810首次发表:更新:

发表机构

Myeongseongsimjae AX Institute; Korea University; Electronics and Telecommunications Research Institute (ETRI)(明成心斋AX研究所; 高丽大学; 韩国电子通信研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究将实时决策器与零延迟教师的差距分解为时间与质量成分,通过ViZDoom游戏和传送带模拟器实验,发现质量成分显著且可通过重训练改善,延迟匹配控制指导补救措施选择。

AI 中文摘要

实时智能体常常在世界暂停的情况下被评估,或者将决策时间四舍五入到整个时间片(时间片转换)。时间片转换预测,在异步游戏中,一个学习型决策器在相同游戏上的命中准确率下降14.8个百分点(pp),但几乎没有损失。我们转而将决策器与零延迟教师之间的差距分解为时间和质量两个组成部分。在ViZDoom中,我们通过模仿脚本教师训练一个小型决策器,按墙钟时间运行游戏,并添加一个控制组,其中教师等待对决策器服务器的调用后再做决策。在Windows主机上,三项预注册研究将时间组成部分定为6.5-8.7个百分点,质量组成部分定为6.3-8.6个百分点。在跳过0.08-0.09%时间片的Linux主机上,时间组成部分消失(-0.1个百分点;配对95%自助法区间[-0.4, 0.0]),而质量组成部分仍然存在(5.6个百分点[3.7, 7.5])。在决策器自身游戏产生的教师标记状态上重新训练(DAgger)在两个主机上的新游戏中均有所改进(2.6个百分点[0.4, 4.9]和3.1个百分点[1.1, 5.2]),这是一个预注册的部分成功。在实时游戏中,对每个臂施加一个延迟调度,在100个新游戏上留下5.8个百分点[4.0, 7.5]的质量组成部分,重新训练将决策器与其自身状态上的教师分歧从22.7%降至14.9%。在一个具有400毫秒截止时间的传送带模拟器中,超过截止时间的延迟消除了质量组成部分。一个延迟与决策器匹配的延迟匹配控制显示了应尝试哪种补救措施。

英文摘要

Real-time agents are often evaluated with the world paused, or with decision time rounded to whole ticks (tick conversion). Tick conversion predicts almost no loss for a learned decider whose hit accuracy falls by 14.8 percentage points (pp) on the same games in asynchronous play. We instead split the decider's gap to a zero-latency teacher into time and quality components. In ViZDoom, we train a small decider by imitating a scripted teacher, run the game on wall-clock time, and add a control in which the teacher waits for a call to the decider's server before deciding. On a Windows host, three preregistered studies put the time component at 6.5-8.7 pp and the quality component at 6.3-8.6 pp. On a Linux host that skips 0.08-0.09% of ticks, the time component disappeared (-0.1 pp; paired 95% bootstrap interval [-0.4, 0.0]) while the quality component remained (5.6 pp [3.7, 7.5]). Retraining on teacher-labelled states from the decider's own play (DAgger) improved it on new games on both hosts (2.6 pp [0.4, 4.9] and 3.1 pp [1.1, 5.2]), a preregistered partial success. Imposing one delay schedule on every arm in live play left a quality component of 5.8 pp [4.0, 7.5] on 100 new games, and retraining cut the decider's disagreement with the teacher on its own states from 22.7% to 14.9%. In a conveyor simulator with a 400 ms deadline, delay past it erased the quality component. A latency-matched control whose delays match the decider's shows which remedy to try.

Comments80 pages including supplementary material. Submitted to Knowledge-Based Systems

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑