arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21398cs.LGcs.AIcs.CYcs.HCcs.MA

星际争霸II中AlphaStar的AI控制之运行时动作干预

Runtime Action Interference for AI Control of AlphaStar in StarCraft II

发表机构新南威尔士大学 · 澳大利亚联邦科学与工业研究组织
查看机构详情
  • University of New South Wales(新南威尔士大学)
  • CSIRO(澳大利亚联邦科学与工业研究组织)

机构由 AI 辅助整理,请以论文原文为准。

Jaymari Chua, Chen Wang, Liming Zhu, Lina Yao

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出运行时动作干预(RAI)机制,将其应用于AlphaStar复现版并开展《星际争霸II》人类实验,发现向用户披露AI对手能力会显著影响人类对其公平性、毒性的感知,需将执行栈控制与能力披露分开评估。

中文摘要 AI 辅助

训练好的强化学习策略并不决定用户所遇到的全部行为:部署代码仍会对其提出的动作进行调度、允许、抑制或替换。本文提出了运行时动作干预(Runtime Action Interference, RAI),这是一种AI控制机制,可在保留策略参数的同时,在推理后调节动作节奏并过滤配置的动作模式。RAI仅在动作满足冷却条件且内容检测器未标记该动作时,才会释放该动作;否则,会发送空操作。检测器涵盖特定的有害行为,包括对工人单位的骚扰,而冷却机制则控制动作速率。我们在AlphaStar的复现版本中实现了RAI,并通过开源代码库提供该实现及可复现性材料。我们将RAI部署到《星际争霸II》的人类参与者研究中,该研究对同一高能力、速率受限动作的对手的两种呈现方式进行了比较:一种呈现方式中我们隐瞒其能力声明,另一种则予以披露。在1至5分的响应量表下,我们观察到,在隐瞒能力声明时,公平性、信任度和毒性的汇总均值分别为3.90、3.50和2.00;而在披露时,上述均值分别为2.62、4.31和2.85。披露与所有专业知识组感知到的公平性降低、毒性升高相对应,而信任度在新手和专家中有所提升,但在中级参与者中下降。因此,我们的人类评估表明,即使配置的控制保持不变,通过RAI控制的对手的感知会随向用户呈现的能力信息而显著变化。我们得出结论,人机评估必须将执行栈内的控制与能力披露分开,并将公平性、信任度和毒性作为人类体验的不同维度进行评估。

英文摘要

A trained reinforcement learning policy does not determine the complete behavior that users encounter: deployment code still schedules, admits, suppresses, or replaces its proposed actions. We contribute \emph{runtime action interference} (RAI), an AI control mechanism that preserves policy parameters while regulating action pacing and filtering configured action patterns after inference. RAI releases a proposed action only when its cooldown condition is satisfied and its content detector does not flag the action; otherwise, it dispatches a no-op. The detector covers specified toxic behaviors, including worker-unit harassment, while the cooldown controls action rate. We implement RAI in a replication of AlphaStar actor.py and make the implementation and reproducibility materials available through an open source code repository. We deployed RAI in a \textit{StarCraft~II} human participant study that compared two presentations of the same opponent with high capability and rate limited actions; we withheld its capability claim in one presentation and disclosed it in the other. On response scales from 1 to 5, we observed pooled fairness, trust, and toxicity means of 3.90, 3.50, and 2.00 under claim withholding, compared with 2.62, 4.31, and 2.85 under disclosure. Disclosure corresponded with lower perceived fairness and higher perceived toxicity across every expertise group, whereas trust increased among novices and experts but decreased among intermediate participants. Our human evaluation therefore shows that perceptions of an opponent controlled through RAI can vary substantially with the capability information presented to users, even when the configured control remains constant. We conclude that human-computer evaluations must separate control within the execution stack from capability disclosure and assess fairness, trust, and toxicity as distinct dimensions of human experience.

↑