评估用于视频游戏QA中自主代理驱动的几何裁剪检测的视觉语言模型
Evaluating VLMs for Autonomous Agent-Driven Geometry Clipping Detection in Video Game QA
查看机构详情
- Sony Interactive Entertainment(索尼互动娱乐)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
研究在代理驱动游戏QA中用VLMs检测几何裁剪异常,通过自定义代理收集视觉观察,自动注释提供标签,在零样本提示下对六个VLMs基准测试,分析对提示变体的敏感性,结果显示VLMs适合作多阶段QA管道的高召回候选过滤器。
中文摘要 AI 辅助
在这项工作中,我们研究视觉语言模型(VLMs)在以几何裁剪为重点的代理驱动游戏质量保证(QA)管道中的异常检测应用。在评估中,一个自定义探索代理在游戏关卡中导航以收集视觉观察结果,同时自动注释管道提供帧级裁剪标签。此设置使我们能够在无需人工注释的情况下,在受控的异常检测任务中评估近期的VLMs。我们在零样本提示设置下对六个近期的VLMs(Gemini、GPT、Qwen、Gemma、Llama和Ministral)进行基准测试,并分析它们对四种提示变体的敏感性。结果表明,VLMs虽能捕捉与几何裁剪相关的视觉线索,但在视觉模糊的帧(如近接触几何和部分遮挡)上会产生大量误报。Gemini-3.1-Flash总体准确率最高,对提示变化最稳健,而开源模型的精确率-召回率会因提示设计有较大波动。这些发现表明,当前的VLMs最适合作为多阶段QA管道中的高召回候选过滤器,而非独立的错误检测器。
英文摘要
In this work, we study the use of Vision-Language Models (VLMs) for anomaly detection in an agent-driven game Quality Assurance (QA) pipeline focusing on geometry clipping. In this evaluation, a custom exploration agent navigates a game level to collect visual observations, while the automatic annotation pipeline provides frame-level clipping labels. This setup allows us to evaluate recent VLMs on a controlled anomaly detection task without manual annotation. We benchmark six recent VLMs (Gemini, GPT, Qwen, Gemma, Llama, and Ministral) under a zero-shot prompting setting and analyse their sensitivity to four prompt variants. Our results show that while the VLMs can capture visual cues associated with geometry clipping, they all produce substantial false positives on visually ambiguous frames such as near-contact geometry and partial occlusions. Gemini-3.1-Flash achieves the best overall accuracy and is the most robust to prompt variation, while open-source models exhibit large precision--recall swings depending on the prompt design. These findings suggest that current VLMs are best suited as high-recall candidate filters within multi-stage QA pipelines rather than as standalone bug detectors.