DoublesEval:通过专业双打羽毛球诊断视觉语言模型的多智能体战术推理能力
DoublesEval: Diagnosing Multi-Agent Tactical Reasoning in Vision-Language Models via Professional Doubles Badminton
- The Hong Kong University of Science and Technology(香港科技大学)
- University of Macau(澳门大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出DoublesEval框架,以专业双打羽毛球为测试平台评估VLMs的多智能体战术推理能力,发现现有模型存在明显瓶颈,所提TacticCheck可提升模型表现但仍有差距,强调需改进VLMs的评估范式。
AI中文摘要:
视觉语言模型(VLMs)在描述可见场景内容方面表现出色,但在动态多智能体交互推理方面存在困难,这类交互中动作语义依赖于协调角色与时空依赖关系。我们将该能力定义为多智能体战术推理,并推出DoublesEval——一个利用专业双打羽毛球作为结构可处理测试平台的诊断评估框架。DoublesEval采用基于关键时刻的协议,将回合分解为战术显著时刻,从四个可解释维度探测模型:原子识别、段内复合理解、段间因果推理及高级战术抽象。该设计可定位推理失败的具体环节,而非仅测量答案正确性。为解决观测到的失败模式,我们提出TacticCheck——一种轻量的约束导向测试时一致性检查器,利用模型自身的低级战术预测对候选答案进行重排序,推理时无需参数更新或真值标签。通过零样本协议在60个精心整理的回合(产生约9600个结构化实例)上评估四个代表性开源VLMs,我们发现模型在所有诊断水平上仍表现薄弱,在空间状态、交互绑定及终端证据方面瓶颈尤为明显。TacticCheck在所有被评估模型上均实现了一致提升,但与稳健的战术推理仍存在显著差距。这些结果凸显了下一代VLMs需要结构化、感知交互的评估范式。源代码可在我们的GitHub仓库获取。
英文摘要:
Visual Language Models (VLMs) excel at describing visible scene content but struggle to reason about dynamic multi-agent interactions, where action semantics depend on coordinated roles and spatial-temporal dependencies. We formalize this capability as \textbf{multi-agent tactical reasoning} and introduce \textbf{DoublesEval}, a diagnostic evaluation framework that leverages professional doubles badminton as a structurally tractable testbed. DoublesEval employs a key-moment-based protocol that decomposes rallies into tactically salient instants and probes models across four interpretable dimensions: atomic recognition, intra-segment composite understanding, cross-segment causal reasoning, and high-level tactical abstraction. This design isolates \emph{where} reasoning fails, rather than merely measuring answer correctness. To address observed failure modes, we propose \textbf{TacticCheck}, a lightweight constraint-guided test-time consistency checker that reranks candidate answers using the model's own lower-level tactical predictions, requiring no parameter updates or ground-truth labels at inference time. Evaluating four representative open-source VLMs on 60 curated rallies (yielding $\sim$9.6K structured instances) via a zero-shot protocol, we find that models remain weak across all diagnostic levels, with especially clear bottlenecks in spatial state, interaction binding, and terminal evidence. TacticCheck delivers consistent gains across all evaluated models, while still leaving a substantial gap to robust tactical reasoning. These results highlight the need for structured, interaction-aware evaluation paradigms for next-generation VLMs. The source code is available in \href{https://github.com/Chengjt1999/DoublesEval}{\textcolor{blue}{our GitHub repository}}.