发表机构
Google DeepMind; Toyota Technological Institute at Chicago; University of Bristol; University of Oxford(谷歌DeepMind; 芝加哥丰田理工学院; 布里斯托尔大学; 牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文总结2026年ECCV研讨会举办的第四届感知测试挑战赛,新增城市级视听推理赛道,发现复杂空间多模态推理需昂贵智能体流水线,独立多模态模型仍难胜任。
AI 中文摘要
延续感知测试挑战赛系列,我们在瑞典马尔默举办的2026年欧洲计算机视觉会议(ECCV)上,以研讨会形式组织了第四届赛事。本届赛事聚焦空间智能,设置了四个不同赛道:来自原感知测试基准的统一多项选择题视频问答(videoQA)和 grounded 视频问答,以及两个基于城市级步行游览视频的新赛道——KilometerAudio 和 KilometerVision。本报告描述了城市级赛道使用的新基准,并总结了所有赛道的获奖方案,包括一个参与所有赛道且表现令人满意的通用模型。新增城市级赛道的获奖方案表明,复杂的空间与多模态推理可通过昂贵的智能体(agentic)流水线解决,但对于独立使用的多模态模型而言,该任务仍具难度。
英文摘要
Continuing the Perception Test challenge series, we organised the fourth edition as a workshop at the European Conference on Computer Vision (ECCV) 2026 in Malmö, Sweden. This edition focused on spatial intelligence and featured four different tracks: unified multiple-choice videoQA and grounded videoQA from the original Perception Test benchmark, alongside two new tracks based on city-scale walking-tour videos (KilometerAudio and KilometerVision). In this report, we describe the new benchmarks used for the city-scale tracks and summarise the winning solutions across all tracks, including a generalist model that competed across all tracks with satisfactory performance. The winning solutions in the newly added city-scale tracks demonstrated that complex spatial and multimodal reasoning can be solved by expensive agentic pipelines, but remains difficult for multimodal models used standalone.