未书写基准:多模态机器学习在抽象感知推理领域的新挑战
The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning
浏览论文内容
中文总结 AI 辅助
本文提出The Unwritten Benchmark基准,针对多模态模型在抽象感知推理中的不足,开展声-运动学文字推理任务评估,发现模型性能远逊于人类且存在模态融合悖论。
中文摘要 AI 辅助
当前多模态模型在识别静态视觉和听觉内容方面已展现出卓越能力,但它们在抽象感知推理(即从动态生成过程中推断未观测信息)方面的能力仍是一个关键且未被充分探索的前沿领域。本文中,我们提出了一个名为The Unwritten Benchmark的新挑战,旨在探究这种抽象感知与认知能力。我们将核心任务定义为“声-运动学文字推理”:模型必须仅通过笔刮擦的音频和手部动作的视频,在没有任何可见墨水痕迹的情况下,识别3种不同书写风格的文字。我们的评估结果揭示了人类与机器性能之间存在巨大差距:人类参与者实现了较高的有序字母准确率(超过80%),而包括GPT-4o和Gemini 2.5-Pro在内的领先多模态机器学习模型表现极差,未能超过10%。此外,我们在模型中发现了一种矛盾的融合效应:提供两种模态信息通常会降低性能而非提升性能。这一发现表明,在该认知任务中,模型综合互补感知线索的能力存在根本性缺陷。这些发现凸显了模型在跨模态因果推理以及对这类认知和直觉感知推理所必需的微运动学理解方面存在的重大局限。
英文摘要
Current multimodal models have demonstrated remarkable proficiency in recognizing static visual and auditory content. However, their capacity for abstract perceptual reasoning, inferring unseen information from dynamic, generative processes, remains a critical and underexplored frontier. In this paper, we introduce The Unwritten Benchmark, a new challenge designed to probe this abstract perceptual and cognitive ability. We define the core task as acousto-kinematic word inference: models must decipher words, across 3 different writing styles, being written solely from the audio of pen scratches and the video of hand movements, without any visible ink trace. Our evaluation results reveal a profound gap between human and machine performance: while human participants achieve high ordered letter accuracy (over 80%), leading Multimodal Machine Learning Models, including GPT-4o and Gemini 2.5-Pro, struggle significantly, failing to surpass 10%. Furthermore, we identify a paradoxical fusion effect in the models, where providing both modalities often degrades performance rather than improving it. This finding indicates a fundamental breakdown in their ability to synthesize complementary perceptual cues for this cognitive task. These findings highlight significant limitations in both cross-modal causal reasoning and the understanding of the micro-kinematics essential for such cognitive and intuitive perceptual reasoning.
发表机构
- Arizona State University(亚利桑那州立大学)
机构由 AI 辅助整理,请以论文原文为准。