视觉感知不等于决策:多模态大语言模型能成为有效的首席执行官吗?
Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs?
浏览论文内容
中文总结 AI 辅助
该研究构建C-SUITEBENCH多模态基准,评估9个前沿模型作为CEO的决策能力,发现多模态输入可提升证据推理但会损害受限资源分配,揭示多模态智能体的视觉感知与受限行动是可分离瓶颈。
中文摘要 AI 辅助
大语言模型正越来越多地被用作自主决策智能体。然而,在高管商业决策场景中,现有基准仅限于纯文本设置,这使得模型能否感知视觉商业证据并有效整合以提升决策质量尚不明确。我们推出C-SUITEBENCH,这是一个受控多模态基准,包含50个场景下的5项决策任务,涵盖纯文本和多模态配对条件。我们让9个前沿模型扮演首席执行官角色,评估其决策能力。多模态输入持续提升以证据为中心的推理能力,在风险预测和面向董事会的论证中,提升幅度最大且最可靠。但我们发现了多模态整合悖论:添加视觉商业信息会降低所有9个模型的受限资源分配表现,尽管视觉 grounding( grounding 指视觉 grounding,即视觉信息与文本的关联匹配)本身有所改善。消融实验显示,这种失效源于信号拥挤:每个视觉通道单独都有帮助,但它们的组合会在解码过程中破坏约束满足。这些发现表明,视觉感知和受限行动是多模态智能体的可分离瓶颈,无差别地添加视觉信息会损害高风险决策,为未来的高管AI系统提出了选择性 grounding 策略的方向。
英文摘要
Large language models are increasingly applied as autonomous decision-making agents. However, in executive business decisions, existing benchmarks are limited to textonly settings. This makes it unclear whether models can perceive visual business evidence and effectively integrate it to improve decision quality. We introduce C-SUITEBENCH, a controlled multimodal benchmark that includes five decision tasks under paired text-only and multimodal conditions across 50 scenarios. We place nine frontier models in the role of a chief executive officer and evaluate their decision-making ability. Multimodal inputs consistently improve evidence-centric reasoning, with the largest and most reliable gains appearing in risk forecasting and board-facing justification. However, we uncover a multimodal integration paradox: adding visual business information degrades constrained resource allocation for all nine models, even as visual grounding itself improves. Ablation experiments reveal that this failure emerges from signal crowding, although each visual channel helps individually, their combination disrupts constraint satisfaction during decoding. These findings demonstrate that visual perception and constrained action are separable bottlenecks in multimodal agents, and that indiscriminate visual augmentation can harm high-stakes decision making, motivating selective grounding strategies for future executive AI systems.