MATEO:一种多模态基准,用于LVLMs中的时间推理和规划
MATEO: A Multimodal Benchmark for Temporal Reasoning and Planning in LVLMs
- Signals and Interactive Systems Lab, University of Trento, Italy(特伦托大学信号与交互系统实验室)
- University of Trento(特伦托大学)
- Amazon(亚马逊)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
MATEO是一个多模态基准,用于评估和提升大型视觉语言模型在时间推理和规划方面的能力。
中文摘要 AI 辅助
人工智能代理需要规划以实现涉及协调感知、子目标分解和执行的复杂目标。这些计划由按时间执行顺序(TEO,一种有向无环图,确保每一步在满足其前提条件后才执行)结构化的有序步骤组成。现有研究在基础模型对时间执行的理解上仅限于自动推导的注释、将TEO近似为线性链的近似值或文本-only输入。为解决这一差距,我们引入MATEO(多模态时间执行顺序),一个旨在评估和提升大型视觉语言模型(LVLMs)的时间推理能力的基准,这些能力对于现实世界规划至关重要。我们获得了一个高质量的专业多模态食谱语料库,该语料库通过标准化的编辑流程编纂,将指令分解为离散的步骤,每个步骤都配以相应的图像。我们通过设计并使用可扩展的众包流程收集TEO注释作为图。使用MATEO,我们评估了六种最先进的LVLMs,在模型规模、语言上下文、多模态输入结构和微调策略上均有所不同。
英文摘要
AI agents need to plan to achieve complex goals that involve orchestrating perception, sub-goal decomposition, and execution. These plans consist of ordered steps structured according to a Temporal Execution Order (TEO, a directed acyclic graph that ensures each step executes only after its preconditions are satisfied. Existing research on foundational models' understanding of temporal execution is limited to automatically derived annotations, approximations of the TEO as a linear chain, or text-only inputs. To address this gap, we introduce MATEO (MultimodAl Temporal Execution Order), a benchmark designed to assess and improve the temporal reasoning abilities of Large Vision Language Models (LVLMs) required for real-world planning. We acquire a high-quality professional multimodal recipe corpus, authored through a standardized editorial process that decomposes instructions into discrete steps, each paired with corresponding images. We collect TEO annotations as graphs by designing and using a scalable crowdsourcing pipeline. Using MATEO, we evaluate six state-of-the-art LVLMs across model scales, varying language context, multimodal input structure, and fine-tuning strategies.