AI 中文总结
提出EngIntervene基准,含3,229个问题,四级评估多模态工程状态理解与干预推理,发现强基础能力不等于强推理,开放模型落后封闭模型14.7个百分点。
AI 中文摘要
多模态工程基准主要评估静态理解,如识别组件、解读图表或回答技术问题。这留下了工程感知与完整设计生成之间的空白:模型能否利用已理解的系统状态来推理关系、约束以及设计变更的后果。我们引入了EngIntervene,一个针对此能力的基准。它包含七个工程领域的3,229个问题,并将评估组织为四个层级:状态基础(T1)、关系与机制推理(T2)、约束感知诊断(T3)和干预推理(T4),其中T4询问提出的修改是否在保持所需约束的同时实现其目标。这些任务实例化了跨越对象、关系、约束和设计目标的统一工程状态表示,T2--T4根据具有原子标准的结构化参考答案进行评分。在开放权重和封闭权重的多模态模型中,更强的基础能力并不能可靠地转化为更好的诊断或干预,最佳开放权重模型在T2--T4平均值上落后最佳封闭模型14.7个百分点。移除或打乱视觉证据持续降低性能,而针对基准的监督微调提高了T1但未提高T2--T4。T4进一步暴露了满足个别修订标准与产生完全有效干预之间的巨大差距。因此,工程推理不仅需要恢复当前状态,还需要可靠地利用它来推理约束和干预后后果。代码和基准工件可在https URL获取。
英文摘要
Multimodal engineering benchmarks largely evaluate static understanding, such as recognizing components, interpreting diagrams, or answering technical questions. This leaves a missing middle between engineering perception and full design generation: whether a model can use an understood system state to reason about relations, constraints, and the consequences of design changes. We introduce \textsc{EngIntervene}, a benchmark for this capability. It contains 3,229 questions across seven engineering domains and organizes evaluation into four levels: state grounding (T1), relational and mechanistic reasoning (T2), constraint-aware diagnosis (T3), and intervention reasoning (T4), which asks whether a proposed modification achieves its target while preserving required constraints. The tasks instantiate a unified engineering state representation spanning objects, relations, constraints, and design objectives, and T2--T4 are scored against structured reference answers with atomic criteria. Across open- and closed-weight multimodal models, stronger grounding does not reliably translate into better diagnosis or intervention, and the best open-weight model trails the best closed model by 14.7 percentage points on the T2--T4 average. Removing or shuffling visual evidence consistently degrades performance, while benchmark-specific supervised fine-tuning improves T1 but not T2--T4. T4 further exposes a large gap between satisfying individual revision criteria and producing a fully valid intervention. Engineering reasoning thus requires not only recovering the current state, but also reliably using it to reason about constraints and post-intervention consequences. Code and benchmark artifacts are available at https://github.com/changcv2021/EngIntervene