ExBind:一个用于视觉到可执行对应关系的受控诊断基准
ExBind: A Controlled Diagnostic Benchmark for Visual-to-Executable Correspondence
浏览论文内容
中文总结 AI 辅助
ExBind是用于视觉到可执行对应关系的受控诊断基准,含多类案例及特定模型的性能数据,可用于定位多模态编辑系统的指称错误根源。
中文摘要 AI 辅助
多模态编码与编辑系统必须将可见或语义指称对象映射到可编辑的精确可执行对象。错误的指称可能会选择有效但不正确的DOM节点、SVG元素、图端点、层次结构成员或表格单元格,而仅最终执行成功并不能揭示失败的根源。ExBind将这种视觉到可执行的对应层分离出来,作为语义定位与动作执行之间的受控诊断基准。它采样与表示无关的潜在绑定实例,并将其编译为SVG、DOM、画布、树、图和表格案例,这些案例具有到可执行指称的确定性映射。模型仅输出严格的指称;评估器将预测映射回潜在结构,并对结构约束进行评分,无需推理轨迹。该基准包含250个案例的通用套件、240个案例的不相交目标套件以及50个配对潜在组。Qwen2.5-VL-3B达到98.4%的候选有效性,但精确准确率为76.4%;Qwen3-VL-4B达到100.0%的有效性,精确准确率为98.8%。在目标表格套件中,Qwen2.5-VL-3B的所有残留错误均为有效的行正确、列错误的选择。候选顺序扰动会改变案例级结果,但保留该错误模式。ExBind专为受控诊断而设计,而非用于总体规模排名或端到端编辑评估。代码和基准记录可在指定的两个URL获取。
英文摘要
Multimodal coding and editing systems must map a visible or semantic referent to the exact executable object that can be edited. A wrong reference may select a valid but incorrect DOM node, SVG element, graph endpoint, hierarchy member, or table cell, while final execution success alone does not reveal the source of the failure. ExBind isolates this visual-to-executable correspondence layer as a controlled diagnostic benchmark between semantic localization and action execution. It samples representation-independent latent binding instances and compiles them into SVG, DOM, canvas, tree, graph, and table cases with deterministic mappings to executable references. Models output only a strict reference; the evaluator maps predictions back to latent structure and scores structural constraints without requiring reasoning traces. The release contains a 250-case broad suite, a disjoint 240-case targeted suite, and 50 paired latent groups. Qwen2.5-VL-3B achieves 98.4% candidate validity but 76.4% exact accuracy, while Qwen3-VL-4B achieves 100.0% validity and 98.8% exact accuracy. In the targeted table suite, all Qwen2.5-VL-3B residual errors are valid correct-row/wrong-column selections. Candidate-order perturbations change case-level outcomes while preserving this error pattern. ExBind is designed for controlled diagnosis rather than population-scale ranking or end-to-end editing evaluation. Code and benchmark records are available at https://github.com/Daerwang2020/Exbind and https://huggingface.co/datasets/Ziqianwwww/ExBind.
发表机构
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。