arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ERUnderstand:在结构化实体关系图上评估视觉语言模型

ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams

Ali Ansari, Yasmin Mohammadi, Farnoush Nili, Parsa Esmaeilkhani, Longin Jan Latecki, Eduard Dragut

arXiv 2607.24707首次发表:更新:

发表机构

Temple University(天普大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对实体关系图仅为图像限制AI辅助数据库工程的问题,引入ERUnderstand基准测试。通过对2960个图表及标准化表示评估视觉语言模型,发现常见元素恢复良好,但特定元素性能差,推理增强模型可提升性能,为评估多模态理解提供基准。

AI 中文摘要

实体关系图(ERD)是概念数据库设计的核心,但通常仅以渲染图像形式存在,而非机器可读模式,这限制了人工智能辅助数据库工程。我们引入了ERUnderstand,这是首个用于结构化理解ER图的大规模基准测试,包含从精心挑选的教育资源、真实世界模式以及跨越不同领域、符号、复杂度级别和扩展实体关系(EER)构造的合成示例中收集的2960个图表。每个图表都与标准化的机器可读表示配对,用于对模式元素进行细粒度评估。评估当前最先进视觉语言模型(VLM)时发现,虽然常见ERD元素能可靠恢复(F1>0.74),但在弱实体(F1低至0.28)、多值属性(0.14 F1)和N元关系(0.07 F1)上性能大幅下降。推理增强模型将整体性能提高了15%-25%,但仍对语言先验和图表复杂度增加敏感。ERUnderstand为评估概念数据库模式的多模态理解提供了标准化基准。基准测试、数据集、评估工具包和生成代码可在指定网址公开获取。

英文摘要

Entity-Relationship Diagrams (ERDs) are central to conceptual database design, yet they are typically available only as rendered images rather than machine-readable schemas, limiting AI-assisted database engineering. We introduce ERUnderstand, the first large-scale benchmark for structured understanding of ER diagrams, comprising 2,960 diagrams collected from curated educational sources, real-world schemas, and synthetically generated examples spanning diverse domains, notations, complexity levels, and Extended Entity-Relationship (EER) constructs. Each diagram is paired with a standardized machine-readable representation for fine-grained evaluation of schema elements. Evaluating state-of-the-art Vision-Language Models (VLMs), we find that while common ERD elements are recovered reliably (F1 > 0.74), performance drops sharply on weak entities (as low as 0.28 F1), multivalued attributes (0.14 F1), and N-ary relationships (0.07 F1). Reasoning-augmented models improve overall performance by 15-25% but remain sensitive to linguistic priors and increasing diagram complexity. ERUnderstand provides a standardized benchmark for evaluating multimodal understanding of conceptual database schemas. The benchmark, dataset, evaluation toolkit, and generation code are publicly available at https://github.com/salinaria/ERUnderstand.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑