GraphDecide:在图形任务上基准测试System One模型
GraphDecide: Benchmarking System One Models on Graph Tasks
浏览论文内容
中文总结 AI 辅助
GraphDecide是一个模型无关的基准测试,通过结构任务概况、图形-文本输入对比和启发式提案控制来诊断System One模型在图形决策上的性能,发现Jev等模型在准确邻接识别之外存在结构正确性、输入一致性和解决方案质量方面的局限。
中文摘要 AI 辅助
大型语言模型(LLMs)在图形理解和决策方面的应用日益增多,而System One模型(如Jev)直接从提供的选项中进行选择。然而,System One模型在图形相关任务上的能力仍不明确。我们引入了GraphDecide,一个与模型无关的基准测试,它结合了结构任务概况、匹配的图形-文本输入对比和启发式提案控制,以诊断图形决策性能。我们评估了Jev及相关的基于选择的模型,并与语言模型基线进行比较,涵盖了十四种模型-接口配置。Jev的结果展示了该基准测试的核心区别:准确的邻接识别并不能保证更广泛的结构正确性,联合图形-文本输入并不能持续提高预测性能,可行的构建并不能确立高解决方案质量。其任务契约、候选接口和评分规则支持在原生选择器和语言模型适配器之间进行比较。代码和汇总结果可在该https URL获取。
英文摘要
Large language models (LLMs) are increasingly explored for graph understanding and decision-making, while System One models such as Jev select directly from supplied options. However, the capabilities of System One models on graph-related tasks remain unclear. We introduce GraphDecide, a model-independent benchmark that combines structural task profiles, matched graph-text input contrasts and heuristic-proposal controls to diagnose graph decision performance. We evaluate Jev and related choice-based models alongside language-model baselines, covering fourteen model-interface configurations. Jev's results illustrate the benchmark's central distinctions: accurate adjacency recognition does not guarantee broader structural correctness, joint graph-text input does not consistently improve prediction, and feasible construction does not establish high solution quality. Its task contracts, candidate interfaces and scoring rules support comparison across native selectors and language-model adapters. Code and aggregate results are available at https://github.com/VictorYXL/JevGraphBench.
发表机构
- Microsoft(微软)
- Beijing University of Technology(北京工业大学)
机构由 AI 辅助整理,请以论文原文为准。