arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CypherTurn:面向对话式文本到Cypher评估的多轮基准与自主性分歧

CypherTurn: A Multi-Turn Benchmark for Conversational Text-to-Cypher Evaluation and the Autonomy Divergence

Yuzhe Zhang, Weijie Zhu, Haolin Yang, Ziyun Zhang, Xianwei Xue, Mengke Chen, Qiutong Pan, Huaqian Cai

arXiv 2609.36987首次发表:更新:

发表机构

Peking University; National Key Lab of Data Space Technology and System; Baidu Inc.(北京大学; 数据空间技术与系统全国重点实验室; 百度公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出CypherTurn,首个对话式文本到Cypher多轮基准,含721个会话,评估15个模型,发现最佳执行准确率仅64.7%,并揭示自主性分歧现象。

AI 中文摘要

图数据库日益通过自然语言进行查询,然而现有的每个基准都只评估孤立的单轮查询,而非分析师实际工作所依赖的多轮会话。我们引入了CypherTurn,这是首个用于对话式文本到Cypher评估的基准,包含721个会话和5,927轮对话,覆盖7个知识图谱和13种对话现象。我们在引导式预言机协议和完全自主的智能体协议下评估了15个模型,得出四项发现。首先,最佳模型仅达到64.7%的执行准确率,会话级正确率低于5%。其次,尽管总体排名相关性很强,前沿模型在自主操作下表现出排行榜顶部的显著重排,我们称之为自主性分歧,这揭示了错误管理作为部分独立于原始生成能力的能力。第三,将行动预算从x3扩展到x10未能缩小自主性差距,因为最强的前沿模型无论可用预算如何,每轮自我限制为大约两个行动。第四,单轮Cypher微调降低了多轮指令遵循能力,而架构适配的专门化优于多个前沿模型。这些结果确立了CypherTurn作为对话式图数据库推理的开放挑战。代码和数据可在该https URL获取。

英文摘要

Graph databases are increasingly queried through natural language, yet every existing benchmark evaluates isolated single-turn queries rather than the multi-turn sessions through which analysts actually work. We introduce CypherTurn, the first benchmark for conversational Text-to-Cypher evaluation, comprising 721 sessions and 5,927 turns across 7 knowledge graphs and 13 conversational phenomena. We evaluate 15 models under a guided oracle protocol and a fully autonomous agentic protocol, yielding four findings. First, the best model reaches only 64.7% execution accuracy, and session-level correctness remains below 5%. Second, despite strong overall rank correlation, frontier models exhibit a consequential reordering of the top of the leaderboard under autonomous operation, a phenomenon we term the Autonomy Divergence, which reveals error-management as a partially independent capability from raw generation skill. Third, scaling action budgets from x3 to x10 fails to close the autonomy gap, as the strongest frontier models self-limit to approximately two actions per turn regardless of available budget. Fourth, single-turn Cypher fine-tuning degrades multi-turn instruction following, while architecture-appropriate specialization outperforms several frontier models. These results establish CypherTurn as an open challenge for conversational graph database reasoning. Code and data are available at https://github.com/BarryQ/CypherTurn.

CommentsAccepted as an oral paper at EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑