什么驱动了智能体文本到Cypher查询的恢复?LAST-CQ:一种LLM智能体自精炼框架
What Drives Recovery in Agentic Text-to-Cypher? LAST-CQ: An LLM Agent Self-Refinement Framework
浏览论文内容
中文总结 AI 辅助
LAST-CQ五智能体框架通过反事实实验证明,文本到Cypher生成中驱动恢复的关键是失败检测与重试路由,而非反馈复杂度或采样数量,且能恢复91.7%的失败查询。
中文摘要 AI 辅助
用于结构化查询生成的智能体流水线正在迅速扩展,但循环中哪一部分产生了增益尚不清楚。我们使用LAST-CQ——一个五智能体、无需训练、基于执行结果的文本到Cypher框架——作为仪表化测试平台,在2,471个实时数据库查询和六个覆盖三个供应商规模层级的骨干模型上运行了三个反事实实验。移除纠错环节相对于单次通过系统在聚合执行BLEU上损失3.1%,相对于无精炼反事实损失12.3%(对于最弱的骨干模型最高达80.7%)。将基于模式、由LLM合成的反馈替换为原始数据库错误字符串几乎不造成损失(朴素精确匹配20.9%对19.9%;端到端差异小于0.2%;通过两次单侧检验在±0.075的集合F1范围内等价)。将相同的调用预算用于并行采样会使质量下降10-11%。真正有效的是检测失败并将其路由到重试,而非反馈的复杂性或样本数量。LAST-CQ本身恢复了单次通过生成下失败的查询中的91.7%,而首次即成功的查询仍然恰好花费一次LLM调用。我们还表明,序列化结果上的n-gram重叠不是一个双向界限:它在65.9%的结果上相对于集合等价性过度评分,而在相对于人工评判语义时则评分不足。最后,我们将LLM评判器与盲人人工标签进行校准,发现其乐观偏差为9个百分点。
英文摘要
Agentic pipelines for structured-query generation are rapidly expanding, but it is unclear which part of the loop produces the gain. We use LAST-CQ -- a five-agent, training-free, execution-grounded Text-to-Cypher framework -- as an instrumented testbed, running three counterfactuals over 2,471 live-database queries and six backbones spanning three vendor scale tiers. Removing correction is worth between 3.1% aggregate execution-BLEU against the single-pass system and 12.3% against a no-refinement counterfactual (up to 80.7% for the weakest backbone). Replacing schema-grounded, LLM-synthesised feedback with raw database error strings costs almost nothing (20.9% vs. 19.9% naive exact match; <0.2% end-to-end; equivalent within $\pm 0.075$ set-F1 by two one-sided tests). Spending the same call budget on parallel sampling degrades quality by 10-11%. What works is detecting failure and routing it to a retry, not the feedback sophistication or number of samples. LAST-CQ itself recovers 91.7% of queries that fail under single-pass generation, while a query that succeeds first time still costs exactly one LLM call. We also show that n-gram overlap on serialised results is not a bound in either direction: it over-scores against set equivalence on 65.9% of results while under-scoring against judged semantics. Finally, we calibrate our LLM judge against blind human labels and find it optimistic by 9 points.
发表机构
- Athens University of Economics and Business(雅典经济与商业大学)
- Orfium
机构由 AI 辅助整理,请以论文原文为准。