arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越正确性:解决智能体文本到SQL中的欠明确性问题

Beyond Correctness: Resolving Underspecification in Agentic Text-to-SQL

Wen-Zhi Li, Yue Gong, Konstantinos Kanellis, Balakrishnan Murali Narayanaswamy

arXiv 2610.02739首次发表:更新:

发表机构

Cornell University; Amazon Web Services(康奈尔大学; 亚马逊云服务)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对智能体文本到SQL中因过早终止澄清导致的静默假设问题,提出PlanPool外部化题库机制,强制处理已识别问题,在三个基准上提升歧义覆盖率并减少静默失败,同时保持执行准确性。

AI 中文摘要

智能体文本到SQL系统可以在生成SQL之前与用户交互,以澄清欠明确的查询。然而,正确的执行结果并不一定意味着智能体已充分解决了潜在的欠明确性:智能体可能静默地做出未经验证的假设,而这些假设恰好与预期答案相符。我们表明,这种行为部分由过早终止澄清所驱动。尽管强制智能体提出更多问题能提高执行准确性,但歧义集中在较早的交互中,使得暴力提问效率低下。更重要的是,即使明确提示智能体规划其澄清过程,它也经常放弃已经识别为相关的问题。为解决这一失败模式,我们引入了PlanPool,它将澄清计划外部化为一个可变的题库。每个计划中的问题必须在提交前明确提出或删除,而新发现的歧义可以在交互过程中添加。在源自BIRD-Interact和Spider的三个基准上,PlanPool持续提高了歧义覆盖率,并减少了静默失败,相比于无约束和基于提示的替代方案,同时保持了竞争性的执行准确性。我们的结果强调了智能体推理中的一个重要区别:识别缺失信息是不够的,智能体还必须在提交答案之前可靠地维护并解决这些信息。

英文摘要

Agentic Text-to-SQL systems can interact with users to clarify underspecified queries before generating SQL. However, a correct execution result does not necessarily imply that the agent has adequately resolved the underlying underspecification: the agent may silently make unverified assumptions that happen to match the intended answer. We show that this behavior is driven in part by premature clarification termination. Although forcing an agent to ask more questions improves execution accuracy, ambiguities are concentrated in earlier interactions, making brute-force questioning inefficient. More importantly, even when explicitly prompted to plan its clarification process, the agent frequently abandons questions that it has already identified as relevant. To address this failure mode, we introduce PlanPool, which externalizes the clarification plan as a mutable question pool. Every planned question must be explicitly asked or dropped before submission, while newly discovered ambiguities can be added during interaction. Across three benchmarks derived from BIRD-Interact and Spider, PlanPool consistently improves ambiguity coverage and reduces silent failures over unconstrained and prompt-based alternatives, while maintaining competitive execution accuracy. Our results highlight an important distinction in agentic reasoning: identifying missing information is not sufficient, and the agent must also reliably maintain and resolve it before committing to an answer.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑