发表机构
Xidian University; Shanghai Jiao Tong University(西安电子科技大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出文本到SQL的结晶问题,通过受控评估发现存储已验证修正查询可提升BIRD数据集的保留首次尝试准确率,确定数据库特定内容是关键要素,明确验证与检索覆盖的作用。
AI 中文摘要
测试时扩展可修正困难的文本到SQL查询,但额外计算通常在每个答案后被丢弃。系统越来越多地保留已验证的修正片段,但评估仍只报告一个端到端分数,无法区分对重复问题的重放与对未见过问题的帮助,也无法识别负责的记忆选择,我们将测量这种未来价值的问题称为结晶问题。我们的受控评估固定了单次求解器,每次仅改变一个记忆选择,分别测量重放、跨问题保留和保留的同数据库迁移。在BIRD数据集上,存储已验证的修正查询可将保留的首次尝试准确率提高4.34个百分点,该增益占同一问题按需修正提供的准确率空间的44.4%。受控干预确定数据库特定内容是主要操作要素,可靠的验证和更广泛的检索覆盖产生了可支持的增益,而更丰富的格式和精心设计的检索器则没有。开源代码、评估人工制品和复现说明可在该https URL获取。
英文摘要
Test-time scaling can correct difficult text-to-SQL queries, but the extra computation is normally discarded after each answer. Systems increasingly retain verified repair episodes, yet evaluations still report one end-to-end score. It cannot distinguish replay on recurring questions from help on unseen questions, or identify the responsible memory choice. We call measuring this future value the crystallization problem. Our controlled evaluation holds the single-shot solver fixed and varies one memory choice at a time. We separately measure replay, cross-question retention, and held-out same-database transfer. On BIRD, storing verified corrected queries improves held-out first-attempt accuracy by 4.34 percentage points. This gain captures 44.4% of the accuracy headroom provided by on-demand repair on the same questions. Controlled interventions identify database-specific content as the main operating ingredient. Reliable verification and broader retrieval coverage yield supported gains; richer formats and elaborate retrievers do not. Open-source code, evaluation artifacts, and reproduction instructions are available at https://github.com/ai-jiaqian/text-to-sql-memory-crystallization.
Comments18 pages, 6 figures. Open-source code, evaluation artifacts, and reproduction instructions: https://github.com/ai-jiaqian/text-to-sql-memory-crystallization