发表机构
Atlassian; Carnegie Mellon University(Atlassian; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对电子表格问答,提出VIGIL框架,通过验证器门控改进循环在持续学习中仅更新校准器与选择器,实现有界适配,前向准确率提升至82.4%,并验证门控泛化局限。
AI 中文摘要
企业智能体应当能够从延迟反馈中改进,而不允许每次修正都重写系统行为。我们研究面向语料级电子表格问答的持续框架学习。在FiCo(先查找后计算)这一静态检索与执行骨干的基础上,我们引入了VIGIL(验证器知情门控改进循环)。在单个问题内部,VIGIL对多种提示生成的结构化查询语言(SQL)候选进行验证与修复。跨回合中,延迟标签仅更新一个契约校准器以及查询/结果列选择器。基础模型、提示、检索器和已记录候选池在持续学习协议中保持固定。在提供金标准工作簿和预期类型的情况下,无门控和双门控全重放变体在79个文档上的前向准确率分别达到82.4%和82.0%,而静态基线为75.8%。双门控具有更大的回溯增益(2.7对2.3个百分点),其6.3个百分点的前向增益的95% t区间为5.6-6.9。在另一个更严格的划分中,排除16个文档和380个问题不参与拟合与在线提升,仅准确率门控将十个最终框架的平均保留准确率从80.3%提高到86.7%。仅校准器适配获得5.9个百分点的增益,接近合并的6.4个百分点增益。然而,26个门控批准的更新中有三个相对于其原有版本降低了保留准确率,因此重放缓冲区的非回归并不意味保留集上的非回归。MiMoTable和外部任务案例研究测试了任务内验证,并采用相同的“提升或保留”纪律。总体而言,结果支持有界、可审计的框架适配,同时揭示了有限重放门控在超越其提升缓冲区时无法泛化的场景。
英文摘要
Enterprise agents should improve from delayed feedback without allowing every correction to rewrite system behavior. We study continual harness learning for corpus-level spreadsheet question answering. Building on FiCo (Find-then-Compute), a static retrieval-and-execution backbone, we introduce VIGIL (Verifier-Informed Gated Improvement Loop). Within a question, VIGIL verifies and repairs diversely prompted Structured Query Language (SQL) candidates. Across episodes, delayed labels update only a contract calibrator and query/result-column selector. The base model, prompts, retriever, and recorded candidate pool remain fixed in the continual-learning protocol. With the gold workbook and expected type supplied, the ungated and dual-gated full-replay variants reach 82.4% and 82.0% forward accuracy over 79 documents, from a 75.8% static baseline. The dual gate has the larger retrospective gain (2.7 versus 2.3 points), and its 6.3-point forward gain has a 95% t-interval of 5.6-6.9. In a separate stricter split that excludes 16 documents and 380 questions from fitting and online promotion, the accuracy-only gate raises mean held-out accuracy across ten final harnesses from 80.3% to 86.7%. Calibrator-only adaptation gains 5.9 points, close to the combined 6.4-point gain. Yet three of 26 gate-approved updates reduce held-out accuracy relative to their incumbents, so replay-buffer non-regression does not imply held-out non-regression. MiMoTable and external-task case studies test within-task verification and reuse the same promote-or-retain discipline. Overall, the results support bounded, auditable harness adaptation while revealing where finite replay gates fail to generalize beyond their promotion buffers.