发表机构
Shenzhen University(深圳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出固定预算修订协议,用确定性验证器提供精确反馈,测试19个LLM,发现模型间成功率差异大,精确反馈虽暴露错误但无法保证闭环修订可靠。
AI 中文摘要
闭环修订越来越多地用于大型语言模型(LLM)应用中,但失败可能源于反馈不完整或对正确反馈的无效响应。我们引入了一种固定预算的修订协议,使用确定性验证器,报告精确长度、词汇和组合约束下的所有剩余违规。固定反馈的正确性和完整性,以隔离模型侧的修订行为。在19个开源和闭源模型中,控制器级别的平均最终联合成功率从17.4%到99.8%不等,在相同的初始草稿下,跨模型的显著差距持续存在。受控实验揭示了模型对精确反馈的可复现的特定响应。后训练和规模调整重塑了这些响应,但并未一致地使其更接近精确修正。在所有约束族中,失败的轨迹常常重复先前的输出,且先前的重复与后续的可恢复性降低相关。匹配状态干预表明,在保持当前草稿和反馈不变的情况下,移除先前的对话会改变重复逃逸,但并未可靠地提高最终成功率;效果取决于模型、任务和触发状态的组成。精确反馈使修订错误可观察,但并未使闭环变得可靠。代码和复现说明:https://github.com/kevinjiang0121-cyber/exact-feedback-code。
英文摘要
Closed-loop revision is increasingly used in large language model (LLM) applications, but failures may reflect incomplete feedback or ineffective responses to correct feedback. We introduce a fixed-budget revision protocol with deterministic verifiers that report all remaining violations across exact-length, lexical, and compositional constraints. Fixing feedback correctness and completeness isolates model-side revision behavior. Across 19 open- and closed-source models, controller-level mean final joint success ranges from 17.4% to 99.8%, with substantial cross-model gaps persisting under identical initial drafts. Controlled experiments reveal reproducible model-specific responses to exact feedback. Post-training and scale reshape these responses without consistently bringing them closer to exact correction. Across all constraint families, failed trajectories often repeat earlier outputs, and prior recurrence is associated with lower subsequent recoverability. Matched-state interventions show that removing earlier dialogue while holding the current draft and feedback fixed changes recurrence escape without reliably improving final success; effects depend on the model, task, and trigger-state composition. Exact feedback makes revision errors observable, but does not make the closed loop reliable. Code and reproduction instructions: https://github.com/kevinjiang0121-cyber/exact-feedback-code.
Comments35 pages, 18 figures, 25 tables, including appendices