在稳态下测量学习:BIRD-SQL 一级方程式案例研究
Measure Learning at Steady State: A BIRD-SQL Formula 1 Case Study
浏览论文内容
中文总结 AI 辅助
本研究在 BIRD-SQL 任务稳态下评估 ICL 学习,发现其后期探测效率下降而成本上升,短期收益掩盖了成本反转,表明无界 ICL 不适合作为学习机制。
中文摘要 AI 辅助
持续学习基准将学习评分视为相对于重置基线的短期收益,并发现在其测试的记忆中,朴素的全上下文 ICL 最强。我们将 ICL 视为一个学习系统,并在更长的共享世界时间表上对其进行评分。稳态学习是相对于预设后期窗口(174 个 BIRD-SQL 一级方程式问题中的最后 40 个)基线的差距。我们将分数分为探索效率(SQL 探测)、任务奖励(命中)和交付成本(API 美元和上下文大小)。在 gpt-5.6-luna 上,后期探测从 4.6-5.6 降至 0.95,而命中仅适度上升,ICL 上下文增长至约 95k 个 token,成本大约翻倍。短期收益低估了后期探测的节省,并忽略了成本反转,因此我们发现无界 ICL 不是学习机制的良好候选者。
英文摘要
Continual Learning Bench scores learning as short-horizon gain versus a reset baseline and finds naive full-context ICL strongest among the memories it tested. We treat ICL as one learning system and score it on a longer shared-world schedule. Steady-state learning is the gap versus baseline on a pre-set late window (last 40 of 174 BIRD-SQL formula-1 questions). We split the score into exploration efficiency (SQL probes), task reward (hits), and delivery cost (API dollars and context size). On gpt-5.6-luna, late probes fall from 4.6-5.6 to 0.95 while hits rise only modestly and ICL context grows to about 95k tokens with cost roughly doubling. Short-horizon gain understates the late probe saving and misses the cost inversion, so we find that unbounded ICL is a poor candidate for the learning mechanism.