arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28011cs.AI

覆盖度而非信用:零阶扰动预算的失败-信用路由并不能提升LLM智能体的池内样本效率

Coverage, Not Credit: Failure-Credit Routing of Zeroth-Order Perturbation Budgets Does Not Improve On-Pool Sample Efficiency for LLM Agents

  • University of York(约克大学)

机构由 AI 辅助整理,请以论文原文为准。

Yuxu Ge

AI总结:

该研究探究将ZO/ES扰动预算按LLM智能体失败信用路由是否提升池内样本效率,发现多数场景下无显著增益,仅在未见过的BFCL函数的保留端点上有一处例外,还记录了三种相关实验失败模式。

AI中文摘要:

轨迹级信用分配可仅通过可验证信号定位使用工具的LLM智能体中导致失败的模块。我们探究此类失败信用是否应路由固定的零阶/进化策略(ZO/ES)扰动预算。在合成环境及冻结的Qwen2.5-1.5B/3B、SmolLM2-1.7B智能体,三类任务族、六种分配方案、信用噪声扫描、配对种子、精确符号翻转测试的设置下,我们发现所有池内比较中,均匀分配均无统计学可检测的提升(无至少2个百分点的增益)。联合soft-plus-sigma方案在1.5B和3B模型上,AUC在±0.02范围内与均匀分配等价;将全部预算集中于信用argmax在1.5B模型上(该模块为已验证瓶颈)边际等价,在3B模型上则显著更差。逆倾向去偏无法挽救路由,在BFCL衍生任务族上,错误路由的内部AUC损失达-0.074,端到端损失达-0.118。在六个固定步长调度中,损失与瓶颈饥饿率呈线性关系(R²=0.94,具描述性),且预注册的无信用覆盖度基线消除了检测到的危害。匹配预算的突发和步长补偿追赶调度与危害源于累积参数移动不足而非更新频率一致。我们的主要估计量是固定任务池上的优化效率。在未见过的BFCL函数上,本研究的唯一例外是软路由在保留端点上优于均匀分配(+0.047,p=0.031,n=6)。一种合理但未验证的解读是,路由偏好的调用者改进可迁移,而均匀分配的池内增益反映了特定于我们框架的合成器行为。我们明确报告此例外,并记录了三种可能使冻结LLM的ZO/ES实验无效的失败模式。

英文摘要:

Trajectory-level credit assignment can localize which module of a tool-using LLM agent causes failures using only verifiable signals. We ask whether such failure credit should route a fixed zeroth-order/evolution-strategies (ZO/ES) perturbation budget. Across a synthetic environment and frozen Qwen2.5-1.5B/3B and SmolLM2-1.7B agents, three task families, six allocation schemes, a credit-noise sweep, paired seeds, and exact sign-flip tests, we find no statistically detectable improvement over uniform allocation in any on-pool comparison (no gain of at least 2 percentage points). The joint soft-plus-sigma scheme is equivalent to uniform within a +/- 0.02 AUC margin on 1.5B and 3B; concentrating the full budget on the credit argmax is marginally equivalent on 1.5B, where that module is the verified bottleneck, and significantly worse on 3B. Inverse-propensity debiasing does not rescue routing, and misrouting costs up to -0.074 AUC in-house and -0.118 end-to-end on the BFCL-derived family. Across six fixed-step schedules, loss is linear in bottleneck starvation rate (R^2 = 0.94, descriptive), and a preregistered credit-free coverage floor removes detected harm. Matched-budget burst and step-compensating catch-up schedules are consistent with harm arising from insufficient cumulative parameter movement rather than update frequency. Our primary estimand is optimization efficiency on a fixed task pool. On unseen BFCL functions, the study's one exception is that soft routing exceeds uniform on held-out endpoints (+0.047, p = 0.031, n = 6). A plausible but untested reading is that routing-favored caller improvements transfer while uniform's on-pool gains reflect a synthesizer behavior specific to our harness. We report this exception explicitly and document three failure modes that can silently invalidate ZO/ES experiments on frozen LLMs.

补充信息

↑