arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于工具结构化大语言模型推理方法的仅幅度前馈网络干预:门控评估协议及跨模型实证结果

Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method: Gated Evaluation Protocol, and Cross-Model Empirical Results

Sheng Xu, Zhen Chen, Junhua Wang, Boyuan Huang, Ke Jia, Jiadun Zhu, Yiming Xu

arXiv 2607.11183首次发表:更新:

发表机构

Alibaba Cloud; University of Toronto(阿里云; 多伦多大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究推理时FFN干预改善大语言模型结构化输出,提出无损的幅度门控方法,定义细粒度干预系统和评估协议。在多个模型上实验,AG在工具结构化任务中效果最佳,不同AG变体各有优劣,确定工具结构化推理是FFN级推理优化首要目标。

AI 中文摘要

大语言模型越来越多地作为使用工具的智能体运行,格式、参数或函数调用中的小错误可能使原本合理的响应无效。我们研究推理时前馈网络(FFN)干预,以在不重新训练模型权重的情况下改善结构化输出。项目始于正交残差投影(ORP),虽揭示了敏感的SwiGLU FFN干预位点,但常弊大于利。因此提出幅度门控(AG),一种无损替代方法,保留预训练FFN权重方向,仅在生成时调制激活幅度。定义了跨越P​​1/P2/P3和特定分支的P1s/P2a/P2b位点的细粒度干预系统,并引入评估协议,将组合预言机余量与固定配置和学习门分开,实施样本级核算,并使用任务感知指标处理二元和部分信用数据集。在Qwen3.5-9B、Qwen3-8B和Qwen2.5-7B上,AG总体呈微弱正向,但在工具结构化任务上最强。在Qwen3.5-9B上,类别级学习门将工具/结构化/智能体性能从38.66%提高到42.92%(提高4.27个百分点),Hermes函数调用任务提高约7.6个百分点。在Qwen3-8B上,Hermes JSON模式提高了11.36个百分点。Qwen2.5-7B保留了预言机余量,但当前学习门未能捕捉到,表明部署需要特定于模型和类别的路由。熵AG与牛顿-舒尔茨窗口AG的比较表明,两者都不具有统一优势。这些结果表明,工具结构化推理是安全FFN级推理优化最可靠的首要目标,同时前瞻性在线验证和更广泛的跨模型评估仍然必要。

英文摘要

Large language models increasingly operate as tool-using agents, where small format, argument, or function-call errors can invalidate otherwise plausible responses. We study inference-time feed-forward network (FFN) intervention as a way to improve structured outputs without retraining model weights. An earlier project-specific approach, Orthogonal Residual Projection (ORP), exposed sensitive SwiGLU FFN sites and non-monotonic energy effects, but its direction-changing operation produced more regressions than repairs in a key diagnostic. We therefore propose Amplitude Gating (AG), which preserves pretrained FFN weight directions and modulates activation magnitudes during decoding. AG separates candidate generation, ranking, and a prospective acceptance/fallback decision. We also introduce Per-Sample Fix-Harm Evaluation (PFHE), a paired reporting protocol that complements native task metrics with fixes, harms, preserved-correct cases, and preserved-wrong cases. On the only cross-position union that passes source-alignment audit, an exploratory offline mixed selector raises the descriptive heterogeneous-scorer Qwen3.5-9B tool-route micro-average from 38.66% to 42.92% (+4.27 percentage points); two Hermes function-call endpoints improve by +7.64 and +7.62 points. The same-output PFHE-format view records 48 fixes, 26 harms, 294 preserved-correct cases, and 2,188 preserved-wrong cases over 2,556 units, with positive paired bootstrap intervals for native and strict effects. Protocol-separated Qwen3-8B and Qwen2.5-7B analyses retain oracle headroom but no positive train-selected fixed tool route. A grouped five-fold RF diagnostic suggests weak nonlinear ranking signal but forces intervention, lacks baseline fallback and paired uncertainty, and is not deployment evidence. The results support model- and task-specific selection with strict fallback, not a universal AG switch.

Comments30 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑