arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FORGE:冻结LLM智能体的有据证据的形式最优路由

FORGE: Form-Optimal Routing of Grounded Evidence for Frozen LLM Agents

Xi Xiao, Yunbei Zhang, Chen Liu, Lin Zhao, Jialin Chen, Tianchen Zhao, Xiang Xu, Youngeun Kim, Tianyang Wang, Min Xu

arXiv 2609.34358首次发表:更新:

AI 中文总结

FORGE提出了一种针对冻结LLM智能体的联合路由框架,通过熵正则化玻尔兹曼目标和轻量级路由器,在5个基准和8个主干上以42-45%更低成本提升准确率。

AI 中文摘要

在智能体AI系统中,冻结基础模型越来越多地被部署为闭权重的API端点,使得下游适配只能通过围绕模型的输入和推理过程来实现。因此,对于每个输入查询,两个耦合的决策在很大程度上决定了答案质量和令牌成本:提供什么证据以及分配多少推理预算。这些轴上的固定默认值通常不是最优的,在我们的分析中,大约80%的查询会错误地分配支持形式或推理深度。为了解决这一挑战,我们提出了FORGE,一个统一的框架,通过在一个涵盖支持形式和思考深度的联合动作空间上进行逐查询路由来适配冻结模型。在熵正则化、成本感知的效用目标下,我们推导出一个闭式的玻尔兹曼路由目标,并将策略实例化为一个轻量级的269K参数因子化路由器。路由策略在冻结主机周围训练,无需任何权重访问,通过三阶段流程:离线臂枚举、从玻尔兹曼目标进行监督式Kullback-Leibler(KL)蒸馏,以及带有主机反馈的组相对策略优化(GRPO)细化。在5个知识密集型基准和8个从7B到671B参数的冻结主干上,FORGE在两个主要主机上以42-45%更低的令牌成本提高了准确率,以更低的令牌成本在主机间零样本迁移,并在可用的情况下与内在思考预算组合。

英文摘要

In agentic AI systems, frozen foundation models are increasingly deployed as closed-weight API endpoints, making downstream adaptation possible only through the inputs and inference procedures surrounding the model. As a result, for each input query, two coupled decisions largely determine both answer quality and token cost: what evidence to provide and how much reasoning budget to allocate. Fixed defaults along these axes are often suboptimal, misallocating support form or reasoning depth on roughly 80% of queries in our analysis. To address this challenge, we propose FORGE, a unified framework for adapting frozen models through per-query routing over a joint action space that spans both support form and thinking depth. Under an entropy-regularized, cost-aware utility objective, we derive a closed-form Boltzmann routing target and instantiate the policy as a lightweight 269K-parameter factorized router. The routing policy is trained around the frozen host, without any weight access, through a three-stage pipeline: offline arm enumeration, supervised Kullback-Leibler (KL) distillation from the Boltzmann target, and Group Relative Policy Optimization (GRPO) refinement with host feedback. Across 5 knowledge-intensive benchmarks and 8 frozen backbones ranging from 7B to 671B parameters, FORGE improves accuracy at 42-45% lower token cost on both main hosts, transfers zero-shot across hosts at lower token cost, and composes with intrinsic thinking budgets where available.

Comments34 pages. Project page: https://xixiaouab.github.io/projects/FORGE/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑