发表机构
Alibaba Cloud(阿里云)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM智能体多LoRA适配器存储与路由开销问题,提出BQ-LoRA框架,通过行为商平衡与决策保持压缩在固定秩下组织轨迹更新,实验验证其优于标准LoRA及近期方法。
AI 中文摘要
基于LLM的智能体依赖异构交互能力来完成复杂任务。现有方法通常将这些能力分布在多个LoRA适配器上,这增加了适配器的存储需求,并在推理过程中引入了路由开销。单个LoRA避免了这种开销,但在固定秩预算下从多样化的智能体轨迹中学习面临两个挑战。首先,具有不同交互轨迹和参数梯度的轨迹可能引发决策分布的等效变化,导致重复更新过度强调冗余的行为变化。其次,聚合更新可能超出适配器的秩预算,而在权重空间中对其进行近似可能会扭曲其旨在产生的决策变化。我们提出BQ-LoRA,一种通过局部行为商流形来组织轨迹更新的低秩适配框架。它包含两个模块,即行为商平衡(BQB)和决策保持压缩(DPC)。BQB从决策分布构建商流形,并根据轨迹更新方向在商切空间中的局部密度对其重新加权。DPC将平衡梯度投影到内在固定秩切空间,并通过联合控制有效权重误差和决策分布失真来重构所得目标。在AppWorld和BrowseComp-Plus上的实验将BQ-LoRA与标准LoRA及近期低秩适配方法进行了比较,同时单独的消融实验评估了两个组件的互补贡献。
英文摘要
LLM-based agents rely on heterogeneous interaction capabilities to accomplish complex tasks. Existing approaches often distribute these capabilities across multiple LoRA adapters, which increases adapter storage requirements and introduces routing overhead during inference. A single LoRA avoids this overhead, but learning from diverse agent trajectories under a fixed rank budget presents two challenges. First, trajectories with different interaction traces and parameter gradients can induce equivalent changes in decision distributions, causing repeated updates to overemphasize redundant behavioral changes. Second, an aggregated update may exceed the rank budget of the adapter, and approximating it in weight space can distort the decision changes that it is intended to produce. We propose BQ-LoRA, a low-rank adaptation framework that organizes trajectory updates through a local behavior quotient manifold. It contains two modules, i.e., behavior quotient balancing (BQB) and decision preserving compression (DPC). BQB constructs the quotient manifold from decision distributions and reweights trajectory update directions according to their local density in the quotient tangent space. DPC projects the balanced gradient onto the intrinsic fixed rank tangent space and refactorizes the resulting target by jointly controlling effective weight error and distortion of decision distributions. Experiments on AppWorld and BrowseComp-Plus compare BQ-LoRA with standard LoRA and recent low-rank adaptation methods, while separate ablations evaluate the complementary contributions of both components.