arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多头潜在控制:大语言模型智能体决策的统一接口

Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making

Amirhosein Ghasemabadi, Ruichen Chen, Bahador Rashidi, Di Niu

arXiv 2607.14277首次发表:更新:

发表机构

University of Alberta; Huawei Technologies Canada Co., Ltd.(阿尔伯塔大学; 华为加拿大技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大语言模型作智能体时的决策控制问题,提出多头潜在控制方法,通过读取冻结模型的隐藏状态轨迹生成控制信号,能在不修改模型的情况下进行事后适配,改善多模型系统质量成本权衡,减少大模型使用并提高工具使用决策质量。

AI 中文摘要

大语言模型越来越多地被用作智能体,但可靠的智能体行为需要的不仅仅是下一个token预测。在推理时,智能体最好能决定是继续当前推理、听从更强的模型、请求更多信息、调用外部工具还是弃权。现有方法通过提示级路由、外部编排或特定任务微调来处理这些决策,主要依赖输入端信号,且随着模型主干的发展成本高且难以维护。本文提出多头潜在控制,从冻结的大语言模型或视觉语言模型中读取隐藏状态轨迹以生成部署时控制信号,包括能力头和分辨率头,仅在相同冻结模型主干的潜在轨迹上训练,能在不修改模型的情况下进行事后适配。跨语言和视觉语言设置,该方法持续改善多模型系统的质量成本权衡,减少大模型使用,提高工具使用决策质量。

英文摘要

Large language models are increasingly deployed as agents, but reliable agentic behavior requires more than next-token prediction. At inference time, it is preferred that an agent can decide whether to proceed with its current reasoning, defer to a stronger model, request additional information, invoke external tools, or abstain under the given setup. Existing approaches address these decisions through prompt-level routing, external orchestration, or task-specific fine-tuning, which primarily rely on input-side signals, and are often costly and difficult to maintain as model backbones evolve. We ask whether such control decisions can be inferred directly from a model's latent generation process. We introduce Multi-Head Latent Control, a lightweight layer that reads hidden-state trajectories from a frozen LLM or VLM to produce deployment-time control signals. A Capability Head predicts whether the current model can solve the instance or should defer to a stronger collaborator, while a Resolution Head predicts appropriate resolution decision Clarification, Tool Use, Abstention, or Direct Answering. Both heads are trained only on latent traces from the same frozen LLM backbone, enabling post hoc adaptation without modifying the model. Across language and vision-language settings, Multi-Head Latent Control consistently improves the quality-cost tradeoff of multi-model systems, enabling early handoff from partial generations and more accurate intervention decisions. In routed execution (small + large model), it reduces large-model usage by up to 90.7 percent on AndroidWorld and 27-53 percent on average across benchmarks, while retaining most of large-model performance. Additionally, the learned control signals improve tool-use decision quality, yielding up to +158 percent relative score gain and 65.5 percent fewer missed-required tool calls.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑