arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

短思考、智能延迟、行动与循环:面向边缘大语言模型智能体的校准推理与感知不确定性的延迟机制

Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM Agents

Amirmohammad Farzaneh, Osvaldo Simeone

arXiv 2607.26865首次发表:更新:

发表机构

Institute for Intelligent Networked Systems (INSI); Northeastern University London(智能网络系统研究所; 伦敦东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出TSDS框架,结合收敛探测器与困惑度延迟规则,经LTT流程校准后,可在边缘LLM智能体上减少43%-73%的思考计算量,同时保证奖励与云端调用率,在多类基准任务上表现优异。

AI 中文摘要

遵循ReAct范式的大语言模型智能体有望成为复杂多步任务的有力支撑,包括多跳问答、代码生成以及物理AI系统的控制。然而,当部署在边缘设备时,它们必须严格管理推理预算,同时保持可靠性,仅在本地不确定性过高无法安全行动时才延迟至云端模型。我们提出Think Short, Defer Smart(TSDS)框架,该框架协同集成了轻量收敛探测器(一旦预期行动稳定即终止设备端推理)与基于困惑度的延迟规则(将不确定行动升级至云端模型)。两种机制均通过多目标的Learn-Then-Test(LTT)流程在端到端回合轨迹上联合校准,为预期回合奖励与云端调用率提供有限样本保证。我们在四个ReAct基准上评估TSDS,涵盖算术推理(GSM8K)、多跳问答(HotpotQA)、代码生成(MBPP)及多步具身规划(家用机器人),并与仅思维校准、仅校准延迟的独立基线对比。TSDS在HotpotQA、MBPP及家用机器人任务上,相比仅延迟基线将每回合思考计算量降低43%-73%,同时保持经认证的奖励与云端调用率保证。

英文摘要

LLM agents following the ReAct paradigm are promising enablers of complex multi-step tasks, including multi-hop question answering, code generation, and control of physical AI systems. Yet, when deployed at the edge, they must tightly manage their reasoning budget while remaining reliable and deferring to a cloud-side model only when local uncertainty is too high to act safely. We propose Think Short, Defer Smart (TSDS), a framework that synergistically integrates a lightweight convergence probe, which halts on-device reasoning once the intended action has stabilized, with a perplexity-based deferral rule that escalates uncertain actions to a cloud-side model. Both mechanisms are jointly calibrated on end-to-end episode trajectories via a multi-objective Learn-Then-Test (LTT) procedure, providing simultaneous finite-sample guarantees on expected episode reward and cloud-call rate. We evaluate TSDS on four ReAct benchmarks spanning arithmetic reasoning (GSM8K), multi-hop question answering (HotpotQA), code generation (MBPP), and multi-step embodied planning (household robot), and compare against thought-calibration-only and calibrated-deferral-only standalone baselines. TSDS reduces per-episode thinking compute by 43%-65% over deferral-only baselines across HotpotQA, MBPP, and the household robot task, while maintaining certified reward and cloud-call rate guarantees.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑