arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27911cs.LG

TACIT-Switch:基于审查式监督的大语言模型智能体成本感知模型升级方案

TACIT-Switch: Cost-Aware Model Escalation for LLM Agents from Censored Supervision

  • The Hong Kong Polytechnic University(香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Ji'an Lei, Jian Huang

中文总结 AI 辅助

TACIT-SWITCH从审查式监督学习切换策略,在可比成本下提升大语言模型智能体成功率7.4-11.1个百分点,在ALFWorld和DABench上取得最优保留成功率。

中文摘要 AI 辅助

采用较小语言模型主干的智能体成本更低,但可能陷入持续失败模式;而采用更大主干的智能体通常更可靠但成本更高。这种可靠性-成本权衡促使人们研究路由方法,以决定何时调用更大主干的智能体:在执行前、固定轨迹前缀后,还是在单个步骤本地进行。我们的方法TACIT-SWITCH从累积轨迹证据和教师标注的审查式干预时间(TACIT)中学习永久切换策略,将每个标注表示为累积风险尺度上的区间审查式观测值。所得的混合治愈阈值模型估计配对强推理rollout成功的概率,以及在成功条件下的切换阈值,部署时无需教师参与。在基于机制的多步骤模拟中,TACIT-SWITCH在可比成本下,比任务级、步骤级和固定前缀路由基线的成功率提高了7.4至11.1个百分点。在该受控模拟中,消融实验表明任务特征和累积轨迹风险提供了互补信息。基于开发数据选择操作点后,TACIT-SWITCH在ALFWorld(4B低成本模型下为48.5%,9B低成本模型下为45.5%)和DABench(73.1%)的学习策略中实现了最高的保留成功率。

英文摘要

Agents with smaller language-model backbones are less expensive but can drift into persistent failure modes, whereas those with larger backbones are generally more reliable but more costly. This reliability-cost trade-off motivates routing methods that decide when to invoke an agent with a larger backbone: before execution, after a fixed trajectory prefix, or locally at individual steps. Our method, TACIT-SWITCH, learns permanent handoff policies from accumulated trajectory evidence and Teacher-Annotated Censored Intervention Times (TACIT). It represents each annotation as an interval-censored observation on a cumulative-risk scale. The resulting mixture-cure threshold model estimates the probability that the paired Strong rollout succeeds and, conditional on success, the handoff threshold; no teacher is required at deployment. In a mechanism-based multi-step simulation, TACIT-SWITCH improves success by 7.4-11.1 percentage points over task-level, step-level, and fixed-prefix routing baselines at comparable cost. Within that controlled simulation, ablations show that task features and cumulative trajectory risk provide complementary information. With operating points selected on development data, TACIT-SWITCH achieves the highest held-out success among learned policies on both ALFWorld (48.5% with 4B Cheap; 45.5% with 9B Cheap) and DABench (73.1%).

补充信息

↑