arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2605.22166cs.AI

适配接口而非模型:面向确定性LLM智能体的运行时框架适配

Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents

  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

Tianshi Xu, Huifeng Wen, Meng Li

更新

AI总结:

提出Life-Harness运行时框架,通过从训练轨迹中演化出可复用的环境侧干预,在不修改模型权重或评估环境的情况下,显著提升冻结LLM智能体在确定性任务中的性能。

AI中文摘要:

LLM智能体不仅由其语言模型塑造,还受运行时框架的影响,该框架协调观察、工具使用、动作执行、反馈解释和轨迹控制。虽然现有的智能体适配方法主要更新模型参数,但在确定性、规则主导的领域中,许多失败源于模型-环境接口的不匹配。我们提出Life-Harness,一种生命周期感知的运行时框架,在不改变模型权重或评估环境的情况下改进冻结的LLM智能体。Life-Harness从训练轨迹中演化,通过将重复出现的交互失败转化为跨环境契约、程序技能、动作实现和轨迹调节的可复用干预,并在未见任务上保持固定以进行评估。在来自$\tau$-bench、$\tau^2$-bench和AgentBench的七个确定性环境中,Life-Harness在18个模型骨干上的126个模型-环境设置中改进了116个,平均相对提升88.5%。仅从Qwen3-4B-Instruct轨迹演化出的框架可迁移到其他17个模型,表明Life-Harness捕获的是可复用的环境侧结构而非模型特定行为。这些结果将运行时接口适配定位为以模型为中心的智能体训练的互补替代方案。代码可在https://github.com/Tianshi-Xu/Life-Harness获取。

英文摘要:

LLM agents are shaped not only by their language models, but also by the runtime harness that mediates observation, tool use, action execution, feedback interpretation, and trajectory control. While existing agent adaptation methods mainly update model parameters, many failures in deterministic, rule-governed domains stem from mismatches at the model--environment interface. We propose Life-Harness, a lifecycle-aware runtime harness that improves frozen LLM agents without changing model weights or evaluation environments. Life-Harness evolves from training trajectories by converting recurring interaction failures into reusable interventions across environment contracts, procedural skills, action realization, and trajectory regulation, and remains fixed for evaluation on unseen tasks. On seven deterministic environments from $τ$-bench, $τ^2$-bench, and AgentBench, Life-Harness improves 116 out of 126 model--environment settings across 18 model backbones, with an average relative improvement of 88.5%. Harnesses evolved only from Qwen3-4B-Instruct trajectories transfer to 17 other models, showing that Life-Harness captures reusable environment-side structure rather than model-specific behavior. These results position runtime interface adaptation as a complementary alternative to model-centric agent training. Code is available at https://github.com/Tianshi-Xu/Life-Harness.

补充信息

↑