arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35149cs.AIcs.MAcs.SE

从迁移到校准:跨模型、司法管辖区和规模保持智能体能力

From Migration to Calibration: Preserving Agent Capabilities across Models, Jurisdictions, and Scale

Yaxiao Liu, Pengbo Liu, Yiwen Liu, Yihua Guan, Jiaxing Song

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出智能体校准框架,通过信息保留、工具适应和用户接受三层约束行为适应,确保模型更换、跨境和规模扩展时能力不退化并满足目标合约。

中文摘要 AI 辅助

当部署条件发生变化时,智能体需要进行校准:更换驾驶模型,包括从基础模型到后训练模型的转换;跨越司法管辖区;或在异构市场和来源之间进行扩展。仅接口兼容性并不能确保能力保持或目标合约的满足。我们将智能体校准定义为跨三个相互作用的层面的受约束行为适应:信息保留、工具(harness)适应和用户接受;这些层面适用于每种场景,而非与三种场景一一对应。基本目标是在满足目标要求的同时,在预先指定的能力指标上实现非退化;聚合改进是更强的目标。信息校准保留经过独立验证的、仍适用于目标任务的源内容。工具校准在语义检查点对齐可观察的工件,并通过迭代、工具替换或在明确预算内的局部重规划来修复它们。用户校准强制执行针对接收者的输出合约:模板、模式和章节级偏好。一个全球电子商务示例展示了共享标准如何与站点和市场特定的适配器及验证共存。我们区分可训练策略与冻结骨干配置或控制器优化,以及证据验证与相对判断和DPO/GRPO优化。最近的工具迁移和评判有效性研究推动了目标原生执行记录、任务有效性和接近平局排名的独立审计,以及匹配的目标原生优化控制。我们提出针对模型变更、跨境适应和规模的留出评估,包括源证据和检查点修复的因子测试,以及群体级报告以防止聚合收益掩盖局部失败。这是一项方法论提案;实施和实证验证仍是未来工作。

英文摘要

Deploying, migrating, or scaling an agent can change its model, harness, infrastructure, application, and intended users. We formulate agent calibration as standards-first adaptation: define basic-capability, technical-environment, and user-context standards; diagnose gaps; generate and apply revisions; and recheck the same standards within fixed budgets. These standard families interact across information, harness, and user-acceptance layers. Source behavior is diagnostic, not a perfect reference or capability ceiling: model replacement can turn correct answers into errors or errors into correct answers. Qualification requires all mandatory known tests, actual end-to-end deployment paths, hard predicates, and declared task/user minimums to pass; aggregate gains cannot erase hard failures. Revisions may change tools or harnesses, add demonstrations and task descriptions, or use validated target-native trajectories to train a policy, controller, or compact skill model served through the harness. Semantic checkpoints validate executed artifacts, localize repair, and revalidate dependencies. The loop exports reusable configuration or training artifacts with a qualification record, while final task outputs undergo their own checks. Independent factual evidence precedes relative preference judgment; DPO and GRPO optimize policies rather than establish truth. Frozen held-out evaluation tests generalization and compares equal-budget target-native optimization. We specify an automatic calibration tool using limited authorized user trajectories and tests as future work. The framework and tool remain proposals; confirmatory empirical validation is pending.

发表机构

  • PwC China AI Center(普华永道中国人工智能中心)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑