arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08793cs.CL

面向异构编码智能体的技能的证据校准运行时重构

Evidence-Calibrated Runtime Reconstruction for Agent Skills Across Heterogeneous Coding Agents

Xueping Gao

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出Skill Runtime Intelligence系统,针对异构编码智能体重构技能生命周期阶段,在多场景实验中验证其能准确定位失败边界,为适配器资格认定提供支撑。

中文摘要 AI 辅助

智能体技能为使用工具的语言模型智能体封装了可复用的指令与资产。渐进式加载会产生会话、模型或工具中心的追踪无法充分表征的失败边界:技能可被发现但未激活、激活时无指令、看似成功但无独立验证的结果。本文提出Skill Runtime Intelligence,这是一种被动式运行时智能系统,可在异构测试环境中重构受支持的技能生命周期阶段,同时将未受支持的阶段保留为未知状态。其Run Panorama将不可变事件、确定性关系、推断诊断结果与受控结果划分为四个证据等级;可选的追踪导入及OTLP/HTTP导出功能可支持现有可观测性部署。在六个冻结的仓库配置文件、三个编码智能体、七个干净或注入故障的条件下,全部126次执行均保留了源工作树,且各对应恰好一个源会话。然而适配器呈现出三种不同语义:无技能运行、完整运行但无类似失败的事件、或在每一次操作失败与干净会话中均出现类似失败的事件。在一项含七个模板的诊断研究中,语义别名与Panorama定位出相同的六个非干净边界,但在精确性/状态行为上存在差异;原始视图在全部18个干净案例中均输出失败状态,而Panorama未输出任何失败状态。一个基于已知规则的图符合126/126个冻结契约,而第二个模型仅完成228/378次调用。这些观察结果推动了可执行适配器的资格认定,表明事件存在性不等同于边界保真度、复合精确分数掩盖了不同错误,且模型解释不得覆盖确定性事实。

英文摘要

Agent Skills package reusable instructions and assets for tool-using language-model agents. Progressive loading creates failure boundaries poorly represented by session-, model-, or tool-centric traces: a Skill can be discovered but not activated, activated without instructions, or appear successful without an independently verified outcome. We present Skill Runtime Intelligence, a passive runtime-intelligence system that reconstructs supported Skill-lifecycle stages across heterogeneous harnesses while preserving unsupported stages as unknown. Its Run Panorama separates immutable events, deterministic relations, inferred diagnoses, and controlled outcomes with four evidence grades; optional trace import and OTLP/HTTP export support existing observability deployments. Across six frozen repository profiles, three coding agents, and seven clean or fault-injected conditions, all 126 executions preserve source worktrees and each correlates to exactly one source session. Yet adapters expose three distinct semantics: no Skill runs; complete runs but no failure-like events; or failure-like events in every operational-failure and clean session. In a seven-template diagnostic study, semantic aliases and Panorama localize the same six non-clean boundaries but differ in exact/status behavior; both Raw views emit a failure status on all 18 clean cases, while Panorama emits none. A known-rule graph conforms to 126/126 frozen contracts, whereas a second model completes only 228/378 calls. These observations motivate executable adapter qualification and show that event presence is not boundary fidelity, composite exact scores mask distinct errors, and model explanations must not overwrite deterministic facts.

发表机构

  • Alibaba Cloud(阿里云)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑