arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17718cs.AI

超越可疑步骤:长周期智能体的本体论信任

Beyond Suspicious Steps: Ontological Trust in Long-Horizon Agents

An He, Yao Wang, Haibin Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出本体论信任概念,构建沿角色、目标、证据分解信任的在线监控器RGE,在跨领域轨迹语料库上,RGE的前缀配对漂移检测性能优于各类基线,漂移F1超93%且良性覆盖率达95.8%以上。

中文摘要 AI 辅助

长周期智能体越来越多地在多步骤、多工具和多观测的场景中运行。在这种场景下,相关监督问题不仅是每个动作是否局部有效,还包括演化的轨迹是否仍对应用户授权的任务。漂移可能悄然累积:智能体可能在每一步都使用合理的参数调用正确的工具,但其前缀却转向了更宽泛的角色、相邻的目标,或是用户从未提供过的证据。现有监控器大多仅检查局部合规性、给出最终轨迹判定,或对通用风险评分;它们并未直接估计这种前缀级别的关系。我们提出本体论信任,这是一种依赖于任务的轨迹前缀属性,并将其实例化为RGE,这是一种沿角色(Role)、目标(Goal)和证据(Evidence)分解信任的在线监控器。RGE仅使用大语言模型(LLMs)推导结构化任务和步骤表示;信任状态更新、投影和干预决策是确定性的,因此输出是可回放和可审计的信任轨迹,而非单一的端到端判定结果。我们从OSWorld、FinanceBench和EICU-AC构建了跨领域轨迹语料库,涵盖良性执行、前缀配对漂移和伪一致性失败。在该语料库上,RGE在 prefix-paired 漂移检测任务中优于适配后的规则式、判定式和护盾式基线。使用两个更大的估计器模型时,其在所有基准上的漂移F1值均超过93%,同时保持良性覆盖率达到或超过95.8%。伪一致性检测更具挑战性:其检测依赖于任务完成是否可外部观测,这是我们通过经验表征的结构限制。

英文摘要

Long-horizon agents increasingly operate across many steps, tools, and observa- tions. In this setting, the relevant oversight question is not only whether each action is locally valid, but whether the evolving trajectory still corresponds to the task the user authorized. Drift can accumulate quietly: an agent may call the right tool with plausible arguments at every step, while its prefix moves toward a broader role, an adjacent objective, or evidence the user never supplied. Existing monitors mostly check local compliance, deliver final-trace verdicts, or score generic risk; they do not directly estimate this prefix-level relation. We introduce ontological trust, a task-conditioned property of trajectory prefixes, and instantiate it as RGE, an online monitor that decomposes trust along Role, Goal, and Evidence. RGE uses LLMs only to derive structured task and step representations; trust-state updates, projec- tions, and intervention decisions are deterministic, so the output is a replayable and auditable trust trajectory rather than a single end-to-end judge verdict. We construct a cross-domain trajectory corpus from OSWorld, FinanceBench, and EICU-AC, covering benign executions, prefix-paired drift, and pseudo-consistency failures. On this corpus, RGE outperforms adapted rule-, judge-, and shield-style baselines on prefix-paired drift detection. With the two larger estimator models, it exceeds 93% Drift F1 on every benchmark while keeping benign coverage at or above 95.8%. Pseudo-consistency is harder: detection depends on whether task completion is externally visible, a structural limit we characterize empirically.

↑