arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向自主科学智能体的以人工制品为中心、感知声明的可观测性

Artifact-centered Claim-aware Observability for Autonomous Scientific Agents

Xiangyu Yin, Ming Du, Michael H. Prince, Mathew J. Cherukara

arXiv 2608.18312首次发表:更新:

发表机构

Argonne National Laboratory; Advanced Photon Source(阿贡国家实验室; 先进光子源)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对自主科学智能体的审计需求,提出以人工制品为中心、感知声明的可观测性概要,补充现有遥测与溯源标准,实现科学审计关系的有效表达。

AI 中文摘要

自主科学智能体如今越来越多地提出想法、编写代码、运行实验、分析结果,甚至撰写论文。对这些智能体进行观测和审计是必要的,但仅记录每次模型调用是不够的,科学家还需要检查系统生成的人工制品、声明及其之间的关系。这是因为科学智能体系统的故障通常分布在多个对象中:手稿声明可能引用了错误的证据,搜索过程可能选择了退化的候选对象,实验室的新颖性声明可能依赖于未明确说明的规则,或者多智能体计划可能在没有可见触发因素的情况下发生变更。现有的追踪、实验跟踪和溯源归档工具虽有价值,但其原生对象并未将这些科学审计关系作为一等对象。我们认为,自主科学系统应输出可移植的、感知声明的人工制品谱系作为最低审计层。我们提出了围绕个体、操作符、适配记录、谱系、归档、运行、流和控制命令组织的紧凑可观测性概要,其中科学声明是具有明确证据绑定和验证记录的普通个体。该概要旨在作为补充当前遥测和溯源标准的语义层,执行细节可保留在OpenTelemetry中,最终包可导出至PROV-O或RO-Crate标准。

英文摘要

Autonomous scientific agents now increasingly propose ideas, write code, run experiments, analyze results, and even draft papers. Observe and audit those agents are necessary but logging every model call is not enough, scientists also need to inspect the artifacts and claims that the systems produced and their relations. This is driven by the fact that failures in scientific agent systems are often distributed across several objects. A manuscript claim may cite the wrong evidence, a search process may select a degenerate candidate, a laboratory novelty claim may depend on an unstated rule, or a multi-agent plan may change without a visible trigger. Existing tracing, experiment tracking, and archival provenance tools are valuable, but their native objects do not make these scientific audit relations first-class. We argue that autonomous scientific systems should emit portable, claim-aware artifact lineage as a minimum audit layer. We propose a compact observability profile organized around individuals, operators, fitness records, lineage, archives, runs, streams, and steering commands. In this profile, scientific claims are ordinary individuals with explicit evidence bindings and verification records. The profile is intended as a semantic layer that complements current telemetry and provenance standards. Execution details can remain in OpenTelemetry. Final packages can export to PROV-O or RO-Crate standards.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑