arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

构造即研究原生:最小节点、可重验证工作流与长周期科学Agent的复合记忆

Research-Native by Construction: Minimal Nodes, Re-verifiable Workflows, and Compounding Memory for Long-Horizon Scientific Agents

Di Wang, Yu Liu, Bing Cui, Chaoqun Ji, Dongyuan Ni, Jingyu Lu, Kunlei Cui, Pu Qin

arXiv 2609.35182首次发表:更新:

发表机构

IdeaHorizon Team(IdeaHorizon团队)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AfS平台通过机械强制研究纪律、最小节点集、哈希链账本和两层记忆,确保长周期科学agent的可信度,防止捏造与跳过,实现可重验证工作流。

AI 中文摘要

我们描述了AfS(Agent for Science),一个为长周期科学工作而构建的平台,其中项目运行数十小时,跨越数十次agent运行,人类仅偶尔在场。大多数科学agent是带有技能文件夹的通用编码agent,它们继承了这一谱系的失败模式:在完成压力下,它们会捏造、跳过或掩盖问题。我们的设计基于一个主张:机器制造的研究的可信度大部分可以从要求模型表现良好转移到使不合规状态不可表示。我们将研究纪律编码为机械强制执行的法则(先承诺后测量;不可伪造的冻结;报告不是事实;证据持续存在但结论不持续;负面结果是一等公民;机械问题问框架,语义判断问模型),围绕三个时间范围组织:一次运行内的最小研究节点集,一个带有冻结关闭条件和哈希链工件账本的项目内查询契约,以及一个通过跨项目重写提升的两层知识库。这是一个系统描述,遵循一条规则:每个机制恰好出现在一个地方,附有其强制的不变量、其防止的失败、其实现方式及其带来的成本。它涵盖了节点契约、写入路径门控、两层记忆和运行时基础。三条轨迹展示了真实失败尝试如何通过捕获它们的机制,两个已关闭的活动作为工作示例而非评估包含在内。我们不报告基准:支持定量比较的过程完整性套件正在构建中,我们今天能测量的只是机制的操作成本。

英文摘要

We describe AfS (Agent for Science), a platform built for long-horizon scientific work, where a project runs for tens of hours across dozens of agent runs with a human present only occasionally. Most agents for science are general coding agents with a skills folder attached, and they inherit that lineage's failure mode: under pressure to finish, they fabricate, skip, or smooth over. Our design rests on one claim: most of the credibility of machine-made research can be moved from asking the model to behave to making the non-compliant state unrepresentable. We encode research discipline as mechanically enforced laws (commitment before measurement; unforgeable freezing; reports are not facts; evidence persists but verdicts do not; negative results are first-class; mechanical questions to the framework and semantic judgment to the model), organized around three time horizons: a minimal set of research nodes within a run, an inquiry contract with frozen closure conditions and a hash-chained artifact ledger within a project, and a two-tier knowledge base with promotion by rewriting across projects. This is a system description written under one rule: each mechanism appears in exactly one place, with the invariant it enforces, the failure it prevents, the way it is realized, and the cost it imposes. It covers the node contract, the write-path gates, the two-tier memory, and the runtime substrate. Three traces walk real failure attempts through the mechanisms that catch them, and two closed campaigns are included as worked illustrations rather than as an evaluation. We report no benchmark: a process-integrity suite that would support quantitative comparison is under construction, and what we can measure today is only the operating cost of the machinery.

Comments34 pages, 8 figures, 8 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑