arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可靠性的栖身之所:智能体系统中行为属性的实验定位

Where Reliability Lives: Experimental Localisation of Behavioural Properties in an Agent System

Timothy Marsden, Matthew Collecutt, James Marsden

arXiv 2609.03192首次发表:更新:

发表机构

Taniwha AI(Taniwha AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究通过构建带权威账本的模拟定居点,干预制度认知机制与智能体认知,证实测得的行为属性与认知实质性变化可分离,且五个预声明属性未随干预改变。

AI 中文摘要

关于智能体系统的可靠性声明,隐含着将每个属性定位于模型或其周边机制之中。我们构建了一个系统,其中该定位是一个实验问题。研究对象是一个持久的模拟定居点,其权威的只追加账本会针对世界状态裁决每一次尝试的行为;被接受的历史是唯一的现实。在任何实验前,我们将心智、制度与世界分离开。在保持认知固定的情况下,我们对制度的认知机制(证据来源、信念可及性、物理证据的可读性)进行干预;预注册实验两次以相反方向否定了我们的核心预测。随后,一个注册的证伪器提供了几何结构未赋予信念通道的输入——一个经过编排的真实第一手见证,其边际价值在无见证阶段始终非正,转为正:11个随机种子中有9个,未添加任何错误归因。在保持制度执行固定的情况下,我们从四个方面对认知进行干预:消融原生心智的机制、在任务中途将其终止并重置、用冻结的前沿大语言模型(frontier-LLM)面板替代整个原生认知,以及用可信的虚假证词破坏信念。行为发生了显著变化:一个虚假陈述使每个信任的运行产生约900次徒劳动作,而不信任分支无此类动作。五个预先声明的属性在所有测试轨迹中均未改变:被接受的现实保持单一、无效尝试会被拒绝并给出类型化原因、职责在其过程结束后依然存在、未接受两次相同工作,且从未接受虚假完成(2581次面板替代主张,无虚假)。我们的声明仅限于该场景:通过干预确立,测得的行为属性与认知的实质性变化可分离。这是一个设计好的世界,而非制度群体;未测试智能体针对制度进行优化的情况。

英文摘要

Reliability claims about agentic systems implicitly locate each property somewhere: in the model, or in the machinery around it. We built a system where that location is an experimental question. The subject is a persistent simulated settlement whose authoritative append-only ledger adjudicates every attempted act against world state; accepted history is the only reality. Mind, institution and world were separated before any experiment. Holding cognition fixed, we intervened on the institution's epistemic mechanisms (evidence provenance, belief availability, physical-evidence legibility); preregistered experiments refuted our central prediction twice, in opposite directions. A registered falsifier then supplied the input the geometry had denied the belief channel, a staged veridical first-hand witness, and its marginal value, non-positive throughout the witness-free phases, turned positive: 9 of 11 seeds, zero added false attribution. Holding institutional enforcement fixed, we intervened on cognition four ways: ablating the native minds' machinery, killing and resetting them mid-task, substituting a frozen frontier-LLM panel for the entire native cognition, and corrupting beliefs with trusted false testimony. Behaviour changed dramatically: one falsehood cost each trusting run about 900 futile actions and the distrusting arm none. Five pre-declared properties did not move in any tested trajectory: accepted reality stayed singular, invalid attempts were refused with typed reasons, duties outlived their processes, no work was accepted twice, and no false completion was ever accepted (2,581 substituted-panel claims, none false). Our claim is limited to this setting: measured behavioural properties were separable from substantial changes to cognition, established by intervention. One designed world, not a population of institutions; no test of an agent optimising against the institution.

CommentsPreprint. 31 pages, 5 figures. Ancillary files: related-work search appendix, artefact manifest, verification receipts

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑