arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32495cs.CRcs.AIcs.CEcs.SE

传闻:审计员能否信任已部署智能体框架所写入的记录?

Hearsay: Can an Auditor Trust the Record a Deployed Agent Harness Writes?

  • Nanyang Technological University(南洋理工大学)
  • National University of Singapore(新加坡国立大学)
  • Case Western Reserve University(凯斯西储大学)

机构由 AI 辅助整理,请以论文原文为准。

Jiahong Dai, Zhuochen Yang, Pengyang Shao, Kelvin Ng, Zhongyi Liu, Chengquan Ju, Yuting He, Bo Hu

AI总结:

研究智能体框架自写记录的可信度,提出外部追加日志作为第二作者以增强证据效力,实验表明该方法能完整报告遗漏与捏造。

AI中文摘要:

智能体框架(agent harness)是将模型转化为智能体的代码,它会自行写入每次运行的记录,而当一次运行发生争议、被调查或审计时,该记录是后续读者所能获得的全部信息。我们将这样的记录称为“具有证据效力”的,当一位不在现场的读者无需信任写入者即可对其核查时。在十六个已部署的框架中,没有一个完整地写出这样的记录。Hearsay 审视的是记录本身,而非任务:五个框架运行了十四项任务,三位盲审的大语言模型审查员和一个人工小组阅读了这些记录,并且审查员引用的每个片段都通过机械方式核查其写入者。首先,记录能让读者指出故障,但无法证明运行过程如何发生。审查员在140次运行中的74%至91%中正确指出了故障,但该故障只能从基准测试添加的两个文件中得到证明;至于中间发生的情况,每十个引用中不到一个落在框架未写入的内容上,而误报最少的审查员也只能捕捉到我们删除、改写或捏造的条目中的一半。其次,补救措施是增加第二位写入者,而非在第一位写入者上加更坚固的封印。一个仅追加的日志,记录框架与模型之间传递的内容,保存在框架之外并双向对照记录进行读取,报告了我们让框架在运行时犯下的全部28次遗漏和捏造,而对框架自身记录进行哈希链验证则全部28次通过。拿到日志后,审查员保持其故障判定,但将其更多引用建立在框架未写入的内容上。使记录成为证据的是谁写入它,而非捕获了什么。

英文摘要:

An agent harness, the code that turns a model into an agent, writes its own record of each run, and that record is all a later reader gets when a run is disputed, investigated or audited. We call a record evidentiary when a reader who was not there can check it without trusting the writer. Across sixteen deployed frameworks, none writes one in full. Hearsay examines the record, not the task: five harnesses run fourteen tasks, three blinded LLM examiners and a human panel read the records, and every excerpt an examiner quotes is checked mechanically for who wrote it. First, the record lets a reader name the fault but not prove how the run went. Examiners name the right fault in 74 to 91% of 140 runs, but the fault can be proved only from two files the benchmark adds; for what happened in between, fewer than one citation in ten lands on anything the harness did not write, and the examiner with the fewest false alarms catches half of the entries we delete, rewrite or fabricate. Second, the remedy is a second author, not a stronger seal on the first. An append-only log of what passes between harness and model, kept outside the harness and read against the record in both directions, reports all 28 omissions and fabrications we made a harness commit as it ran, where a hash chain over the harness's own record passes all 28. Handed the log, examiners keep their fault verdicts but rest more of their citations on what the harness did not write. What makes a record evidence is who writes it, not what is captured.

补充信息

↑