arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Toronto(多伦多大学)

2026-02-09 至 2026-02-09 共收录 1
2601.21112 2026-02-09 cs.AI cs.SE

How does information access affect LLM monitors' ability to detect sabotage?

信息访问如何影响大语言模型监视器检测破坏行为的能力?

Rauno Arike, Raja Mehta Moreno, Rohan Subramani, Shubhorup Biswas, Francis Rhys Ward

机构 * Aether Research(Aether研究机构) Vector Institute(Vector研究所) Trajectory Labs(Trajectory实验室) University of Toronto(多伦多大学) Imperial College London(伦敦帝国学院)

AI总结 本文研究了信息访问对大语言模型监视器检测破坏行为的影响,提出了一种新的分层监控方法EaE,通过提取和评估被监视代理的轨迹片段来提升检测效果。

Comments 54 pages, 34 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏