Factor(U,T): Controlling Untrusted AI by Monitoring their Plans
Factor(U,T): 通过监控其计划来控制不可信的AI
专题命中 红队测试 :safety(abstract);分类 cs.AI
AI总结 Factor(U,T)通过监控AI分解计划来控制不可信AI,实验表明在仅观察自然语言指令时,监控恶意活动效果有限,而结合子任务监控则能实现高区分度和安全性。
Comments Accepted to AAAI 2026 Workshop on Trust and Control in Agentic AI (TrustAgent). 6 pages body, 8 pages total, 3 figures