arXivDaily arXiv每日学术速递 周一至周五更新

大厂专区

Anthropic

2026-08-27 至 2026-08-27 共收录 2
2608.25460 2026-08-27 cs.AI cs.LG 新提交

Training Alignment Auditors via Reinforcement Learning

通过强化学习训练对齐审计员

Paul Rosu, Rowan Wang

机构 * Anthropic

AI总结 本研究用强化学习训练LLM审计员,通过成对奖励等优化提升审计质量与真实性,误报率低于1%,且审计能力可跨框架泛化。

Comments 82 pages, 15 figures. Code, prompts, and evaluation data are available at this https URL (https://github.com/paulrosu11/training-auditing-agents-public)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.19380 2026-08-27 cs.SE cs.LG 版本更新

ClayBuddy: A Framework, Evaluation, & Mitigation of Coding Agent Failures

AgentArmor:编码代理失败的框架、评估与缓解

Kenneth Ge, Andre Assis

机构 * Anthropic Fellows Program(Anthropic Fellow 项目) Constellation

AI总结 提出AgentArmor框架,通过系统提示增强、命令分类器、三振政策等机制,缓解编码代理因规范不足、能力错误和工具错误导致的失败,显著提升安全性。

Comments Accepted at The Third Annual Conference on Language Modeling 2026 Workshop on Agent Behavior

详情

展开后加载摘要…

URL PDF HTML 收藏