Oversight Has a Capacity: Calibrating Agent Guards to a Subjective, Fatiguing Human
监督具有容量:将智能体守卫校准到主观且易疲劳的人类
专题命中 其他LLM :LLM(summary_cn,abstract);分类 cs.AI、cs.LG
AI总结 针对LLM智能体动作审批中人类评审者主观且易疲劳的问题,提出将守卫建模为成本敏感的选择性分类,并引入负载感知策略,发现过度监督反而降低安全性,形成倒U型曲线。
Comments 12 pages, 4 figures. Code and interactive demo: https://github.com/turangenesis/headroom