利用推理时计算资源与部署框架提升评估真实性
Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds
浏览论文内容
中文总结 AI 辅助
针对模型对齐评估中存在的评估感知问题,提出批评细化与DISH两种技术,二者组合使用可更有效提升对齐评估的真实性。
中文摘要 AI 辅助
对齐评估的核心障碍是评估感知:能力较强的模型能判断出自己是在接受测试而非实际部署,这会削弱安全评估所能支撑的结论。我们提出两种技术,让模拟对齐评估更难与真实部署区分。第一种技术是批评细化,它为每个模拟器动作投入额外的推理时计算资源:模拟器生成多个候选动作,利用目标模型的反馈细化这些动作以使其更真实,再选择最接近部署场景的候选动作继续评估。第二种技术是DISH(Deployment-Imitating SWE-Agent Harness,部署模仿型软件工程智能体框架),它将目标模型封装在智能体框架中,缩小编码场景下模拟与真实部署环境的差距。我们在多个目标模型上测试这些技术,发现它们可组合使用:同时应用两种技术比单独使用任何一种都能获得更大的真实性提升。我们的结果表明,自动化方法可提升对齐评估的真实性,且这些改进比延长审计时长更能有效利用额外计算资源。
英文摘要
A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are being tested rather than deployed, weakening the conclusions a safety evaluation can support. We present two techniques that make simulated alignment evaluations harder to distinguish from real deployments. Our first technique, critique refinement, spends additional inference-time compute on each simulator action: the simulator generates multiple candidate actions, refines them using feedback from an instance of the target model on how to make them more realistic, and continues the evaluation with the most deployment-like candidate. Our second technique, DISH (Deployment-Imitating SWE-Agent Harness), wraps the target in an agent harness, reducing the gap between simulated and real deployment environments in coding settings. We test the techniques on multiple target models and find that they compose: applying both yields larger realism gains than either alone. Our results show that automated approaches can improve the realism of alignment evaluations, and that these improvements use additional compute more effectively than making the audits longer.
发表机构
- Meridian Visiting Researcher Programme(梅里登访问研究员计划)
- Cambridge Boston Alignment Initiative(剑桥-波士顿对齐倡议)
- UK AI Security Institute(英国人工智能安全研究所)
- Meridian Labs(梅里登实验室)
- Anthropic(Anthropic公司)
机构由 AI 辅助整理,请以论文原文为准。