发表机构
MIT; NVIDIA; Stanford University(麻省理工学院; 英伟达; 斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CARE是一个端到端人机协同决策框架,通过自适应校准与修正保证安全并持续学习,在四个安全关键数据集上减少25-81%人工查询。
AI 中文摘要
在人机协同决策中,人工审查可以防止不安全的AI决策,但每次人工判断成本高昂。将AI弃权(不执行)后的人工干预视为一次性回退,会错失改进未来AI决策以实现更高自动化的机会,然而AI从选择性查询的人工反馈中自适应学习,会破坏为旧模型校准的安全护栏。我们通过CARE——校准自适应修正与升级——应对这一挑战,这是一个端到端流水线,结合AI模型与人工审查者,以保证安全、符合人类对齐的决策,同时持续从人工反馈中学习,以更少的人工查询实现更高自动化。CARE是有原则的、通用的、模块化的,适用于任何黑盒AI模型。我们新颖的自适应校准模块保证任何修正模块在每个时间步的风险控制。我们进一步展示了当AI模型训练良好且人机不对齐具有清晰结构时,CARE如何提高查询效率。在涵盖驾驶、语言和机器人的四个安全关键真实世界数据集上的实验表明,CARE实现了符合人类对齐的决策,同时相对于基线将人工查询减少了25-81%。
英文摘要
In human-AI collaborative decision making, human review can prevent unsafe AI decisions, but each human judgment is costly. Treating human intervention after AI abstention as a one-off fallback misses the opportunity to improve future AI decisions for greater automation, yet AI adaptively learning from selectively queried human feedback breaks safety guardrails calibrated for old models. We approach this challenge with CARE---calibrated adaptive rectification and escalation---an end-to-end pipeline that combines AI models and human reviewers to guarantee safe, human-aligned decisions, while continuously learning from human feedback to achieve greater automation with fewer human queries. CARE is principled, general, modular, and works with any black-box AI model. Our novel adaptive calibration module guarantees risk control at every time step for any rectification module. We further show how CARE improves query efficiency when the AI model is well trained and the human-AI misalignment has a clear structure. Experiments on four safety-critical real-world datasets spanning driving, language, and robotics demonstrate that CARE achieves human-aligned decisions while reducing human queries by 25-81% relative to baselines.