Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
通过迭代偏好对齐平衡医疗AI助手的安全性与帮助性
机构 * University of Maryland(马里兰大学) ; Oracle Labs(Oracle实验室) ; Oracle Health AI(Oracle健康AI)
专题命中 医学数据与评测 :healthcare AI(title)
AI总结 本文提出通过迭代偏好对齐框架提升医疗AI助手的安全性与帮助性,通过KTO和DPO优化改进模型,提升有害查询检测性能并平衡安全与效用。
Comments ML4H 2025 Proceedings, Best Paper Award