DA-DPO: Cost-efficient Difficulty-aware Preference Optimization for Reducing MLLM Hallucinations
DA-DPO:面向减少多模态大语言模型幻觉的高效难度感知偏好优化
机构 * ShanghaiTech University(上海科技大学) ; Lingang Laboratory(灵冈实验室) ; Shanghai Engineering Research Center of Intelligent Vision and Imaging(上海智能视觉与成像工程技术研究中心)
专题命中 跨模态检索 :MLLM(title);multimodal(abstract);分类 cs.AI
AI总结 DA-DPO通过难度感知机制优化多模态大语言模型的偏好学习,有效减少幻觉并提升模型鲁棒性与泛化能力。
Comments Accepted by TMLR