arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越二元偏好:面向肢体动作描述的分级偏好优化

Beyond Binary Preferences: Graded Preference Optimization for Limb-Motion Captioning

Yanan Wang, Tingsong Li, Kaixun Jiang, Chongyang Zhong, Chenwei Xoe, Zhaohe Liao

arXiv 2609.36628首次发表:更新:

发表机构

Zhejiang University; Alibaba Group; University of Science and Technology of China; Fudan University; Shanghai Jiao Tong University(浙江大学; 阿里巴巴集团; 中国科学技术大学; 复旦大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出FlexBench基准和GPA评分,并基于此设计GM-DPO优化方法,以解决VLM在肢体动作描述中的错误识别问题,显著提升动作准确性和降低幻觉率。

AI 中文摘要

视觉语言模型(VLMs)能够生成丰富的视频描述,但常常错误识别执行动作的人物或涉及的肢体,尤其在镜头切换时。改进这些细节需要能够区分信息缺失与错误断言的评估和训练方法。我们引入了FlexBench基准,涵盖3,105个镜头和18,161个评估查询,具有人工验证的身份和系统性的逐人覆盖细粒度肢体动作和状态。其参考派生的检查清单支持在人物和镜头上下文中对完整描述进行自动评估。我们的分级物理对齐分数(GPA)对正确内容给予奖励,并对错误或捏造的动作扣分,使这些错误在总分中明确体现。基于这一评分标准,我们提出了分级边际直接偏好优化(GM-DPO),该方法为更严重的动作错误分配更强的偏好边际和更大的训练权重。在三个VLM骨干网络上,GM-DPO在评估的偏好目标中取得了最高的实质性动作和GPA分数,相较于DPO将GPA提高了2.02至3.40分。在Qwen3-8B上,它相对于DPO将加权幻觉率降低了21.3%。这些提升伴随着持续的长格式输出、改进的镜头结构,以及在三个额外多模态基准上的竞争性表现。

英文摘要

Vision-Language Models (VLMs) can generate rich video captions, yet often misidentify which person performs an action or which limb is involved, particularly across camera cuts. Improving these details requires evaluation and training that distinguish missing information from incorrect assertions. We introduce FlexBench, a benchmark spanning 3,105 shots and 18,161 evaluation queries, with human-verified identities and systematic per-person coverage of fine-grained limb actions and states. Its reference-derived checklists support automated assessment of complete captions in their person and shot contexts. Our Graded Physical Alignment score (GPA) awards credit for correct content and deducts points for incorrect or fabricated actions, making these errors explicit in the aggregate score. Building on this rubric, we propose Graded Margin Direct Preference Optimization (GM-DPO), which assigns stronger preference margins and greater training weight to more severe action errors. Across three VLM backbones, GM-DPO achieves the highest substantive-action and GPA scores among the evaluated preference objectives, improving GPA over DPO by 2.02-3.40 points. On Qwen3-8B, it reduces the weighted hallucination rate by 21.3% relative to DPO. These gains accompany sustained long-form output, improved shot structure, and competitive performance on three additional multimodal benchmarks.

Comments18 page, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑