SemReward-VL:面向发育行为评估的语义奖励引导视频-语言适配
SemReward-VL: Semantic Reward-Guided Video-Language Adaptation for Developmental Behavior Assessment
AI总结:
针对发育筛查视频中行为描述缺失的问题,提出SemReward-VL,利用视觉-语言模型生成描述,并通过冻结语言模型评分和GRPO优化LoRA适配器,在13,379个视频上提升了准确率与召回率。
AI中文摘要:
发育筛查视频展示了儿童如何执行特定行为,但临床记录通常只包含结果而非行为过程的描述。我们提出了SemReward-VL,该方法从这些结果中学习描述项目特定行为。一个视觉-语言模型生成描述,一个冻结的语言模型对其与临床结果的一致性、与项目的相关性、对无关视频-项目对的弃权(不执行)以及清晰度进行评分。组相对策略优化(GRPO)利用这一语义奖励更新LoRA适配器。在覆盖41个项目的13,379个视频上,该方法提高了总体准确率和召回率高于0.5的项目数量。错误仍存在于时间方向、持续时间和特定年龄的行为解释中。
英文摘要:
Developmental screening videos show how children perform specific behaviors, but clinical records usually contain outcomes rather than descriptions of what happened. We present SemReward-VL, which learns to describe item-specific behavior from these outcomes. A vision-language model generates a description, and a frozen language model scores its agreement with the clinical outcome, relevance to the item, abstention on unrelated video-item pairs, and clarity. Group relative policy optimization (GRPO) updates LoRA adapters using this semantic reward. On 13,379 videos covering 41 items, the method improves aggregate accuracy and the number of items with recall above 0.5. Errors remain in temporal direction, duration, and age-specific interpretations of behavior.