AI 中文总结
提出参考绑定的 Visual Jev 奖励,通过 Qwen3.5-4B 验证器将视觉判断转为 GRPO 训练信号,在 MICo-Bench 上提升 GPT-5.4 得分至 52.50。
AI 中文摘要
多主体图像生成需要奖励机制来验证所请求的属性、动作和关系是否针对指定的参考主体成立。仅存在主体并不能确定正确的主体参与了所请求的交互。我们提出了参考绑定的 Visual Jev 奖励,将这些视觉决策转化为生成器训练信号。每个与主体相关的问题仅在请求的条件和相关的参考身份同时成立时才获得正标签。我们离线构建固定问题,使用二元监督训练一个 Qwen3.5-4B 验证器,并直接从其语言模型头读取 Yes 概率。它们的均值提供 GRPO 奖励,同时保留个体判断以供检查。使用 200 个 MICo-150K 训练任务和 30 次更新,该框架在手动选择的 897 任务 MICo-Bench 子集上将 GPT-5.4 综合得分从 41.78 提升至 52.50;直接 27B 奖励得分为 51.84。每个奖励在单次 GRPO 运行中测试,离线人工评估未建立相对于直接评分的统计显著优势。该研究提供了 Visual Jev 作为多主体图像生成参考绑定奖励的初步实现和评估。
英文摘要
Multi-subject image generation requires rewards that verify whether requested attributes, actions, and relations hold for the specified reference subjects. Subject presence alone does not establish that the correct subjects participate in a requested interaction. We present reference-bound Visual Jev rewards that turn these visual decisions into generator training signals. Each subject-related question receives a positive label only when the requested condition and the relevant reference identities hold jointly. We construct fixed questions offline, train a Qwen3.5-4B verifier with binary supervision, and directly read Yes probabilities from its language-model head. Their mean supplies a GRPO reward while retaining individual judgments for inspection. Using 200 MICo-150K training tasks and 30 updates, the framework raises a GPT-5.4 composite score from 41.78 to 52.50 on a manually selected 897-task MICo-Bench subset; direct 27B rewards yield 51.84. Each reward is tested in one GRPO run, and offline human evaluation does not establish a statistically significant advantage over direct scoring. The study provides an initial implementation and evaluation of Visual Jev as a reference-bound reward for multi-subject image generation.