发表机构
Qwen Business Unit of Alibaba; Southeast University(阿里巴巴通义千问业务单元; 东南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对对话系统的图像创作场景,提出三阶段多模态推荐框架,经测试可降低视觉不一致率并显著提升推荐相关指标。
AI 中文摘要
对话助手越来越多地推荐后续编辑以帮助用户完成任务,现有系统主要针对纯文本交互,图像创作对话领域探索不足。在图像创作任务中,有用的后续编辑建议必须反映用户偏好、提供多样方向且可在当前图像上执行。我们从Qwen App收集了10万条真实的多轮图像创作对话样本,发现80.1%的样本依赖图像,凸显了多模态推荐的必要性。我们采用三阶段框架解决该场景:阶段1,利用真实在线数据构建经人工审核的合适后续编辑意图表,创建SFT目标并微调多模态策略;阶段2,为使规则引导的SFT建议与实际用户选择对齐,通过用户点击反馈,采用多目标强化学习优化策略;阶段3,为减少建议编辑与当前图像间的视觉不一致,引入视觉验证器作为额外训练监督。大量实验表明,我们的框架在自动和人工评估中均显著优于基线。在覆盖数百万用户的在线用户随机A/B测试中,最终框架将视觉不一致率从3.7%降至0.9%,还使推荐点击率提升32.70%、图像采纳率提升16.32%、每位用户平均对话轮次提升39.90%(所有p值均小于0.05)。
英文摘要
Conversational assistants increasingly recommend follow-up edits to help users continue a task. Existing systems primarily target text-only interactions, leaving image-creation conversations underexplored. In image-creation tasks, useful follow-up edit suggestions must reflect user preferences, offer diverse directions, and remain executable on the current image. We collected 100,000 real multi-turn image-creation conversation samples from Qwen App and found that 80.1% are image-dependent, underscoring the need for multimodal recommendation. We address this setting with a three-stage framework. In Stage 1, we use real online data to build a human-reviewed table of appropriate follow-up editing intents, then create SFT targets and fine-tune a multimodal policy. In Stage 2, to align rule-guided SFT suggestions with actual user choices, we use user click feedback to optimize the policy through multi-objective reinforcement learning. In Stage 3, to reduce visual inconsistencies between suggested edits and the current image, we introduce a visual verifier as additional training supervision. Extensive experiments demonstrate that our framework significantly outperforms baselines on both automatic and human evaluations. In a live user-randomized A/B test with millions of users, our final framework reduces visual inconsistency from 3.7% to 0.9%. Furthermore, it significantly improves recommendation CTR by 32.70%, image take-away rate by 16.32%, and average conversation turns per user by 39.90% (all p<0.05). Project page: https://what-to-edit-next.github.io/