UniCreative: Unifying Long-form Logic and Short-form Sparkle via Reference-Free Reinforcement Learning
UniCreative:通过无参考强化学习统一长格式逻辑与短格式闪耀
机构 * Beihang University(北京航空航天大学) ; Baidu Inc.(百度公司) ; Renmin University of China(中国人民大学)
AI总结 本文提出UniCreative框架,通过自适应约束奖励模型和ACPO算法,实现无需监督细调的长短期创作统一,提升多任务表现并展现模型的元认知能力。
Comments Accepted to Findings of ACL 2026