接下来是什么?Omni-StoryBench:用于评估故事驱动的全模态生成的基准
What Comes Next? Omni-StoryBench for Evaluating Story-Grounded Omnimodal Generation
浏览论文内容
中文总结 AI 辅助
提出Omni-StoryBench基准,评估模型跨图像、叙述和语音连贯续写故事的能力,发现编排范式优于原生全模态模型,图像生成是主要瓶颈。
中文摘要 AI 辅助
全模态评估应超越独立的文本、图像和语音生成:单独看来合理的输出可能无法表达连贯的共享事件。我们引入了Omni-StoryBench,一个故事驱动的全模态基准,用于评估模型能否在图像、叙述和语音之间连贯地继续故事。每个实例提供当前故事书页面和结构化的下一页条件,要求模型生成下一页插图、叙述和角色口语话语。Omni-StoryBench包含900个经过严格验证的、来自开放许可儿童读物的故事转换,并配有真实下一页参考和语音元数据。我们使用模态特定指标和以一致性为中心的LLM-as-a-judge评分标准来评估系统,涵盖上下文保持、条件遵循、参考一致性和跨模态连贯性。在涵盖编排、半编排和原生任意到任意范式的32种基线配置中,我们发现带有强大VLM规划的编排最可靠,而当前原生全模态模型常在输出完整性和可控性上表现不佳。我们的分析表明,文本侧性能与图像和语音质量相关,但在所评估的配置中,图像生成和视觉连续性构成了最明显的瓶颈。这些结果将Omni-StoryBench定位为一个系统级基准,用于衡量超越孤立模态质量的连贯全模态生成。
英文摘要
Omnimodal evaluation should go beyond independent text, image, and speech production: individually plausible outputs may not express a coherent shared event. We introduce Omni-StoryBench, a story-grounded omnimodal benchmark evaluating whether models can coherently continue stories across image, narration, and speech. Each instance provides a current storybook page and structured next-page conditions, requiring models to generate the next illustration, narration, and spoken character utterance. Omni-StoryBench contains 900 rigorously validated story transitions from openly licensed children's books, with ground-truth next-page references and speech metadata. We evaluate systems with modality-specific metrics and consistency-centered LLM-as-a-judge rubrics for context preservation, condition following, reference consistency, and cross-modal coherence. Across 32 baseline configurations spanning orchestration, semi-orchestration, and native any-to-any paradigms, we find orchestration with strong VLM planning most reliable, while current native omnimodal models often struggle with output completeness and controllability. Our analysis shows text-side performance is associated with image and speech quality, but image generation and visual continuity form the clearest observed bottleneck among the evaluated configurations. These results position Omni-StoryBench as a system-level benchmark measuring coherent omnimodal generation beyond isolated modality quality.
发表机构
- IPAI, Seoul National University(首尔大学IPAI研究所)
机构由 AI 辅助整理,请以论文原文为准。