arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HALLELUAI:一个用于大规模超逼真图像到视频生成的幻觉感知人工智能系统

HALLELUAI: A Hallucination-Aware AI System for Ultra-Realistic Image-to-Video Generation at Scale

Aniket Sakpal, Yang Jiang, Rouzbeh Davoudi, Shayan Hassantabar, Mani Najmabadi

arXiv 2607.22959首次发表:更新:

发表机构

Expedia Group(亿客行集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对AI生成视频时质量控制难的问题,HALLELUAI系统集成视频调节与智能再生模块,能评估风险并迭代修复。经人工参与评估,该系统可输出超逼真视频,推动了可信AI视频内容发展,实现大规模图像到视频生成。

AI 中文摘要

人工智能生成的视频在营销、产品故事讲述和创意工作流程中越来越常用,但自动化高精度质量控制仍是扩大生产的主要限制。我们提出了HALLELUAI,一个端到端系统,它能调节并重新生成图像到视频的输出,以达到专家级创意标准,并大规模提供具有一致最终用户体验质量(QoE)的超逼真视频。该系统集成了一个视频调节模块和一个智能再生模块,调节模块评估帧级美学、时间运动保真度和相对于源图像的细粒度幻觉风险,再生模块通过提示细化、可控相机调整、目标模型或图像切换以及结构化重试策略迭代修复失败。调节逻辑符合特定领域的创意指南,并产生可直接驱动再生的详细、机器可操作的反馈。在与创意专家的人工参与评估中,HALLELUAI表现出很强的一致性,能够可靠地输出适合大规模产品和营销展示的超逼真、生产级视频。这个框架通过强制视觉逼真、品牌安全和严格的输入图像保真度,同时实现大规模图像到视频生成,推动了可信人工智能生成视频内容的发展。

英文摘要

AI-generated video is increasingly used across marketing, product storytelling, and creative workflows, yet automated; high-precision quality control remains a major constraint to scaling production. We present HALLELUAI, an end-to-end system that moderates and regenerates image-to-video outputs to meet expert-level creative standards and deliver ultra-realistic videos with consistent end-user quality of experience (QoE) at scale. The system integrates a video moderation module that evaluates frame-level aesthetics, temporal motion fidelity, and fine-grained hallucination risks relative to the source image, with an agentic regeneration module that iteratively fixes failures through prompt refinement, controlled camera adjustments, targeted model or image switching, and structured retry strategies. The moderation logic is aligned with domain-specific creative guidelines and produces granular, machine-actionable feedback that directly drives regeneration. In human-in-the-loop evaluations with creative experts, HALLELUAI shows strong alignment and reliably outputs ultra-realistic, production-grade videos suitable for product and marketing placements at scale. This framework advances trustworthy AI generated video content by enforcing visual realism, brand safety, and strict input-image fidelity while enabling image-to-video generation at scale.

Comments10 pages, 4 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑