arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于内容的游戏玩法视频旁白:利用视觉语言模型

Content Based Video Narration of Gameplay with Vision Language Models

Mathew Varghese

arXiv 2608.14016首次发表:更新:

发表机构

University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出一种基于内容的游戏玩法视频旁白系统,利用视觉语言模型等技术生成电竞风格解说,无需游戏专用工具,支持本地化语音,还发布了可复现基线。

AI 中文摘要

实时游戏解说十分稀缺,仅存在于职业电子竞技赛事直播中,几乎没有其他场景的相关内容。我们提出一种基于内容的视频旁白系统,该系统使用通用视觉语言模型(VLM)和文本转语音后端,为任意游戏玩法录制生成电子竞技风格的口语解说,无需游戏专用工具、引擎遥测数据或任务特定训练。该系统包含三种机制:时间马赛克打包将9个均匀采样的帧排列成单个3×3图像,使原生图像VLM能推理运动,同时每个片段仅消耗1个图像载荷而非9个;上下文条件提示将最近K条解说作为助手角色历史重播,抑制静态场景逐片段字幕生成中普遍存在的重复现象;时长条件生成与弹性对齐在提示中约束解说长度,随后对合成音频进行时间缩放或对称填充,使每个话语恰好填充其片段时隙,无需强制对齐器即可实现帧级精确复用。该实现支持云TTS或Apple硅上的6位量化40亿参数设备端TTS模型,使语音阶段完全本地化。我们报告了针对实时策略素材的定性案例研究、显示马赛克将每分钟图像载荷减少9倍的成本模型,以及观察到的故障模式——游戏状态幻觉、马赛克导致的分辨率损失、时间缩放导致的韵律伪影。我们将该系统作为可复现基线发布,并提供定量研究的评估协议,完整版本将报告相关研究结果。

英文摘要

Live game commentary is scarce: it exists for professional esports broadcasts and almost nowhere else. We present a content-based video narration system that produces spoken, esports-style commentary for arbitrary gameplay recordings using a general-purpose vision-language model (VLM) and a text-to-speech back end, with no game-specific instrumentation, no engine telemetry, and no task-specific training. Three mechanisms carry the system. Temporal mosaic packing arranges nine uniformly sampled frames into a single 3x3 image, letting an image-native VLM reason about motion while consuming one image payload per segment instead of nine. Context-conditioned prompting replays the K most recent narrations as assistant-role history, suppressing the repetition that dominates per-segment captioning of static scenes. Duration-conditioned generation and elastic alignment constrain narration length in the prompt, then time-scale or symmetrically pad the synthesized audio so each utterance fills its segment slot exactly, giving frame-accurate muxing without a forced aligner. The implementation supports either cloud TTS or a 6-bit quantized 4B-parameter on-device TTS model on Apple silicon, making the speech stage fully local. We report a qualitative case study on real-time strategy footage, a cost model showing the mosaic reduces per-minute image payloads by 9x, and a candid account of observed failure modes - hallucinated game state, resolution loss from mosaicking, and prosody artifacts from time-scaling. We release the system as a reproducible baseline, with an evaluation protocol for the quantitative study a full version will report.

Commentshttps://mathewvarghese.space/ai-powered-game-commentary-auto-narrating-gameplay-videos-with-gpt-4o/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑