arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14408cs.AI

动态学习解决方案:个性化教育视频生成系统

Dynamic Learning Solutions: A System for Personalized Educational Video Generation

Siddhanth Sridhar, Shreya Chaurasia, Baddela Sai Yaswantha Reddy, Deepak Parmar, Shylaja S S

首次发表
浏览论文内容

中文总结 AI 辅助

该系统利用RAG检索教科书内容并生成脚本,结合Stable Diffusion、DynamiCrafter和语音合成,将静态PDF转化为个性化交互式视频讲解。

中文摘要 AI 辅助

我们提出了一种自动化流水线,可将NCERT教科书转换为直接响应用户查询的交互式视频讲解。用户上传PDF并提出问题;系统随后生成基于视频的讲解作为输出,处理PDF中的文本和视觉元素,以进行多模态检索和响应生成。该流水线结合了检索增强生成(RAG)模型与生成式多媒体组件。RAG阶段针对NCERT教科书的结构进行了优化,并在这些书籍的内容上表现最佳。给定用户查询,RAG模型从PDF中检索相关内容,并生成包含叙事性解释和结构化视觉提示的多场景脚本,这些提示与教科书的解释风格保持一致。这些提示被传递给Stable Diffusion模块,该模块逐层实现以实现可解释性和控制,生成上下文相关的图像。随后,图像由DynamiCrafter处理以生成动画序列。最后,Google文本转语音模块生成同步的旁白,通过基于时间的控制使语音与视觉场景对齐。结果是一个连贯的视频讲解,集成了动画、旁白和与教科书对齐的视觉内容,将静态教育材料转变为引人入胜的学习体验。通过结合多模态文档检索、生成式视觉模型、动画框架和语音合成,该流水线展示了一种可扩展的方法,用于提供交互式、个性化的数字教育内容。

英文摘要

We present an automated pipeline that converts NCERT textbooks into interactive video explanations that respond directly to user queries. A user uploads a PDF and asks a question; the system then generates a video-based explanation as output, handling both text and visual elements from the PDF for multi-modal retrieval and response generation. The pipeline combines a Retrieval-Augmented Generation (RAG) model with generative multimedia components. The RAG stage is optimized for the structure of NCERT textbooks and performs best on content from those books. Given a user query, the RAG model retrieves relevant content from the PDF and generates a multi-scene script containing narrative explanations and structured visual prompts aligned with the textbook's explanatory style. These prompts are passed to a Stable Diffusion module, implemented layer by layer for interpretability and control, which generates contextually relevant images. The images are then processed by DynamiCrafter to produce animated sequences. Finally, a Google Text-to-Speech module generates synchronized narration, aligning speech with the visual scenes through time-based control. The result is a coherent video explanation integrating animation, narration, and textbook-aligned visuals, transforming static educational material into an engaging learning experience. By combining multi-modal document retrieval, generative visual models, animation frameworks, and speech synthesis, this pipeline demonstrates a scalable approach to delivering interactive, personalized digital education content.

发表机构

  • PES University(PES大学)

机构由 AI 辅助整理,请以论文原文为准。

↑