面向多讲座教育推理的证据 grounded 多模态知识图谱构建
Evidence-Grounded Multimodal Knowledge Graph Construction for Multi-Lecture Educational Reasoning
浏览论文内容
中文总结 AI 辅助
该研究针对讲座视频知识保留不全问题,提出基于证据的多模态知识图谱构建流程,在神经网络讲座实验中取得高覆盖率与检索准确率,贡献可审计的知识图谱构建方法。
中文摘要 AI 辅助
讲座视频将知识分布在语音、幻灯片文本、图表、公式和演示顺序中,仅转录检索无法完全保留这些信息。本文提出一种基于证据的多模态流程,该流程对讲座进行转录、选择语义锚点、应用光学字符识别(OCR),并使用视觉语言模型仅提取有转录、OCR或视觉证据支持的概念和类型化关系。提及内容经验证并规范化为具有溯源信息的知识图谱。在三个神经网络讲座上,该流程处理了3118帧、756个转录片段和559个锚点,保留了1022个概念和312个关系提及,生成了172个规范化概念和282个关系,端点覆盖率达90.38%。初步的三问题检索测试实现了100%的top-1和top-3准确率,以及100%的平均top-5召回率。其贡献在于一种可审计的构建方法,而非追求超越现有技术的性能提升。
英文摘要
Lecture videos distribute knowledge across speech, slide text, diagrams, equations, and presentation order, which transcript-only retrieval does not fully preserve. This paper presents an evidence-grounded multimodal pipeline that transcribes lectures, selects semantic anchors, applies optical character recognition (OCR), and uses a vision-language model to extract only concepts and typed relationships supported by transcript, OCR, or visual evidence. Mentions are validated and canonicalized into a provenance-rich knowledge graph. On three neural-network lectures, the pipeline processed 3,118 frames, 756 transcript segments, and 559 anchors. It retained 1,022 concept and 312 relationship mentions, yielding 172 canonical concepts and 282 relationships with 90.38% endpoint coverage. A preliminary three question retrieval test achieved 100% top-1 and top-3 accuracy and 100% mean top-5 recall. The contribution is an auditable construction method rather than a state-of-the-art performance claim.