AI 中文总结
研究面向富含图表的技术会议视频的问答问题,核心方法是开发基于大语言模型的多模态问答系统LMVQA,主要贡献是显著提高答案准确性,降低响应时间和成本,获领域专家认可。
AI 中文摘要
软件工程越来越依赖异步通信工件,如记录的会议,其中利益相关者讨论问题、基本原理和决策。这些会议常包含基于图表的需求、系统行为等表示。获取其中的知识具有挑战性。本文报告开发和评估LMVQA的行业经验,它是基于大语言模型的技术会议视频多模态问答系统。与Ciena工程师合作开发,通过音频和视觉证据为答案提供依据,处理视频构建可重用的带时间戳证据语料库。在两个数据集上,其显著提高了答案准确性,大幅降低响应时间和成本,领域专家也看重它在定位软件工程相关信息等方面的价值。
英文摘要
Software engineering increasingly relies on asynchronous communication artifacts, including recorded meetings where stakeholders discuss concerns, rationale, and decisions. These meetings often include diagram-based representations of requirements, system behavior, component interactions, and trace dependencies. Accessing knowledge from these meetings is challenging because recordings are long and relevant evidence is distributed across speech, slides, and technical diagrams. This paper reports our industrial experience developing and evaluating LMVQA, an LLM-based multimodal question-answering system for technical meeting videos. Developed in collaboration with engineers at Ciena, LMVQA supports the understanding of requirements and design intent by grounding answers in audio and visual evidence, with explicit handling of diagram-rich content such as requirements and UML diagrams. It processes each video once to build a reusable time-stamped evidence corpus for grounded question answering. Across a Ciena dataset and a public dataset, we show that LMVQA significantly improves answer accuracy compared to a state-of-the-art baseline, from 31% to 94% on the Ciena dataset and from 21% to 88% on the public dataset, with larger gains on diagram-rich videos. We further show that, after one-time indexing, LMVQA reduces average response time from 81.3s to 3.3s on Ciena and from 98.4s to 9.2s on the public dataset, while lowering average token-based LLM API cost by about 75%. Finally, our interviews with three domain experts show that engineers particularly value LMVQA for locating software-engineering-relevant information, revisiting rationale, and tracing answers to specific video segments.
CommentsAccepted for publication in ICSME 2026