发表机构
UC Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Concord 提出视频关系代数(VRA)及近似优化,通过处理转录或轨迹级连接减少 MLLM 使用,在真实视频上降低高达 87% 成本并提升 F1 分数。
AI 中文摘要
语义视频查询允许用户嵌入自然语言提示,并使用多模态大语言模型(MLLM)来解释视频。此类查询在查询视频数据方面日益流行。然而,其表达能力带来了高昂的代价:MLLM 可能处理数小时的媒体,却只返回几秒钟的相关输出,使得朴素执行变得缓慢、昂贵且不准确。我们提出了 Concord,一个用于表达和优化语义视频查询的系统。我们做出了三项贡献。首先,我们引入了视频关系代数(VRA),这是一种基于视频、转录文本、帧和对象轨迹的嵌套代数,能够捕获常见的语义视频操作。其次,我们推导出一组近似优化,通过重写 VRA 查询来减少 MLLM 的使用,同时提高结果质量。对于有叙述的视频,Concord 要么处理转录文本而非视频,要么利用转录文本识别出供 MLLM 处理的视频片段。对于无叙述的跨摄像头查询,检测和跟踪取代了全视频 MLLM 连接,转而使用轨迹级关系连接。第三,我们在真实世界视频上评估了 Concord。在 4.59 小时的足球转播和 3.92 小时的讲座中,转录到视频查询仅将源视频时长的 5.32% 和 2.47% 发送给 MLLM,并将 MLLM 成本降低了高达 87%。在两个各五秒的高速公路片段中,包含 18 辆经人工裁决的跨摄像头车辆,Detect-Track-Join 查询将 F1 分数从 0.364 提升到 0.813,且未调用任何 MLLM。请参阅我们的项目,网址为 https://this URL。
英文摘要
Semantic video queries let users embed natural language prompts and use multimodal large language models (MLLMs) to interpret the video. Such queries are increasingly popular for querying video data. However, their expressiveness comes at a steep cost: an MLLM may process hours of media to return only seconds of relevant output, making naive execution slow, expensive, and inaccurate. We propose Concord, a system for expressing and optimizing semantic video queries. We makes three contributions. First, we introduce Video Relational Algebra (VRA), a nested algebra over videos, transcripts, frames, and object tracks that captures common semantic video operations. Second, we derive a set of approximate optimizations that rewrite VRA queries to reduce MLLM usage while improving result quality. For narrated video, Concord either processes transcripts instead of video or uses them to identify video clips for MLLM processing. For cross-camera queries without narration, detection and tracking replace a whole-video MLLM join with a track-level relational join. Third, we evaluate Concord on real-world videos. Across 4.59 hours of soccer broadcasts and 3.92 hours of lectures, transcript-to-video queries send only 5.32% and 2.47% of source-video duration to the MLLM and reduce MLLM cost by up to 87%. In two five-second highway clips with 18 manually adjudicated cross-camera vehicles, a Detect-Track-Join query improves F1 from .364 to .813 while making no MLLM calls. See our project at https://concord-db.github.io/.