arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DAEP:面向医学视频语料时序答案 grounding 的难度感知证据规划

DAEP: Difficulty-Aware Evidence Planning for Medical Video Corpus Temporal Answer Grounding

Tianjian He, Yujie Liu, Zhiping Huang, Changbo Xu

arXiv 2608.06869首次发表:更新:

发表机构

TikTok, ByteDance; Beijing Institute of Graphic Communication; Lingnan University(字节跳动TikTok; 北京印刷学院; 岭南大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

BIGC团队提出DAEP方案,通过多模态证据排名与难度感知规划,在NLPCC 2026共享任务中以0.2728的平均得分获第一名,提升了医学视频时序答案定位的性能。

AI 中文摘要

我们介绍了DAEP,它是BIGC团队提交给NLPCC 2026共享任务1赛道3:医学视频语料中难度感知时序答案定位(DA-TAGVC)的方案。该任务要求从50个候选视频中检索目标视频并定位支持答案的片段。DAEP利用字幕、视觉和程序上下文证据对视频进行排名,将高分锚点扩展为时序片段,并对片段重新排名以生成最终输出。其核心设计是将任务提供的简单/复杂输入标签转换为推理时的证据规划,以此控制模态权重、Top-K聚合、边界阈值、扩展长度和重排序强度。在官方评估中,BIGC团队在10个系统中排名第一,平均得分为0.2728。验证性 ablation实验显示,视觉证据、程序上下文以及难度感知规划提升了排名质量,在复杂问题上的增益最大。

英文摘要

We describe DAEP, team BIGC's submission to NLPCC 2026 Shared Task 1 Track 3: Difficulty-Aware Temporal Answer Grounding in Video Corpus (DA-TAGVC). The task requires retrieving the target video from 50 candidates and localizing the answer-supporting span. DAEP ranks videos with subtitle, visual, and procedural-context evidence, expands high-scoring anchors into temporal spans, and reranks spans for final output. Its main design is to convert the task-provided simple/complex input label into an inference-time evidence plan controlling modality weights, Top-K aggregation, boundary threshold, expansion length, and reranking strength. In the official evaluation, BIGC ranks first among ten systems with an Average score of 0.2728. Validation ablations show that visual evidence, procedural context, and difficulty-aware planning improve ranking quality, with the largest gain on complex questions.

Comments12 pages, 2 figures, 5 tables, accepted by NLPCC 2026 Shared Task Track 3

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑