AI 中文总结
HyperClaim是一种用于视频虚假信息检测的判别式时间超图框架,通过细粒度跨模态超图推理,在FactGuard协议下的三个数据集上准确率优于基线方法。
AI 中文摘要
视频虚假信息检测常采用全局多模态融合或自由形式多模态推理方法,但这两种范式往往无法充分表示查询短语、上下文文本与帧的短时间跨度之间耦合交互产生的局部真实性线索,而此类交互本质上是高阶的,成对图结构不足以捕捉多向跨模态依赖,超图则能很好地表示这些关系。我们提出HyperClaim,一种用于样本级真实性分类的判别式时间超图框架:使用标题或基准提供的配对文本作为类查询的主张,HyperClaim在查询标记、证据标记和采样帧上构建稀疏异构超图;应用感知置信度的过滤和源预算法形成紧凑的文本-帧及短时间证据单元;执行带残差文本-视频校准的自适应软关联推理;通过感知差异的读出聚合文本、视觉和超边状态。HyperClaim不依赖生成的理由或外部工具调用,保留了全局融合易扁平化的细粒度跨模态和时间结构。在FactGuard时间协议下,它在FakeSV、FakeTT和FakeVV上分别达到83.7%、82.0%和87.3%的准确率,优于强大的判别式和推理式基线,学习到的关联和注意力权重进一步揭示了标记和帧级结构。
英文摘要
Video misinformation detection is often approached through global multimodal fusion or free-form multimodal reasoning. Both paradigms can under-represent localized authenticity cues that arise from coupled interactions among query phrases, contextual text, and short temporal spans of frames. Because such interactions are inherently higher-order, pairwise graph formulations are insufficient to capture multi-way cross-modal dependencies, whereas hypergraphs offer a suitable representation for these relations. We propose HyperClaim, a discriminative temporal hypergraph framework for sample-level authenticity classification. Using the title or benchmark-provided paired text as a claim-like query, HyperClaim constructs a sparse heterogeneous hypergraph over query tokens, evidence tokens, and sampled frames; applies confidence-aware filtering and source budgeting to form compact text-frame and short-range temporal evidence units; performs adaptive soft-incidence reasoning with residual text-video calibration; and aggregates textual, visual, and hyperedge states through a discrepancy-aware readout. Without relying on generated rationales or external tool calls, HyperClaim preserves fine-grained cross-modal and temporal structure that global fusion tends to flatten. Under the FactGuard temporal protocol, it achieves 83.7%, 82.0%, and 87.3% accuracy on FakeSV, FakeTT, and FakeVV, respectively, outperforming strong discriminative and reasoning-centric baselines. Learned incidence and attention weights further reveal token- and frame-level structure.
Comments13 pages, including supplementary material