发表机构
Virginia Tech(弗吉尼亚理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对大视觉语言模型的事实性问题,提出无训练的IntroConformal共形风险控制框架,利用模型内省信号实现事实性保障,减少弃权情况且性能优于外部验证器基线。
AI 中文摘要
大视觉语言模型(LVLMs)已取得出色的多模态性能,但确保生成内容的事实正确性仍具挑战性。现有能为事实性提供统计保障的方法通常依赖外部验证器或生成时的置信度信号,这会引入辅助依赖,且对于置信度高但不正确的输出往往失效。我们认为可靠的事实性控制可通过模型自身衍生的内省信号实现。我们提出IntroConformal,这是一个无训练的共形风险控制(CRC)框架,可提供有限样本、无分布的事实性保障。我们首先用分层语义稳定性(一种从隐藏状态表示衍生的一致性分数)实例化该框架,随后提出验证概率,这是一种更强的分数,能捕捉模型对主张事实性的自我判断。在多个LVLM架构上,IntroConformal满足共形风险保障,同时大幅减少弃权(不执行)情况,且相对于基于外部验证器的基线,在主张级判别上达到或优于性能。
英文摘要
Large Vision-Language Models (LVLMs) have achieved strong multimodal performance, yet ensuring the factual correctness of generated content remains challenging. Existing methods that provide statistical guarantees on factuality typically rely on external verifiers or generation-time confidence signals, which introduce auxiliary dependencies or often fail for confident but incorrect outputs. We argue that reliable factuality control can instead be achieved through introspective signals derived from the model itself. We introduce IntroConformal, a training-free Conformal Risk Control (CRC) framework that provides finite-sample, distribution-free factuality guarantees. We first instantiate it with layer-wise semantic stability, a conformity score derived from hidden-state representations, and then propose verification probability, a stronger score capturing the model's self-administered judgment on claim factuality. Across multiple LVLM architectures, IntroConformal satisfies the conformal risk guarantee while substantially reducing abstention and achieving competitive or superior claim-level discrimination relative to external verifier-based baselines.
CommentsEMNLP 2026 main conference