arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向视频问答的目标校验可靠性分数精化

Target-Checked Reliability Score Refinement for Video Question Answering

Guoxiang Ren, Rohitash Chandra

arXiv 2609.13288首次发表:更新:

发表机构

UNSW Sydney(新南威尔士大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视频问答中模型高置信但错误的问题,提出基于响应图和目标校验的分数精化方法,在不重训模型下显著降低AURC并提升可靠性。

AI 中文摘要

视频语言模型能够以高置信度回答多项选择题,但答案可能是错误的。我们研究在目标偏移下,是否可以在不重新训练模型或改变其答案的情况下改进答案级可靠性分数。我们从三个固定的视频语言模型中,在四种确定性视频采样下收集选项概率列表,并将跨视图变化和跨模型一致性表示为响应图。利用带标签的目标试点,我们将原始分数(定义为所选答案的概率)与在开发数据集上训练的基于直方图的梯度提升(HGB)分数以及在目标试点上训练的正则化逻辑回归分数进行比较。仅当重复的视频级检查表明有积极且稳定的改进时,候选分数才替换原始分数。我们在公开的VideoQA基准和视频幻觉诊断(VHD)(一个用于共享高置信度错误的受控诊断数据集)上开发了这一规则。排序质量通过风险覆盖率曲线下面积(AURC)衡量,数值越低越好。在保留的963个问题的HERBench划分上,该方法将三个模型的平均AURC降低了16.64%(95%置信区间(CI),12.12至22.61%);最小的模型级增益为11.39%。在另一个保留的911个问题的感知测试划分上,平均降低为18.87%(95% CI,15.43至22.14%)。对于InternVL3.5,目标检查保留了原始分数。使用相同的输出,该方法在两个数据集上的平均AURC上优于七种无训练基线。它还提高了AUROC,减少了校准误差,并将50%覆盖率下的错误率降低了6.50和6.58个百分点。

英文摘要

Video-language models can answer multiple-choice questions with high confidence yet be wrong. We study whether answer-level reliability scores can be improved under target shift without retraining the models or changing their answers. We collect option-probability lists from three fixed video-language models under four deterministic video samplings and represent cross-view changes and cross-model agreement as a response graph. Using a labeled target pilot, we compare the original score, defined as the probability assigned to the chosen answer, with a histogram-based gradient-boosting (HGB) score trained on the development datasets and a regularized logistic-regression score trained on the target pilot. A candidate replaces the original score only when repeated video-level checks indicate a positive, stable improvement. We develop this rule on public VideoQA benchmarks and Video Hallucination Diagnosis (VHD), a controlled diagnostic dataset for shared high-confidence errors. Ranking quality is measured by the area under the risk-coverage curve (AURC), where lower is better. On a held-out 963-question HERBench split, the method reduces mean AURC across the three models by 16.64% (95% confidence interval (CI), 12.12 to 22.61%); the smallest model-level gain is 11.39%. On a separate held-out 911-question Perception Test split, the mean reduction is 18.87% (95% CI, 15.43 to 22.14%). For InternVL3.5, the target check retains the original scores. Using the same outputs, the method outperforms seven training-free baselines in mean AURC on both datasets. It also improves AUROC, reduces calibration error, and lowers the error rate at 50% coverage by 6.50 and 6.58 percentage points.

Comments18 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑