发表机构
Duke University(杜克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究指出视频多项选择评估中帧相位偏移导致答案不稳定,提出PHASEFUSION方法融合多相网格以稳定答案,并建议报告相位约定或边缘化处理。
AI 中文摘要
视频语言模型通过在统一网格的帧上进行多项选择准确率来排名。该网格有两个参数:速率和相位,而基准测试仅报告速率。相位会改变答案:两个仅相差半步相位偏移的已部署采样器,在23.6%的问题上给出不同答案,同时得分仅相差一分;在来自两个系列的四个版本中,在控制选项顺序后,仅改变相位就会改变约五分之一的答案。PHASEFUSION解码三个偏移网格并平均选项后验概率。这些网格是密集网格的多相分量。融合在预定义余量内(对数得分)达到与32帧单次通过相当的准确率,并将三个网格全部半步偏移所改变的答案从18.2%降至10.1%。选项顺序仅改变呈现方式,通过单次通过的答案余量来标记。在预算中报告相位约定,或将其边缘化。
英文摘要
Video-language models are ranked by multiple-choice accuracy on frames from a uniform grid. The grid has two parameters, a rate and a phase, and benchmarks report only the rate. The phase moves answers: two deployed samplers differing only by a half-step phase offset answer 23.6% of questions differently while scoring within a point, and across four releases from two families shifting only the phase changes roughly one answer in five after controlling option order. PHASEFUSION decodes three offset grids and averages the option posteriors. The grids are the polyphase components of the dense grid. Fusion matches a 32-frame single pass in accuracy within a prespecified margin (logit-scored) and cuts the answers a half-step shift of all three grids changes from 18.2% to 10.1%. Option order, which changes only the presentation, is flagged instead by a one-pass answer margin. Report the phase convention with the budget, or marginalize it.