发表机构
National Research Council Canada(加拿大国家研究委员会)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对音频、视频、文本冲突场景,提出阈值策略等方法,在三得分基准实验中验证了其决策准确率与效用,揭示选择价值受预算和成对排名影响。
AI 中文摘要
当音频、视频和文本信息不一致时,仅靠准确率无法判断是否需要获取另一模态的信息或弃权(不执行)。我们在可控三得分基准中研究这些选择:一个策略观察两个带符号得分,可请求第三个得分并产生一定代价,也可选择弃权(不执行)。主要奖励机制与特定歧义机制相关:仅针对指定的歧义机制时弃权(不执行)是正确的,在混合损坏场景下弃权(不执行)会受到惩罚。匹配控制实验显示,阈值策略与始终请求决策的请求次数更少;其相比始终回答融合策略的优势取决于分配给该歧义的奖励。在部分保留的合成划分上,阈值策略在83个随机种子下达到0.789±0.006的目标决策准确率和0.481±0.014的效用。三得分多数参考策略达到0.626±0.008的目标决策准确率和0.252±0.016的效用,但使用了更多信息。在匹配预算测试中,仅训练的值选择器在10%和25%预算下,相比无查询和匹配随机策略提升了效用,而成对不确定性在所有预算下都有更高的效用。在50%和63.7%预算下,尽管阈值策略的非歧义准确率略高,但值选择器的效用降低。若所有弃权(不执行)都被判定为错误,多数策略的效用超过阈值策略。在中心时间设置下,全轨迹控制与神经模型匹配,而位置扰动可区分它们。在保留的演员情绪片段上,八帧融合对两个编码器对呈现出相反符号的准确率差异,且两个演员区间均包含零;匹配请求路由的增益较小且不确定。这些结果将全模态准确率与请求前选择价值区分开,表明选择价值取决于预算和观察到的成对排名。
英文摘要
When audio, video, and text disagree, accuracy alone does not show whether to acquire another source or abstain. We study these choices in a controlled three-score benchmark: a policy observes two signed scores, may request the third at a cost, and can abstain. The primary reward is mechanism-specific: abstention is correct only for one designated ambiguity mechanism and is penalized under mixed corruption. Matched controls show that a threshold policy matches always-request decisions with fewer requests; its advantage over always-answer fusion depends on the reward assigned to that ambiguity. On a partially held-out synthetic split, the threshold policy reaches 0.789 +/- 0.006 targeted decision accuracy and 0.481 +/- 0.014 utility across 83 seeds. A three-score majority reference reaches 0.626 +/- 0.008 and 0.252 +/- 0.016, but uses more information. In a matched-budget test, a train-only value selector improves utility over no-query and matched-random policies at 10% and 25% budgets, while pair uncertainty has higher utility at every budget. At 50% and 63.7% budgets, the selector lowers utility despite slightly higher non-ambiguous accuracy. If all abstentions are scored incorrect, majority outranks the threshold policy in utility. At a central temporal setting, full-trace controls match the neural models while position perturbations separate them. On held-out-actor emotion clips, eight-frame fusion has opposite-signed accuracy differences for two encoder pairs, with both actor intervals containing zero; matched-request routing gains are small and uncertain. These results separate full-modality accuracy from pre-request selection value and show that selection value depends on budget and the observed-pair ranking.