arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17490cs.CV

当更多基础模型意味着更少时:诊断与解决多视图融合失效问题

When More Foundation Models Means Less: Diagnosing and Addressing Multi-View Fusion Failure

Yibo Liu, Bowen Jiang

首次发表
浏览论文内容

中文总结 AI 辅助

针对多视图融合中融合编码器数量与性能非单调的问题,提出KAGES方法选择任务对齐的紧凑视图集,在多场景下提升AULC,优于DPP等选择方法。

中文摘要 AI 辅助

基础模型中心将多视图融合转化为选择问题:从大量异构编码器池中,应选择哪些视图进行融合,以及选择多少个?我们发现,下游任务性能与融合编码器的数量呈非单调关系;后续视图可能存在冗余或与任务不匹配,导致准确率饱和或下降。我们将此场景形式化为视图集组合问题,并提出KAGES(核对齐贪心编码器选择器),这是一种感知标签的方法,通过冻结编码器在中心化核目标对齐中的边际增益对其进行排序。KAGES在选择过程中无需下游分类器训练,评估每个候选的时间复杂度为O(n²),且与编码器维度无关;在单调性和正子模性比率的条件下,可获得条件性(1-e^(-γ))前缀保证。在五种识别场景及少样本、大池、全数据协议下,KAGES较全融合分别提升平均AULC 3.9、5.8和3.3个百分点,且在平均AULC上优于DPP和设施选址选择方法。图像检索在KAGES排序下呈现出更晚的、依赖任务的饱和现象,且在冻结LLM融合中重现了“先升后降”的规律。这些结果表明,有效的大池融合依赖于选择紧凑、与任务对齐的视图集,而非不加区分地融合更多编码器。

英文摘要

Foundation-model hubs turn multi-view fusion into a selection problem: from a large heterogeneous encoder pool, which views should be fused, and how many? We show that downstream performance is non-monotonic in the number of fused encoders; later views can be redundant or task-misaligned, causing accuracy to saturate or decline. We formalise this setting as view-set composition and propose KAGES (Kernel-Alignment Greedy Encoder Selector), a label-aware method that orders frozen encoders by their marginal gain in centred kernel-target alignment. KAGES requires no downstream classifier training during selection, evaluates each candidate in $\mathcal{O}(n^2)$ time independent of encoder dimension, and admits a conditional $(1-e^{-γ})$ prefix-wise guarantee under monotonicity and a positive submodularity ratio. Across five recognition regimes and low-shot, larger-pool, and full-data protocols, KAGES improves average AULC over full fusion by 3.9, 5.8, and 3.3 points, respectively, and exceeds DPP and facility-location selection in average AULC. Image retrieval exhibits later, task-dependent saturation along the KAGES ordering, while peak-then-decline reproduces in frozen-LLM fusion. These results show that effective large-pool fusion depends on selecting a compact, task-aligned set of views rather than indiscriminately fusing more encoders.

发表机构

  • Beijing University of Posts and Telecommunications(北京邮电大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑