arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12099cs.CV

超越Argmax:冻结基础模型组合中语义保留的机制研究——面向广义少样本3D分割

Beyond Argmax: A Mechanistic Study of Semantic Retention in Frozen Foundation-Model Composition for Generalized Few-Shot 3D Segmentation

Silas Kwabla Gah, Ebenezer Owusu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过冻结模型组合的受控干预实验,证明保留多种语义替代方案而非仅取argmax,可显著提升广义少样本3D分割性能,并诊断出过早语义崩溃是关键信息瓶颈。

中文摘要 AI 辅助

经典的分类器组合研究区分了分数级融合与硬决策级投票。我们重新审视这一区分,其中独立预训练的冻结基础模型在推理时被组合用于广义少样本3D分割。我们提出疑问:当异构来源在交互之前被压缩为单一类别时,有多少有用的语义信息会丢失?我们通过一种同输入语义保留干预来回答这一问题。密集的RegionPLC和稀疏的跨视图SAM3证据、模型权重、掩码、几何、词汇表和融合规则均被冻结;仅在交互之前保留的语义替代数量通过匹配的top-k阶梯进行变化。在156个保留的ScanNet200场景上,top-1达到28.47调和平均(HM)IoU,而全分布融合达到34.87 HM(+6.40,95%置信区间[+5.24,+7.64])。该模式在50个ScanNet++场景上重复出现:23.02对比26.50 HM(+3.48,95%置信区间[+1.64,+5.93])。结论是稳健的:全分布HM在稀疏来源权重0.3–0.7范围内保持稳定;替代算子(最大值、几何池化)也优于top-1;并且GroundingDINO–SAM2.1来源替换诊断显示,随着完全保留,HM从14.77单调增加到18.75。校准诊断揭示了两个来源的相反误校准,但校正校准并未消除保留优势。跨数据集和来源堆栈,通过保留一组紧凑的合理替代方案可恢复大部分信息。贡献在于对过早语义崩溃作为异构冻结模型组合中可重复信息瓶颈的受控诊断。

英文摘要

Classical classifier-combination work distinguishes score-level fusion from hard decision-level voting. We revisit this distinction where independently pretrained, frozen foundation models are composed at inference time for generalized few-shot 3D segmentation. We ask: how much useful semantic information is lost when heterogeneous sources are collapsed to a single class before they can interact? We answer with a same-input semantic-retention intervention. Dense RegionPLC and sparse cross-view SAM3 evidence, model weights, masks, geometry, vocabularies, and fusion rules are frozen; only the number of semantic alternatives retained before interaction is varied via a matched top-k ladder. On 156 held-out ScanNet200 scenes, top-1 reaches 28.47 harmonic-mean (HM) IoU while full distribution fusion reaches 34.87 HM (+6.40, 95% CI [+5.24,+7.64]). The pattern replicates on 50 ScanNet++ scenes: 23.02 vs. 26.50 HM (+3.48, 95% CI [+1.64,+5.93]). The conclusion is robust: full-distribution HM is stable across sparse-source weights 0.3--0.7; alternative operators (max, geometric pooling) also outperform top-1; and a GroundingDINO--SAM2.1 source-replacement diagnostic shows monotonic HM increase from 14.77 to 18.75 with full retention. Calibration diagnostics reveal opposite miscalibration of the two sources, yet correcting calibration does not eliminate the retention advantage. Across datasets and source stacks, most information is recovered by retaining a compact set of plausible alternatives. The contribution is a controlled diagnosis of premature semantic collapse as a repeatable information bottleneck in heterogeneous frozen-model composition.

发表机构

  • University of Ghana, Legon(加纳大学莱贡校区)

机构由 AI 辅助整理,请以论文原文为准。

↑