arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

共享编码器并非共享任务:深度专家池的条件比较

A Shared Encoder Is Not a Shared Task: Conditional Comparison for Deep Expert Pools

Kentaro Oda

arXiv 2609.27866首次发表:更新:

发表机构

Kagoshima University(鹿儿岛大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出条件双判别器嵌入差异,解决深度专家池中共享编码器的任务比较混淆,以双轴门实现更优决策质量,并在广义类别发现中达到高AUROC。

AI 中文摘要

共享深度编码器本身并不能解决任务比较分数中的核心混淆问题。我们证明,在冻结的共享表示上进行交叉评估的头会继承浅层交换分数的外推混淆:固定标签的纯输入旋转将深度交换分数从约0提升至0.80,而表示新颖性分数在互补方向上失明(在完全改变任务的标签排列下保持平坦)。将条件双判别器差异移植到嵌入空间可同时解决两个盲点:功能轴在旋转下保持在±0.001以内,并单调跟踪标签排列漂移质量。在混合头生命周期中,双轴门在匹配的训练预算下,以更少的头实现了比交换或新颖性触发器更好的决策质量。在广义类别发现中,相同的块级功能轴以AUROC 0.98-0.99区分语义新颖性与光度偏移,而逐输入OOD分数(MSP、Energy、Mahalanobis、KNN)在该区分上接近随机水平。所有发现均在冻结的ImageNet-21k ViT-B/16和自监督DINOv2骨干网络上于CIFAR-100上复制,并扩展到具有循环的残差适配器池,其中零校准的新颖性触发器在机制变化时从不触发,而双轴门以完全循环重用处理它们。我们明确陈述了嵌入空间结论迁移到原始机制的公共因子条件。

英文摘要

Sharing a deep encoder does not, by itself, fix the central confound of task-comparison scores. We show that cross-evaluated heads on a frozen shared representation inherit the extrapolation confound of shallow exchange scores: pure input rotations with fixed labels inflate a deep exchange score from about 0 to 0.80, while representation-novelty scores are blind in the complementary direction (flat under label permutations that change the task completely). Transplanting a conditional two-discriminator discrepancy into the embedding space resolves both blind spots: the functional axis stays within +-0.001 under rotations and tracks label-permutation drift mass monotonically. Built into a mixture-of-heads lifecycle, the two-axis gate attains better decision quality with fewer heads than exchange or novelty triggers at a matched training budget. On generalized category discovery, the same chunk-level functional axis separates semantic novelty from photometric shift with AUROC 0.98-0.99 where per-input OOD scores (MSP, Energy, Mahalanobis, KNN) sit near chance for that distinction. All findings replicate across frozen ImageNet-21k ViT-B/16 and self-supervised DINOv2 backbones on CIFAR-100, and extend to residual adapter pools with recurrence, where a null-calibrated novelty trigger never fires on mechanism changes while the two-axis gate handles them with full recurrence reuse. We state explicitly the common-factoring condition under which embedding-space conclusions transfer to the original mechanism.

Comments6 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑