arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

查询条件原型自适应用于跨域少样本学习:单查询推理、受控比较与失败模式

Query-Conditioned Prototype Adaptation for Cross-Domain Few-Shot Learning: Single-Query Inference, Controlled Comparisons, and Failure Modes

Rushab Rasik Karania, Tomas Maul

arXiv 2609.30769首次发表:更新:

发表机构

University of Nottingham Malaysia(诺丁汉大学马来西亚校区)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出WIPT方法,通过单查询测试时原型自适应改善跨域少样本分类,在部分目标上提升精度,但收益非普遍,且支持流式推理。

AI 中文摘要

跨域少样本学习要求在没有目标时参数更新的情况下,从极少量标注样本中将分类器适应到新的视觉域。我们聚焦一个问题:在固定全局表示下,联合查询-支持自适应对原型构建有何贡献?实例内原型Transformer(WIPT)通过联合变换一个未标注查询和标注支持嵌入,形成查询特定的类均值,实现单查询测试时原型自适应。使用共享冻结ViT-S/16编码器、miniImageNet源训练以及CUB、EuroSAT和ISIC目标,我们在五个独立训练种子上复现了关键比较。在1-shot评估中,WIPT在CUB(+0.21个百分点)和EuroSAT(+2.07)上的每次运行均优于冻结ProtoNet,但在ISIC上降低(-0.22)。在5-shot评估中,ProtoNet整体仍最强,而WIPT在ISIC上持续优于容量匹配的仅支持Transformer(+0.99)。联合处理多达五个查询未带来可靠的精度提升;在仅头部5-shot基准中,g=5相对于g=1减少了73%的分析注意力令牌对和29%的峰值分配内存,尽管延迟非单调。在所有目标/样本条件下,WIPT对不确定的ProtoNet决策的改变远多于自信决策,救援/破坏分解解释了观察到的增益和损失。源偏移和评分器控制进一步表明该收益并非普遍。总体而言,WIPT提供了一种流式兼容的测试时原型自适应形式,可在无目标时优化的情况下改善困难的低样本跨域决策。

英文摘要

Cross-domain few-shot learning requires adapting a classifier to a new visual domain from very few labelled examples without target-time parameter updates. We isolate one question: under a fixed global representation, what does joint query-support adaptation contribute to prototype construction? The Within-Instance Prototypical Transformer (WIPT) implements single-query test-time prototype adaptation by jointly transforming one unlabelled query and the labelled support embeddings, then forming query-specific class means. Using a shared frozen ViT-S/16 encoder, miniImageNet source training, and CUB, EuroSAT and ISIC targets, we replicate the key comparisons across five independent training seeds. In 1-shot evaluation, WIPT improves frozen ProtoNet in every run on CUB (+0.21 percentage points) and EuroSAT (+2.07), but decreases ISIC (-0.22). In 5-shot evaluation, ProtoNet remains strongest overall, while WIPT consistently improves a capacity-matched support-only Transformer on ISIC (+0.99). Joint processing of up to five queries yields no reliable accuracy gain; in a head-only 5-shot benchmark, g = 5 reduces analytical attention-token pairs by 73% and peak allocated memory by 29% relative to g = 1, although latency is non-monotonic. Across all target/shot conditions, WIPT changes uncertain ProtoNet decisions far more than confident ones, and rescue/break decomposition accounts for the observed gains and losses. Source-shift and scorer controls further show that the benefit is not universal. Overall, WIPT provides a streaming-compatible form of test-time prototype adaptation that can improve difficult low-shot cross-domain decisions without target-time optimization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑