arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33359cs.CV

几何视图合成何时有助于酒标检索?一个跨自监督与视觉-语言骨干网络的公开单样本基准

When Does Geometric View Synthesis Help Wine Label Retrieval? A Public One-Shot Benchmark Across Self-Supervised and Vision-Language Backbones

发表机构国立东华大学
查看机构详情
  • National Dong Hwa University(国立东华大学)

机构由 AI 辅助整理,请以论文原文为准。

Yueh-Cheng Huang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究通过公开基准评估几何视图合成对酒标检索的增益,发现其对自监督DINO配方提升显著(34.1%至62.6-63.7%),但对冻结SigLIP 2-B仅带来较小改进,且SAM定位优于边缘方法。

中文摘要 AI 辅助

几何视图合成可以将单张酒标照片扩展为训练集,但其与预训练图像编码器结合的价值尚不明确。我们在一个基于WineSensed的公开基准上对此进行研究,该基准包含1,000个类别,每个类别一张注册照片,以及4,295个真实查询。使用较早的DINO视觉变换器(ViT-S/16)配方,几何视图将top-1准确率从34.1%提升至62.6-63.7%,约为二维(2D)增强所获增益的三倍。冻结的SigLIP 2-B已达到94.7%的准确率。在其冻结特征之上添加线性头,通过Segment Anything Model(SAM)定位的两种几何管线可获得1.2-1.3个百分点的增益,而其他管线仅获得不显著的0.3-0.6个百分点增益。低秩适应(LoRA)和基于验证集选择的完全微调在报告的置信区间内未显示明确增益;固定预算的完全微调则损失9-24个百分点。SAM定位为99%的源提供了全部六个视图,而基于边缘的前端仅为43%。两种圆柱体构造之间的识别差异取决于训练配方,并受其裁剪和画布约定的混淆影响。渲染圆柱体测试显示对源倾斜的不同响应,但未校准的轮缘比代理未在真实照片的识别中建立相应趋势。一项经作者确认的50个残差错误审计识别出21个查询-注册外观不匹配,但未确定不可约的错误率。这些结果支持在所测试的自监督配方中使用几何合成,并通过冻结特征适应为文本监督编码器带来较小收益。

英文摘要

Geometric view synthesis can expand a single wine-label photograph into a training set, but its value with pretrained image encoders is unclear. We study this on a public WineSensed-derived benchmark of 1,000 classes, one enrollment photograph per class, and 4,295 real queries. With the earlier DINO vision transformer (ViT-S/16) recipe, geometric views raise top-1 accuracy from 34.1% to 62.6-63.7%, about three times the gain from two-dimensional (2D) augmentation. Frozen SigLIP 2-B already reaches 94.7%. A linear head over its frozen features gains 1.2-1.3 percentage points with the two geometric pipelines localized by the Segment Anything Model (SAM), while the other pipelines gain an inconclusive 0.3-0.6 points. Low-rank adaptation (LoRA) and validation-selected full fine-tuning show no clear gain within the reported confidence intervals; fixed-budget full fine-tuning loses 9-24 points. SAM localization supplies all six views for 99% of sources, compared with 43% for the edge-based front end. Recognition differences between the two cylinder constructions depend on the training recipe and are confounded by their crop and canvas conventions. Rendered-cylinder tests show different responses to source tilt, but an uncalibrated rim-ratio proxy establishes no corresponding trend in recognition on real photographs. An author-confirmed audit of 50 residual errors identifies 21 query-enrollment appearance mismatches, without establishing an irreducible error rate. These results support geometric synthesis for the tested self-supervised recipe and a smaller benefit through frozen-feature adaptation of the text-supervised encoder.

↑