arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33450cs.CV

原生关联:基于基础视觉语言模型的野外置信度感知人类感知

Native Association: Confidence-Aware Human Perception in the Wild with a Foundation VLM

Igal Dmitriev, Ofir Liba

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出原生关联方法,通过微调0.77B的Florence-2模型,以语法约束序列直接输出球员属性,消除身份互换,相比零样本API将错误关联降低4倍,并实现检测F1达0.95,支持置信度拒绝与缩放重读,表明增益源于结构而非权重。

中文摘要 AI 辅助

从广播画面中提取谁在何处、属于哪支队伍、身穿哪个号码,通常是通过拼接检测器、OCR引擎和分类器来实现的——而拼接步骤在遮挡情况下会导致身份互换。我们转而让关联成为原生操作:一个0.77B参数的视觉语言模型(Florence-2)经过微调,以语法约束的序列形式输出每个人的所有属性,每个属性在其所属者的块内生成。因此,输出在每一帧上天然符合模式,且不存在事后绑定步骤将正确读取的号码附加到错误的球员身上:残余的错误关联纯粹是感知错误,其发生概率比零样本提示的前沿API低约4倍(0.057对比0.21-0.24)。在冻结的多运动测试集上,这一单次通过达到了0.95的检测F1分数(API为0.65-0.75)。一次额外的前向传递即可获得每个字段的置信度,支持拒绝选项(在50%覆盖率下,球衣精度从0.71提升至0.96),并为小球员路由无训练的缩放重读。令人惊讶的是,一旦语法被学习,在我们测试的适应配置下,进一步的参数高效调优并未带来可衡量的增益;相同的配方在WIDER-Attribute上达到了93.1的给定框mAP,在其标准测试协议下首次产生了检测耦合的端到端结果(84.5 mAP),并复现了相同的调优结果。在这种机制下,增益存在于结构中,而非增加的权重中。

英文摘要

Extracting who is where, on which team, wearing which number from a broadcast frame is typically done by stitching a detector, an OCR engine, and classifiers together -- and the stitching step swaps identities under occlusion. We make association native instead: a 0.77B vision-language model (Florence-2) is fine-tuned to emit all per-person attributes as one grammar-constrained sequence, with each attribute generated inside its owner's block. Output is therefore schema-valid on every frame by construction, and no post-hoc binding step exists to attach a correctly read number to the wrong player: residual misassociation is pure perception error, $\approx4\times$ rarer than zero-shot-prompted frontier APIs' (0.057 vs. 0.21-0.24). On a frozen multi-sport test set, this single pass reaches 0.95 detection F1 (APIs: 0.65-0.75). A single extra forward pass yields a per-field confidence that supports a reject option (jersey precision $0.71\rightarrow0.96$ at half coverage) and routes a training-free zoom-and-re-read for small players. Surprisingly, once the grammar is learned, further parameter-efficient tuning yields no measurable gain under the adaptation configurations we test; the identical recipe on WIDER-Attribute reaches 93.1 mAP given-box, yields the first detection-coupled end-to-end results under its standard test protocol (84.5 mAP), and reproduces the same tuning result. In this regime, the gains live in the structure, not in added weights.

发表机构

  • WSC Sports(WSC体育公司(或WSC Sports,按原文保留))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑