arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Visual Jev:基于共享视觉上下文的准确高效决策

Visual Jev: Accurate and Efficient Decisions from Shared Visual Context

Guanxu Yu, Yuhang Yao

arXiv 2609.25845首次发表:更新:

发表机构

Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Visual Jev 通过共享图像上下文并批量执行多个独立问题,结合答案监督后训练,将宏平均准确率从70.6%提升至76.1%,同时实现8.9倍加速,验证了简单高效的设计方案。

AI 中文摘要

许多视觉应用会对同一张图像提出多个独立的、强制选择的问题。Visual Jev 仅对图像和公共上下文编码一次,将孤立的问句后缀作为一批并行执行,并从骨干网络的语言模型头读取候选概率。在四个基准测试中,经过答案监督的后训练将等权宏平均准确率从 70.6% 提升至 76.1%,且提升主要集中在训练中涉及的两类任务上。当每张图像有 N=32 个问题时,共享批量执行的热摊销时间比独立串行执行快 8.9 倍,并且比已经批量化的、需重新计算前缀的基线仍快 3.4 倍,但代价是更高的峰值内存。一个匹配的带类型头的对照组在准确率上相比语言模型头读取方式并无一致优势。因此,所支持的设计很简单:调整骨干网络以提升质量,保留现有读取头,并通过共享执行来提高效率。

英文摘要

Many vision applications ask several independent, forced-choice questions about the same image. Visual Jev encodes the image and public context once, executes isolated question suffixes as a batch, and reads candidate probabilities from the backbone's language-model head. Across four benchmarks, answer-supervised post-training raises equal-weight macro accuracy from 70.6% to 76.1%, with the gain concentrated on the two task families represented in training. At N=32 questions per image, shared batched execution is 8.9x faster in warm amortized time than independent serial execution and remains 3.4x faster than an already-batched baseline that recomputes the prefix, at the cost of higher peak memory. A matched typed-head control offers no consistent accuracy advantage over the language-model-head readout. The supported design is therefore simple: adapt the backbone for quality, retain the existing readout, and share execution for efficiency.

CommentsCode: https://github.com/guanxuyu-sv/Visual-Jev

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑