arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在6GB 2011 GPU上的现代多模态助手:针对费米架构的阶段验证、全GPU CUDA推理

A Modern Multimodal Assistant on a 6 GB 2011 GPU: Stage-Validated, All-GPU CUDA Inference for Fermi

A. C. Opus, J. Q. Lu

arXiv 2607.14568首次发表:更新:

发表机构

Department of Physics, University of Puerto Rico(物理系,波多黎各大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究在6GB 2011费米GPU上部署MiniCPM-V-4.6多模态助手,通过构建全GPU引擎、移植验证视觉部分、优化长上下文处理等方法,实现了高效的推理,能在1.7秒内端到端回答图像问题。

AI 中文摘要

一项相关研究在2011年的NVIDIA Tesla C2075(费米,sm_20,6GB)GPU上运行了一个35B专家混合模型,采用GPU预填充/CPU解码混合模式,但由于4位模型不适合设备内存(arXiv:2606.24031)。本报告保持硬件不变,探讨适合该硬件的模型能实现什么:我们在GPU上完全部署了MiniCPM-V-4.6,这是一个现代多模态助手,它将SigLIP2视觉编码器和窗口注意力合并器(16倍视觉令牌压缩)与紧凑的混合门控增量网络主干相结合。研究有三个结果。一是基于实测构建全GPU引擎,包括对8位权重的去量化投影等;二是视觉部分的移植及验证,发现位置嵌入桶化在精确有理关系上存在差异;三是长上下文暴露了短基准测试隐藏的O(N^2)墙,通过特定重写优化了预填充和图像编码等。系统能在1.7秒内端到端回答图像问题。

英文摘要

A companion study ran a 35B mixture-of-experts model on a 2011 NVIDIA Tesla C2075 (Fermi, sm_20, 6GB) as a GPU-prefill/CPU-decode hybrid, because the 4-bit model did not fit in device memory (arXiv:2606.24031). This report keeps the hardware and asks what a model that fits can do: we deploy MiniCPM-V-4.6, a modern multimodal assistant pairing a SigLIP2 vision encoder and window-attention merger (16x visual token compression) with a compact hybrid gated-delta-net backbone, entirely on the GPU. Three results. (i) An all-GPU engine built on measured foundations: projections that dequantize 8-bit weights once and call the vendor SGEMM still in the last Fermi toolchain (64% of FP32 peak; our best hand-written GEMM hit 37%, wrongly called the ceiling); a chunked delta-rule rewrite of the recurrent layers, 2.8x faster than the sequential scan once attribution exposed one bad kernel; and a measured negative: 4-bit weights make decode slower than 8-bit here, since Fermi issues nibble-unpacking shifts at half rate. (ii) The vision side is a port with a proof obligation: we translate tower, merger, and projector to sm_20 CUDA, validating every stage against a locally generated reference forward (full tower 1.4e-5). One failure, position-embedding bucketization differing on exact rational ties, generalizes to a rule: float tie-breaking in index arithmetic is implementation-defined; call the reference operator, do not reimplement it. (iii) Long context exposes an O(N^2) wall short benchmarks hide: prefill falls from 114 tok/s at 2k tokens to 21 at 10k in a naive attention kernel; per-head vendor-GEMM calls writing into the existing score buffer (zero extra memory) restore a flat profile (408 at 2k, 361 at 10k; 17x), verified by exact needle retrieval from 60% depth. The same rewrite cuts image encoding 6x, to 0.93s. The system answers an image question end-to-end in 1.7s.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑