arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VectraYX-Vision-1B:一款具备结构化视觉推理与原生工具使用能力的20亿参数以下西班牙语/拉丁美洲网络安全多模态模型

VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model, and What Limits Its Visual Grounding

Juan S. Santillana

arXiv 2608.08477首次发表:更新:

发表机构

Globant(高博特科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出VectraYX-Vision-1B,一款20亿参数以下西班牙语/拉丁美洲网络安全多模态模型,支持结构化推理与工具调用,针对网络UI优化,同时报告了其视觉定位的负面结果及相关修复方案,开源了模型与数据。

AI 中文摘要

我们提出了VectraYX-Vision-1B,这是一款针对西班牙语/拉丁美洲网络安全图像的20亿参数以下多模态模型(VLM),通过多层感知器(MLP)将冻结的SigLIP-so400m编码器与10.4亿参数的西班牙语/拉丁美洲安全解码器相连。据我们所知,它是首款针对网络用户界面(包括IDA、Ghidra、Wireshark、Nmap、Metasploit、Volatility)的20亿参数以下多模态模型,支持西班牙语回答,通过原生<|think|>标记输出结构化推理,通过模型上下文协议(Model Context Protocol,<|tool_call|>)调用工具,并可导出至该http URL的LLaVA mmproj格式以实现空气隔离部署。我们报告了一项负面的初步视觉定位结果:尽管管道功能完整,但当前视觉监督微调(SFT)步数为400至1900步,约1600万token)产生的B6分数接近零(工具识别得分为0.08),忽略了图像内容。我们明确了修复方案(更长的SFT、≥60%的重放、更低的学习率),并发现了一个检查点加载器错误(未剥离的llm.前缀),该错误伪装成训练崩溃。关键的是,我们引入了一个包含3个变体的 ablation矩阵(V0:每4层无位置编码(NoPE)、V1:全旋转位置编码(RoPE)、V2:无位置编码+学习到的2D编码),以研究在729 token的视觉块上,周期性无位置编码层对注意力的影响。代码、配置和权重已发布,以确立该架构问题的优先权。我们提供了文本主干的B1-B5分数、文本控制、初步B6/B7分数、 wall时间、CPU上的GGUF效率,以及涵盖10个领域的14596个问答对的语料库。我们开源了所有模型和轨迹:jsantillana/vectrayx-1b、jsantillana/vectrayx-vision-1b和jsantillana/vectrayx-vision-1b-checks。

英文摘要

We build VectraYX-Vision-1B, a sub-2B Spanish/LATAM cybersecurity vision-language model for offline use, and measure what limits its visual grounding. A frozen 1.04B-parameter decoder is coupled to a frozen vision encoder through a trainable projector: first SigLIP, then the Qwen2-VL-2B tower with a projector shaped for llama.cpp's mmproj export. With SigLIP, after five fine-tuning defects were repaired, a nine-field extraction gate with a shuffled-image control passes the same 2/9 fields under every configuration that keeps the encoder and adds no text hint. A frozen-feature probe explains why. Transplanting the Qwen2-VL tower with its own merger reads an 8-nibble address SigLIP never read (0.00 to 0.81 exact match). On B8 (2,040 items, 34 fields, 16 templates, shuffled-image and best-constant controls), 9 fields pass on a Qwen-tower checkpoint, all within trained template-field pairs. Screenshot tool identification (B6) sat at exactly 0.000. That was not a perception ceiling. Stock Qwen2-VL-2B transcribes the same images at word recall 0.93 at our pixel budget (0.52 with our letterboxing). Training only the projector on B6's task format and render family lifts B6 tool identification to 0.96 and recall to 0.48-0.51. Perception of rendered text is therefore present and trainable in this frozen-backbone design. The evidence is narrow. B6 is nearly in-distribution for that projector, which also forgets part of B8 (mean field accuracy 0.53 to 0.29 on a retention check); runs are single-seed, and no checkpoint yet has both. We document four harness defects, retract an earlier B6 score, and release code, benchmarks and checkpoints.

Comments32 pages, 1 figure, 14 tables. v4 adds B8 (2,040 items, 34 fields, 16 templates, dual control; 9/34 pass) and a diagnosis of the B6 tool-identification floor: not a perception ceiling; projector-only training lifts B6 to 0.96 on its own render family, at a B8 retention cost. Title changed. Code, benchmarks, checkpoints on HF

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑