arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.01046cs.CRcs.AI

HiveTraceGuard-Pro:一款用于提示注入、越狱和对抗混淆的轻量型生成式护栏

HiveTraceGuard-Pro: A Compact Generative Guardrail for Prompt Injection, Jailbreaks, and Adversarial Obfuscation

Nikita Oblakov, Sabrina Sadiekh, Evgeniy Kokuykin

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出HiveTraceGuard-Pro,一款基于Qwen3-0.6B微调的0.6B参数生成式护栏,在俄语提示注入等任务中表现优异,延迟最低,相关权重已开源。

中文摘要 AI 辅助

生产环境中的大语言模型(LLMs)必须处理试图覆盖系统指令、绕过安全策略或诱导有害响应的输入。常见的缓解措施是使用独立的护栏模型,但现有研究几乎未提供关于俄语提示注入或俄语表面混淆的证据。本文提出HiveTraceGuard-Pro,这是一款基于Qwen3-0.6B进行LoRA微调的0.6B参数生成式护栏。该模型在俄语和英语语料上训练,对最终目标轮次采用二元评分规则(安全/不安全)。其训练语料库将存在对应良性示例的有害示例,与同一领域的良性示例配对,并对两类标签均应用8种混淆变换。在某测试框架中,作者将HiveTraceGuard-Pro与其他34款护栏在19组基准上进行对比,其中16组为公开基准。HiveTraceGuard-Pro的聚合关键指标为0.7432,低于两款得分更高的护栏的0.7641和0.7552;仅在16组公开基准上,其关键指标为0.7153,另有4款其他护栏得分高于它。在15款模型的对比中,HiveTraceGuard-Pro在自有团队构建的俄语数据集上,取得了最高的干净俄语鲁棒性综合F1值(0.88)和俄语提示注入召回率(0.999),且至少27.1%的提示注入数据集与训练语料库重叠;其14.3毫秒的中位延迟是该次运行中15款模型里最低的。在整个测试套件中,其误报率(FPR)为0.268,漏报率(FNR)为0.156。所有报告的响应结果均采用遗留的独立回复序列化方式,而非已发布聊天模板的自然助手角色路径。作者在Hugging Face上以Apache-2.0许可发布了合并后的模型权重,语料库、评估集和评估代码仍为内部使用。

英文摘要

Production LLMs must handle inputs that attempt to override system instructions, bypass safety policies or elicit harmful responses. A common mitigation is a separate guardrail model. Existing reports, however, provide little evidence on Russian prompt injection or Russian surface obfuscation. We present HiveTraceGuard-Pro, a 0.6B generative guardrail LoRA-tuned from Qwen3-0.6B. It is trained on Russian and English and uses one binary scoring rule (safe/unsafe) for the final target turn. Its training corpus pairs harmful examples, where a counterpart exists, with benign examples from the same domain and applies eight obfuscation transforms to both labels. In one harness, we compare HiveTraceGuard-Pro with thirty-four other guards on nineteen benchmark groups, sixteen of which are public. Its aggregate key is 0.7432, behind 0.7641 and 0.7552 for the two higher-scoring guards. Over the sixteen public groups alone, its key is 0.7153 and four of the thirty-four other suite guards score higher. In a fifteen-model comparison, HiveTraceGuard-Pro has the highest clean Russian robustness combined-F1 (0.88) and Russian prompt-injection recall (0.999). Both results use Russian sets assembled by our team, and at least 27.1% of the prompt-injection set overlaps the training corpus. Its 14.3 ms median latency is the lowest among those fifteen models in that run. Across the suite, FPR is 0.268 and FNR is 0.156. All reported response results use a legacy standalone-reply serialization rather than the natural assistant-role path of the shipped chat template. We release the merged weights on Hugging Face under Apache-2.0. The corpus, evaluation sets and evaluation code remain internal.

发表机构

  • HiveTraceLab(蜂巢追踪实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑