arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Semalith v1.4:一个经过校准的1.84亿参数安全分类器,在参数比Llama - Guard - 3 - 8B少44倍的情况下实现了先进的提示注入检测

Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B

Tejasvi C. Addagada

arXiv 2607.22545首次发表:更新:

AI 中文总结

研究针对金融服务等场景中大语言模型安全分类需求,提出Semalith v1.4分类器,能单步实现三轴安全分类,经训练和对比测试,在提示注入检测上表现出色且参数少,还给出不同场景下的部署建议。

AI 中文摘要

在金融服务和智能环境中部署大语言模型需要安全分类器,能同时处理提示注入、监管合规和一般危害,现有开放护栏无法在单次推理中解决这些问题。Semalith v1.4是一个1.84亿参数的DeBERTa - v3 - base分类器,能在单次前向传播中同时进行包括提示注入、一般危害和金融服务监管合规的三轴安全分类。其22类头部在联合加权损失下,由4类辅助超级类别头部训练,基于从49个公共来源挖掘的762,04行语料库,与每个留出的评估集进行SHA - 1去重,22个基准中的21个零污染(最大0.22%)。在22个留出的基准上与Llama - Guard - 3 - 8B对比,Semalith v1.4在所有提示注入评估中获胜(7/7),在18个基准中的11个获胜,参数少44倍,对208个良性智能提示的误报率为0.000,而Llama - Guard - 3 - 8B为0.063。在一般危害基准上Llama - Guard - 3领先。第6节披露了6个测量到的弱点。部署建议:对话审核部署推荐v1.3(ToxicChat F1 0.624);当优先考虑BFSI标签覆盖或良性智能提示的零误报率时,推荐v1.4。

英文摘要

Deploying large language models in financial-services and agentic settings requires safety classifiers that simultaneously handle prompt injection, regulatory compliance, and general harm, a combination no existing open guardrail addresses in a single inference pass. Semalith v1.4 is a 184M-parameter DeBERTa-v3-base classifier performing simultaneous three-axis safety classification including prompt injection, general harm, and financial-services regulatory compliance, in a single forward pass. Its 22-class head (BENIGN, nine prompt-injection sub-types, general-harm, eleven BFSI labels) is trained with a 4-class auxiliary super-category head under jointly weighted loss, on a 76,204-row corpus mined from 49 public sources with SHA-1 deduplication against every held-out evaluation set, with 21 of 22 benchmarks at zero contamination (max 0.22%). Against Llama-Guard-3-8B on 22 held-out benchmarks, Semalith v1.4 wins every prompt-injection evaluation (7/7) and 11 of 18 benchmarks overall at 44x fewer parameters, with FPR = 0.000 on 208 benign agentic prompts vs 0.063 for Llama-Guard-3-8B. On general-harm benchmarks (WildGuardMix, HEx-PHI, HarmBench), Llama-Guard-3 leads; this complementary split is documented in Section 4. Six measured weak spots are disclosed in Section 6. Deployment guidance: v1.3 is recommended for conversational moderation deployments (ToxicChat F1 0.624); v1.4 is recommended when BFSI label coverage or zero-FPR on benign agentic prompts is the priority.

Comments16 pages, 8 tables, no figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑