arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BanglaVeilGuard:孟加拉语大语言模型的跨脚本安全基准测试与轻量安全防护措施

BanglaVeilGuard: Cross-Script Safety Benchmarking and Lightweight Guardrails for Bangla Large Language Models

Md. Rakibul Hassan, Muhammad Iqbal Hossain

arXiv 2608.21880首次发表:更新:

发表机构

BRAC University(BRAC大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对孟加拉语LLM跨脚本安全评估难题,本文提出BanglaVeilGuard基准与轻量提示防护,可降低攻击成功率,提升不安全请求召回率,但存在方言及带噪良性提示过度拒绝的问题。

AI 中文摘要

以英语为中心或标准脚本的基准难以评估孟加拉语大语言模型(LLM)的安全性,因为孟加拉语用户经常混用多种脚本、拼写、代码混合形式及区域语域。本文提出了BanglaVeilGuard,这是一个以孟加拉语为核心的紧凑安全基准,以及针对六种语言形式的轻量提示防护措施:标准孟加拉语、罗马化孟加拉语、孟加拉语-英语混合语(Banglish)、孟加拉语-英语代码混合语、带噪孟加拉语及方言孟加拉语。该基准包含2366条经质量筛选的提示,以及一个保留的354条提示的评估集,涵盖不安全、安全及安全敏感请求。BanglaVeilGuard采用非破坏性多视图归一化,结合提示风险分类器和阈值化预生成门,无需修改目标模型权重即可对异构目标模型的提示进行筛选。在目标模型系列中,采用防护措施后,确定性响应评分下的攻击成功率从93.8%-100.0%降至Claude Opus 4.8、BanglaLLama和TituLLM的6.3%;配备BanglaVeilGuard的TigerLLM-1B准确率达78.2%,攻击成功率(ASR)为8.8%。该提示防护还实现了88.5%的不安全请求召回率,显著高于所评估的仅提示防护基线。主要剩余代价是对方言和带噪良性提示的过度拒绝,这揭示了孟加拉语LLM部署中具体的安全-有用性边界。

英文摘要

Bangla large language model (LLM) safety is difficult to evaluate with English-centric or standard-script benchmarks because Bangla users routinely write across scripts, spellings, code-mixed forms, and regional registers. This paper presents BanglaVeilGuard, a compact Bangla-first safety benchmark and lightweight prompt guard for six language forms: standard Bangla, Romanized Bangla, Banglish, code-mixed Bangla--English, noisy Bangla, and dialectal Bangla. The benchmark contains 2,366 quality-filtered prompts and a held-out 354-prompt evaluation split spanning unsafe, safe, and safe-sensitive requests. BanglaVeilGuard uses non-destructive multi-view normalization with a prompt-risk classifier and thresholded pre-generation gate, allowing it to screen prompts for heterogeneous target models without changing their weights. Across target-model families, guarded runs reduce attack success under deterministic response scoring from 93.8--100.0\% to 6.3\% for Claude Opus 4.8, BanglaLLama, and TituLLM; TigerLLM-1B with BanglaVeilGuard achieves 78.2\% accuracy with 8.8\% ASR. The prompt guard also attains 88.5\% unsafe recall, substantially above the evaluated prompt-only guard baselines. The main remaining cost is over-refusal on dialectal and noisy benign prompts, revealing a concrete safety-helpfulness frontier for Bangla LLM deployment.

CommentsAccepted at the 4th International Conference on Computing Advancements (ICCA 2026). 8 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑