AI 中文总结
本文通过对三类SLM架构、三个领域、两类训练数据及四种微调策略的216组实验,量化了领域适应的可信性代价,发现对抗扰动训练数据可提升适应质量且不降低可信性,部分安全策略反而增加对抗伤害易感性。
AI 中文摘要
小语言模型(Small Language Models, SLMs)的领域适应已成为在医疗、法律服务、金融分析等资源受限且高风险环境中部署强大自然语言处理系统的实用策略。尽管参数高效微调带来的性能提升已得到充分表征,但其对可信性(事实校准与对抗鲁棒性)的相应影响仍知之甚少。本文开展了首个系统的跨领域、跨架构实证研究,量化了三个SLM架构(TinyLlama 1B、Gemma-2 2B、Llama 3.2 1B)、三个领域(医疗、法律、金融)、两种训练数据条件(良性与对抗扰动)及四种微调策略(基线LoRA、Safety-DPO、Dark Experience Replay、Task Arithmetic LoRA,即TA-LoRA)下领域适应的可信性代价。通过TruthfulQA MC2(事实校准)与HarmBench ASR(对抗鲁棒性)对所有216种实验配置(含三个随机种子)评估可信性,得出三个主要发现:其一,基线QLoRA领域适应在所有模型-领域组合中使TruthfulQA MC2变化极小,平均绝对差值|Delta TQA|<0.02;其二,对抗扰动训练数据始终提升领域适应质量(损失差值约为-0.040),且未恶化可信性基准;其三,三种安全保留策略均未降低对抗伤害易感性:Safety-DPO实际效果中性(平均Delta ASR<0.001),而Dark ER与TA-LoRA在安全对齐模型(Gemma-2 2B、Llama 3.2 1B)中使平均HarmBench ASR分别提升0.171与0.155,部分配置超0.45。这些结果挑战了基于重放与算术合并策略可将对齐迁移至领域适应SLMs的假设。
英文摘要
Domain adaptation of small language models (SLMs) has emerged as a practical strategy for deploying capable NLP systems in resource-constrained, high-stakes environments including healthcare, legal services, and financial analysis. While performance gains from parameter-efficient fine-tuning are well characterised, the corresponding impact on trustworthiness (factual calibration and adversarial robustness) remains poorly understood. This paper presents the first systematic cross-domain, cross-architecture empirical study quantifying the trustworthiness cost of domain adaptation across three SLM architectures (TinyLlama 1B, Gemma-2 2B, Llama 3.2 1B), three domains (healthcare, legal, finance), two training-data conditions (benign and adversarially perturbed), and four fine-tuning strategies (baseline LoRA, Safety-DPO, Dark Experience Replay, and Task Arithmetic LoRA, TA-LoRA). Trustworthiness is evaluated through TruthfulQA MC2 (factual calibration) and HarmBench ASR (adversarial robustness) across all 216 experimental configurations with three random seeds. Three principal findings emerge. First, baseline QLoRA domain adaptation produces minimal TruthfulQA MC2 change across all model-domain combinations (mean |Delta TQA| < 0.02). Second, adversarially perturbed training data consistently improves domain adaptation quality (Delta loss approximately -0.040) without worsening trustworthiness benchmarks. Third, none of the three safety-preserving strategies reduced adversarial harm susceptibility: Safety-DPO was effectively neutral (mean Delta ASR < 0.001), while Dark ER and TA-LoRA increased mean HarmBench ASR by +0.171 and +0.155 respectively in safety-aligned models (Gemma-2 2B, Llama 3.2 1B), with individual configurations exceeding +0.45. These results challenge the assumption that replay-based and arithmetic-merge strategies transfer alignment to domain-adapted SLMs.
Comments13 pages, 7 tables, 2 appendices (Reproducibility Checklist; Software and Data Availability). Code, model checkpoints, and datasets publicly available at https://github.com/rbpdlf/slm-trw