arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LANTERN:面向噪声与变换任务的语言模型评估,用于理解误差与鲁棒性细微差异

LANTERN: Language Model Assessment on Noisy and Transformed Tasks for Understanding Error and Robustness Nuances

Vamsi Krishna Kodavali, Rituraj Singh

arXiv 2609.07309首次发表:更新:

发表机构

Samsung(三星)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究构建合成增强数据集,系统评估不同规模及微调变体LLM在词错误率、字符重复、选项修改和指令遵循等扰动下的鲁棒性,揭示响应变异性,为开发更稳健模型和评估方法提供见解。

AI 中文摘要

大型语言模型(LLM)的鲁棒性评估仍然是一个关键挑战,尤其是在评估其对输入数据扰动的敏感性方面。在本工作中,我们系统地评估了LLM在多个维度上的鲁棒性,包括词错误率、字符重复与复制、选项修改以及指令遵循的变异性。为促进这一评估,我们构建了一个合成且增强的数据集,涵盖多种LLM基准,特别针对多项选择题(MCQ)数据集和指令遵循任务。我们对不同规模(小、中、大)的LLM以及基础版本和指令微调版本进行了广泛实验。我们的分析量化了在扰动条件下模型响应的变异性,并突出了与基线模型相比的差异。研究结果为LLM在不同评估场景下的稳定性提供了见解,有助于开发更鲁棒和可靠的语言模型以及稳健的评估方法。

英文摘要

Robustness evaluation of large language models (LLMs) remains a critical challenge, particularly in assessing their sensitivity to perturbations in input data. In this work, we systematically evaluate LLM robustness across multiple dimensions, including word error rate, character repetition and duplication, modifications in choices, and variability in instruction following. To facilitate this evaluation, we construct a synthetic and augmented dataset encompassing a diverse set of LLM benchmarks, specifically targeting multiple-choice question (MCQ) datasets and instruction-following tasks. We conduct extensive experiments on LLMs of varying scales-small, medium, and large-as well as across base and instruction-tuned variants. Our analysis quantifies the variability in model responses under perturbed conditions and highlights discrepancies relative to baseline models. The findings provide insights into the stability of LLMs across different evaluation scenarios contributing to the development of more robust and reliable language models as well as robust evaluation methodologies.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑