arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

STEMMA:用于评估大语言模型(LLM)自我身份一致性的对抗性多智能体框架

STEMMA: An Adversarial Multi-Agent Framework for Evaluating Self-Identity Consistency in LLMs

Nuthakki Siva Gopala Krishna, Kanishka Jain

arXiv 2608.08164首次发表:更新:

发表机构

BML Munjal; Indian Institute of Technology, Delhi(BML Munjal(BML蒙贾尔); 印度理工学院德里分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对LLM自我身份一致性评估问题,提出对抗性多智能体框架STEMMA,结合手动设计的对抗性提示开展实验,发现多数LLM存在自我表征不一致的问题。

AI 中文摘要

知识蒸馏是大语言模型(LLM)训练与微调中广泛采用的技术,可将大型教师模型的结构化信息与功能行为迁移至规模更小的学生模型,同时大幅降低计算成本。然而,随着蒸馏技术在规模与复杂度上的提升,一个重要问题随之产生:教师模型究竟迁移了何种类型的知识?本研究认为,除功能知识外,学生模型还会学习到行为模式,尤其是模型表征自身身份的方式,这引发了对输出同质性、模型偏见及问责制的担忧。为应对这一挑战,我们提出STEMMA,这是一个多模态多智能体框架,其中角色特定的智能体可协同探测不同模型的自我身份表征行为。我们还贡献了一组手动设计的对抗性提示,用于评估LLM的身份一致性。实验结果表明,在一定程度上,大多数模型的自我表征都存在不一致性,易受此类问题影响。

英文摘要

Knowledge Distillation is a widely adopted technique in the training and fine-tuning of large language models (LLMs) enabling transfer of structured information and functional behavior from a large teacher model to a smaller student model while significantly reducing computational costs. However, as the use of distillation increases in both scale and complexity it raises an important question about what kind of knowledge is really transferred from the teacher model. In this work, we argue that apart from the functional knowledge, student models also learn behavioral patterns, specifically how a model represents its own identity raising concerns about output homogeneity, model biases, and accountability. To address this challenge, we introduce STEMMA, a multi-modal and multi-agent framework in which role specific agents collaboratively probe self identification behavior in different models. We also contribute a set of adversarial prompts designed manually to evaluate identity consistency in LLMs. Our results show that to an extent most models are vulnerable to inconsistencies in self-representations.

Comments15 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑