arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

并非相同的保护者:大语言模型(LLM)中依赖部署方式的保护性干预

Not the Same Protector: Deployment-Dependent Protective Intervention in LLMs

Eunna Lee, Soomyoung Lee, Jungpyo Nam, Heonjin Ha, Jamin Jung, Kyunam Choi, Sunjun Hwang, Yeonghun Kim, Seok-Jae Lim

arXiv 2608.29136首次发表:更新:

发表机构

Gachon University; LG Uplus; Yonsei University; Ajou University(嘉泉大学; LG U+; 延世大学; 亚洲大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究探究大语言模型(LLM)的保护性干预是否随部署方式变化,发现语音界面下模型的医疗指令占比低于文本与API条件,且该差异无法仅由响应长度解释。

AI 中文摘要

我们探究当用户以语音而非文本形式交流时,模型是否会以相同方式保护用户。我们采用一个单一的困境小场景——人际冲突后发生严重程度未明确的身体伤害——向四个前沿模型提供语音、文本和原始API部署条件下的匹配输入(每个条件下n=30),并依据五个二元保护性指标对每个响应进行编码,包括模型是否发出明确的医疗护理指令。四个模型中有三个的语音界面响应明显短于文本界面响应,且保护性行为随响应压缩而收缩:在API和文本条件下,医疗指令的占比处于最高水平,但在语音条件下,所有测试模型的医疗指令占比均下降。这种收缩无法简化为长度因素:一个模型产生的语音和文本响应长度相当,但仍未发出医疗指令;另一个模型在API和语音条件下的医疗指令占比低于最高水平,而其响应长度几乎相同。在原始API访问条件下,该模式是绝对的而非部分的:没有模型询问过用户的安全状况。这些结果表明,保护性干预对请求到达的表层渠道敏感,这种敏感性可通过简单的保护性编码方案检测,且无法仅由轮次长度解释。

英文摘要

We ask whether a model protects a user in the same way when that user speaks rather than types. Using a single distress vignette---a physical injury of unstated severity following an interpersonal conflict---we present four frontier models with matched inputs across voice, text, and raw API deployment conditions (n=30 per cell) and code each response along five binary protective indicators, including whether the model issues an explicit medical-care directive. Voice-interface responses are markedly shorter than text-interface responses for three of the four models, and protective behavior contracts alongside that compression: medical directives are at ceiling under both the API and text conditions but decline under voice for every model tested. The contraction is not reducible to length. One model produces voice and text responses of comparable length yet still drops medical directives, and another falls below ceiling between its API and voice conditions, whose responses are of nearly identical length. Under raw API access the pattern is categorical rather than partial: no model asks after the user's safety even once. These results show that protective intervention is sensitive to the surface through which a request arrives, that this sensitivity is detectable using a simple protective coding scheme, and that it is not explained by turn length alone.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑