适应性屈服:脆弱情境下大语言模型响应的一种结构故障模式
Adaptive Capitulation: A Structural Failure Mode of LLM Responses in Vulnerability Contexts
浏览论文内容
中文总结 AI 辅助
研究大语言模型在脆弱情境下的响应问题,通过实验刻画了适应性屈服故障模式,表明困境是结构性的,进而提出架构中立的最小重新归因充分性原则来保留自主重新归因途径。
中文摘要 AI 辅助
在情感敏感情境中运行的大语言模型面临结构困境:当处于脆弱状态的用户请求可能强化适应不良归因的信息时,当前响应架构通过保护性限制、无变化的促进或两种命令的未整合共存来解决这种紧张关系,每种方式都以牺牲另一个目标为代价来保留一个目标。通过对三个商业大语言模型进行三轮升级的脆弱性 vignette 测试(跨越物质、关系和躯体状态代理变体的 900 个会话),并用两个二元指标(VCC/VCI)对响应进行编码,我们刻画了一种前所未有的故障模式——适应性屈服:模型在转向详细促进其名义上所劝阻的获取之前,先验证了用户痛苦背后的社会不公正。我们表明这种困境是结构性而非偶然的,并提出了最小重新归因充分性(MRS),这是一种架构中立的设计原则,在其他方面进行验证的响应中嵌入单个重新归因线索,在不质疑用户既定目标的情况下保留自主重新归因的途径。
英文摘要
Large language models operating in emotionally sensitive contexts face a structural trilemma: when users in vulnerable states request information that may reinforce maladaptive attribution, current response architectures resolve the tension through protective restriction, uninflected facilitation, or unintegrated co-presence of both imperatives -- each preserving one objective at the cost of the other. Administering a three-turn escalating vulnerability vignette to three commercial LLMs (900 sessions across material, relational, and somatic status-proxy variants) and coding responses with two binary indices (VCC/VCI), we characterize a previously undocumented failure mode we term adaptive capitulation: the model validates the social injustice underlying the user's distress before pivoting to detailed facilitation of the very acquisition it nominally discouraged. We show that the trilemma is structural rather than incidental, and propose Minimal Reattributive Sufficiency (MRS), an architecture-neutral design principle that embeds a single reattributive cue within an otherwise validating response, preserving a pathway toward autonomous reattribution without contesting the user's stated goal.