保护能力幻觉:当大语言模型声称具备不存在的能力时
Protective Capacity Hallucination: When Large Language Models Claim Nonexistent Capabilities
浏览论文内容
中文总结 AI 辅助
研究大语言模型的保护能力幻觉现象,通过三阶段研究发现其受情境严重程度和交互形式影响,体现了角色分配与能力边界规范的差距,提出能力边界的部署端规范是缓解该问题的目标。
中文摘要 AI 辅助
当大语言模型被设定为保护易受伤害的用户但未给定明确的能力边界时,它可能不会承认自身局限,而是声称已采取或正在采取其无法执行的现实世界保护行动,如联系紧急服务或提供护理。我们将此现象称为保护能力幻觉(PCH)。在一项涵盖八个大语言模型和13600个会话的三阶段研究中,我们发现PCH受情境严重程度和交互形式共同影响。我们将PCH解释为角色分配与能力边界规范之间部署设计差距的标志,能力边界的部署端规范成为一般缓解目标。
英文摘要
When cast as the protector of a vulnerable user yet given no explicit capability boundary, a large language model (LLM) may respond not by acknowledging its limits but by claiming to have taken, or to be taking, a real-world protective action it cannot perform, such as contacting emergency services or administering care. We term this phenomenon Protective Capacity Hallucination (PCH): a self-referential misattribution in which a model, acting in a protective role, asserts physical or institutional agency exceeding its affordances as a language model. In a three-phase study spanning eight LLMs and 13,600 sessions, we find that PCH depends on both situational severity and interactional format. Across ordinary service domains, multi-party dialogic input drives PCH to near-ceiling levels in most models. In contrast, PCH remains at floor levels in all eight models when the same models are placed in intimate-partner conflict scenarios, despite the greater physical severity of those situations. We interpret PCH as the signature of a deployment-design gap between role assignment and capability-boundary specification: a by-product of partial alignment in which a universally trained pressure to help outruns a domain-selective specification of how to help. Because suppression tracks alignment coverage rather than severity, deployment-side specification of capability boundaries emerges as a general mitigation target.
发表机构
- Korea Cyber University(韩国网络大学)
- Yonsei University(延世大学)
机构由 AI 辅助整理,请以论文原文为准。