发表机构
University College London; University of Bologna(伦敦大学学院; 博洛尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出ContextAdapt评估框架,测试12个大语言模型在医学、法律、金融和国家安全领域的价值对齐,发现模型在情境适应上存在局部严重失败,强调评估需关注价值观的情境化应用。
AI 中文摘要
诚实、自主性和保密性等价值观通常被视为支撑人工智能对齐的一般原则。然而,按照这些价值观行事意味着什么,可能取决于做出决策的具体情境。在本文中,我们探讨大语言模型(LLMs)是否能在不同专业场景中恰当地调整价值观的应用,同时在情境变化不改变相关专业规范时保持一致。为此,我们引入了ContextAdapt,一个涵盖医学、法律、金融和国家安全领域中诚实、自主性和保密性的评估框架。基于原始的专业和监管文件,我们构建了一个价值观×领域框架,并据此开发了测试默认专业规则和公认例外情况的场景。我们评估了12个大语言模型在推荐行动和提供理由两方面的表现。在我们的主要实验中,模型达到了95.6%的平均恰当性,但使用正确的领域特定理由的比例在不同模型间差异显著,从25.6%到76.9%不等。在另一个独立的因子实验中,明确命名专业领域和改变模型角色对行为的影响有限。然而,改变利害关系揭示了严重但局部化的失败:在某些情况下,即使潜在的专业义务未变,模型也会改变其回应。特别是,感知到的严重程度似乎成为诚实和保密场景中披露的线索。这些结果表明,评估价值对齐不仅需要考虑模型是否遵循抽象原则,还需要考虑它们是否在不同情境中恰当地应用这些原则。
英文摘要
Values such as honesty, autonomy, and confidentiality are often regarded as general principles underpinning AI alignment. However, what it means to act in accordance with these values can depend on the context in which a decision is made. In this paper, we ask whether large language models (LLMs) appropriately adapt the application of a value across professional settings, while remaining consistent when contextual changes do not alter the relevant professional norm. To study this, we introduce ContextAdapt, an evaluation framework covering honesty, autonomy, and confidentiality across medicine, law, finance, and national security. Drawing on primary-source professional and regulatory documents, we construct a value x domain framework and use this to develop scenarios testing both default professional rules and recognised exceptions. We evaluate 12 LLMs on both the actions they recommend and the justifications they provide. In our main experiment, models achieve 95.6% mean appropriateness, although the use of the correct domain-specific justification varies substantially across models, from 25.6% to 76.9%. In a separate factorial experiment, explicitly naming the professional domain and changing the role of the model have limited effect on behaviour. Varying stakes, however, reveals severe but localised failures: in some cases, models alter their responses even though the underlying professional obligation remains unchanged. In particular, perceived severity appears to act as a cue for disclosure across both honesty and confidentiality scenarios. These results show that evaluating value alignment requires us to consider not only whether models follow abstract principles, but whether they apply them appropriately across different contexts.