当规范冲突时:一种基于对称性的测量大语言模型(LLM)偏好的框架
When Specifications Conflict: A Symmetry-Based Framework for Measuring LLM Preferences
浏览论文内容
中文总结 AI 辅助
本研究提出基于对称性的框架,通过含550个实例的数学基准等评估,发现LLM在冲突规范下有系统性偏好,排序为形式类>纯自然语言>输入-输出示例,可跨领域应用。
中文摘要 AI 辅助
大语言模型(LLM)日益需要整合多个可能不一致或冲突的信息源,但目前仍缺乏可控且可归因的方法来分析模型如何解决竞争规范间的冲突。我们提出一种受控实验框架,用于研究冲突规范下的模型偏好,通过构建存在明确冲突的规范,可直接观察并分析模型在竞争规范间的选择。该框架采用基于对称性的设计,进一步减少混杂因素,使能系统比较不同表示类型的偏好。我们在包含11个函数族、550个冲突实例的可执行数学基准上评估该框架,对比四种表示类型:纯自然语言、形式语言、归化形式语言及输入-输出示例。结果显示模型存在系统性偏好模式而非随机行为,一致排序为:$ \text{Formal} \text{近似等于} \text{Naturalized Formal} > \text{Pure Natural Language} > \text{Input--Output Examples} $,示例效果还取决于模型能力和函数族。我们将该框架扩展到布尔代数、代码生成及临床领域的异构规范冲突,证明其适用于多样任务和规范形式,为测量LLM如何解决竞争信息源间冲突提供了统一方法。
英文摘要
Large language models (LLMs) are increasingly required to integrate multiple sources of information that may be inconsistent or conflicting. However, there is still a lack of controllable and attributable methods for analyzing how models resolve conflicts between competing specifications. We propose a controlled experimental framework for studying model preferences under conflicting specifications. By constructing specifications with explicit conflicts, the framework enables model choices between competing specifications to be directly observed and analyzed. A symmetry-based design further reduces confounding factors, allowing preferences across representation types to be compared systematically. We evaluate the framework on an executable mathematical benchmark with 550 conflict instances spanning 11 function families, comparing four representation types: pure natural language, formal language, naturalized formal language, and input--output examples. Results show systematic preference patterns rather than random behavior, with a consistent ordering: $ \text{Formal} \approx \text{Naturalized Formal} > \text{Pure Natural Language} > \text{Input--Output Examples} $. Example effects further depend on model capability and function family. We extend the framework to heterogeneous specification conflicts in Boolean algebra, code generation, and the clinical domain, demonstrating its applicability across diverse tasks and specification forms. The framework provides a unified approach for measuring how LLMs resolve conflicts between competing sources of information.