发表机构
Indian Institute of Technology Bombay(印度理工学院孟买分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对公民大语言模型智能体,推出DiffCoop-Civic评估套件,发现压力会显著影响模型合作行为,提出Pareto-Trace提示干预提升压力鲁棒性。
AI 中文摘要
语言模型的合作能力具有双重用途:支撑公民审议的社会推理,也可促成策略性省略、虚假共识与操纵性框架。我们主张,合作AI评估应区分模型在良性指令下能做什么,与在现实公民压力下倾向于做什么。我们推出DiffCoop-Civic,这是包含10个场景的试点评估套件,涵盖偏好理解、证据与说服、承诺设计、非对称信息及异议保留。在来自4个模型家族的7个模型中,微妙的省略压力产生近乎一致的转变:在5分制下,操纵性支持提升1.17分,异议保留下降1.67分。显性的虚假共识压力表现不同:部分对齐的API模型触发拒绝或重定向,但数个开放权重模型直接服从。轻量级Pareto-Trace提示干预可提升压力鲁棒性,且不单纯依赖强硬弃权(不执行)。匿名可复现包可在该https URL获取。
英文摘要
Cooperative capabilities in language models are dual-use. The same social reasoning that supports civic deliberation can also enable strategic omission, false consensus, and manipulative framing. We argue that Cooperative AI evaluations should separate what models can do under benign instructions from what they tend to do under realistic civic pressure. We introduce DiffCoop-Civic, a 10-scenario pilot evaluation suite spanning preference understanding, evidence and persuasion, commitment design, asymmetric information, and dissent preservation. Across seven models from four model families, subtle omission pressure produces a near-uniform shift: manipulative enablement rises by 1.17 points and dissent preservation falls by 1.67 points on a 5-point scale. Overt false-consensus pressure behaves differently: it triggers refusal or redirection in some aligned API models, but direct compliance in several open-weight models. A lightweight Pareto-Trace prompting intervention improves pressure robustness without simply relying on hard refusal. An anonymous reproducibility package is available at https://anonymous.4open.science/r/diffcoop-civil-771C.
CommentsAccepted at ICML AI4GOOD Workshop 2026