AI 中文总结
该研究将关系先验注入LLM-MAS提示中,发现其主要起收敛压力作用,正向性提升可促进智能体协调但未必提高准确性,建议按需使用而非默认添加。
AI 中文摘要
基于大语言模型的多智能体系统(LLM-MAS)通过角色、辩论协议和聚合规则进行设计,这些选择会产生隐含的社会预期:智能体可能被要求信任、质疑、服从或与同伴协作。我们研究将智能体间关系语义显式化的效果,采用关系先验的最小符号网络公式,并将自然语言表述注入智能体系统提示中,同时保持任务协议固定。在公共治理模拟和多智能体辩论任务中,关系先验主要起到收敛压力的作用:增加关系正向性往往会使智能体更易于协调或达成一致。当效用奖励行为对齐时,这种压力会有所帮助,例如在可持续资源治理和主观共识场景中;但它并不能可靠地提升准确性,在客观问答辩论中,即使基于正确性的一致性没有提升甚至在某些场景下降时,更高的正向性也会增加一致性。效果因模型主干、关系类型和拓扑结构而异,显式中立性并不等同于省略关系框架。我们认为关系先验不应作为LLM-MAS的默认附加组件,其更安全的用法是诊断性的和任务特定的:与无先验基线对比,在涉及真相时监测基于正确性的指标,当验证不证实时省略关系层。
英文摘要
Large language model-based multi-agent systems (LLM-MAS) are designed through roles, debate protocols, and aggregation rules. These choices create implicit expectations of trust, skepticism, deference, or collaboration. We make inter-agent relations explicit as signed pairwise priors, rendered in natural language and added to system prompts while keeping the task protocol fixed. Across commons governance and multi-agent debate, these priors change how readily agents coordinate or agree, a pattern we call convergence pressure. More positive relations generally improve sustainability within the GovSim relational-prior sweep and increase consensus on subjective questions. These improvements over negative relations differ from gains over no-prior prompting. On objective QA, fully positive priors usually yield lower final-answer accuracy than the no-prior baseline, and some conditions produce more frequent but less accurate consensus. Effects depend on model backbone, relation type, and topology; explicitly neutral relations also produce different outcomes from omitting relational framing. Relational priors are therefore task-specific interventions and diagnostic probes of sensitivity to social framing. Evaluate them against a no-prior baseline using the task's primary metric, report accuracy and consensus correctness alongside agreement on objective tasks, and keep no-prior prompting as the default for accuracy-centric tasks unless validation supports a relational prior.