基于多层组合融合的上下文价值对齐
Contextual Value Alignment via Multilayer Combinatorial Fusion
浏览论文内容
中文总结 AI 辅助
本研究提出MCF-CVA框架,通过多层组合融合及EAR算法,结合多道德智能体的认知多样性缓解冲突冗余,在标准指标上优于单智能体等基线,提升LLM的上下文价值对齐能力。
中文摘要 AI 辅助
将大语言模型(LLMs)与人类价值观对齐仍是可信赖人工智能领域的重大挑战。尽管RLHF、CAI及其变体等现有方法已取得良好效果,但它们通常依赖单智能体框架和统一奖励系统,这限制了其捕捉伦理多元性、适应不同道德语境以及反映多智能体道德推理动态的能力。本研究提出了一种用于上下文价值对齐的多层组合融合框架(MCF-CVA)。在框架的第一层,它实例化了多个道德智能体,每个智能体都经过微调以代表一种独特的价值观。它们的输出随后通过基于分数和排名的组合以及平均和加权聚合进行组合扩展。这些组合后的模型随后被缩减为与初始道德智能体相同的数量。这种扩展与缩减(EAR)过程会持续多层,直到达到停止准则。MCF-CVA框架利用智能体之间的认知多样性来缓解多个智能体之间的冲突和冗余,产生更好反映上下文人类价值观的响应。该框架使用EAR算法在欧几里得分数空间和Kemeny排名空间的双重架构上执行。实证评估表明,所提出的框架在标准指标上优于单智能体基线、多智能体单层结果以及先前的聚合方法,表明MCF-CVA框架为推进LLMs中的上下文价值对齐提供了一种稳健且有效的机制。
英文摘要
Aligning large language models (LLMs) with human values remains a major challenge, especially for trustworthy AI. While existing approaches such as RLHF, CAI, and their variants have achieved promising results, they often rely on a single-agent framework and a unified reward system. This limits their ability to capture ethical pluralism, adapt to diverse moral contexts, and reflect the dynamics of multi-agent moral reasoning. In this work, we propose a framework that utilizes multilayer combinatorial fusion for contextual value alignment (MCF-CVA). At the first layer of the framework, it instantiates multiple moral agents, each fine-tuned to represent a distinctive value. Their outputs are then expanded combinatorially using both score- and rank-combinations as well as average and weighted aggregations. These combined models are then reduced to the same number of initial moral agents. This expansion and reduction (EAR) process continues for multi-layers until a stopping criterion is reached. The MCF-CVA framework leverages cognitive diversity between agents to mitigate conflicts and redundancies across multiple agents, producing responses that better reflect contextual human values. The framework using the EAR algorithm is performed on the dual architecture of Euclidean score space and Kemeny rank space. Empirical evaluations demonstrated that the proposed framework outperforms single-agent baselines, multi-agent single-layer results, and previous aggregation approaches on standard metrics, showing that the MCF-CVA framework provides a robust and effective mechanism for advancing contextual value alignment in LLMs.