发表机构
University of Wollongong; Adelaide University(伍伦贡大学; 阿德莱德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VACS提出四层组合屏蔽框架,通过逆强化学习推断价值、形式化屏蔽保障安全、核仁与哈密顿共识解决冲突,在三个基准上提升准确率并降低逻辑不一致率。
AI 中文摘要
在高风险领域中的多智能体推理系统必须既准确又安全,然而智能体往往遵循异质的价值优先级(例如严谨性、简洁性、安全性),导致相互冲突的建议。现有方法无法同时提供:(i)从行为中对每个智能体隐含价值的原理性推断,(ii)无需完全在线通信的组合式形式安全保证,以及(iii)具有忠实解释的价值感知冲突解决。我们提出VACS(价值对齐组合屏蔽),一个四层框架,解决上述所有三个问题。第一层使用Bradley-Terry建模从成对偏好中学习价值维度奖励,并通过深度最大熵逆强化学习推断每个智能体的价值权重。第二层在受Lean启发的领域特定语言中编码价值约束,并综合组合式假设-保证屏蔽以实现运行时安全。第三层通过基于核仁的信用分配和长期价值约束下的哈密顿共识优化来解决分歧。第四层从协态敏感性中提取关键推理路径,并生成具有形式基础的自然语言解释。我们的贡献主要是一个统一的系统设计,具有形式化接口和在验证器约束决策层面的操作保证,而非对所有语言模型内部机制的完整端到端形式证明。在受控的概念验证评估中,使用角色条件智能体面板在NEJM-AI问答、MathInstruct-Subset和网络安全事件响应基准(CyberSec-Eval)上,VACS在准确性(85.4%、95.0%和90.0%)上优于强基线,同时将逻辑不一致率降至接近零。
英文摘要
Multi-agent reasoning systems in high-stakes domains must be both accurate and safe, yet agents often follow heterogeneous value priorities (e.g., rigor, conciseness, safety), causing conflicting recommendations. Existing methods do not jointly provide: (i) principled inference of each agent's implicit values from behavior, (ii) compositional formal safety guarantees without full online communication, and (iii) value-aware conflict resolution with faithful explanations. We present VACS (Value-Aligned Compositional Shielding), a four-layer framework addressing all three. Layer 1 learns value-dimension rewards from pairwise preferences using Bradley-Terry modeling and infers per-agent value weights via deep MaxEnt IRL. Layer 2 encodes value constraints in a Lean-inspired DSL and synthesizes compositional assume-guarantee shields for runtime safety. Layer 3 resolves disagreement through nucleolus-based credit allocation and Hamiltonian consensus optimization under long-term value constraints. Layer 4 extracts a critical reasoning path from co-state sensitivities and generates formally grounded natural-language explanations. Our contribution is primarily a unified systems design with formalized interfaces and operational guarantees at the verifier-constrained decision level, rather than a complete end-to-end formal proof of all language-model internals. In controlled proof-of-concept evaluations with role-conditioned agent panels on NEJM-AI QA, MathInstruct-Subset, and a cybersecurity incident-response benchmark (CyberSec-Eval), VACS outperforms strong baselines in accuracy (85.4%, 95.0%, and 90.0%) while reducing logical inconsistency rates to near zero.