arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2509.12179cs.AIcs.MA

协同对齐:重新思考对齐——作为双向人机认知适配的范式

Co-Alignment: Rethinking Alignment as Bidirectional Human-AI Cognitive Adaptation

  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

Yubo Li, Weiyi Song

更新

AI总结:

针对现有RLHF单向对齐范式将人类认知视为固定的局限,本文提出双向认知对齐(BiCA)的协同对齐框架,实验证明该方法可显著提升协作性能、安全性与协同增益,验证了双向对齐范式的有效性。

AI中文摘要:

当前通过RLHF实现的AI对齐遵循单向范式,即AI适配人类偏好,同时将人类认知视为固定不变的。我们提出转向通过双向认知对齐(BiCA)实现的协同对齐,在该框架下人类与AI可相互适配。BiCA采用可学习协议、表征映射和KL预算约束实现受控协同进化。在协同导航任务中,BiCA取得了85.5%的成功率,而基线方法为70.3%;其相互适配性能提升230%,协议收敛性能提升332%。涌现出的协议性能比人工设计的协议高出84%,同时双向适配还意外提升了安全性(分布外鲁棒性提升23%)。46%的协同增益表明最优协作存在于人类与AI能力的交集而非并集,验证了从单向对齐到协同对齐范式转变的合理性。

英文摘要:

Current AI alignment through RLHF follows a single directional paradigm that AI conforms to human preferences while treating human cognition as fixed. We propose a shift to co-alignment through Bidirectional Cognitive Alignment (BiCA), where humans and AI mutually adapt. BiCA uses learnable protocols, representation mapping, and KL-budget constraints for controlled co-evolution. In collaborative navigation, BiCA achieved 85.5% success versus 70.3% baseline, with 230% better mutual adaptation and 332% better protocol convergence. Emergent protocols outperformed handcrafted ones by 84%, while bidirectional adaptation unexpectedly improved safety (+23% out-of-distribution robustness). The 46% synergy improvement demonstrates optimal collaboration exists at the intersection, not union, of human and AI capabilities, validating the shift from single-directional to co-alignment paradigms.

↑