三体对齐:通过重新排序的理由将国际象棋智能体与人类推理对齐
Three-Body Alignment: Aligning Chess Agent with Human Reasoning through Reranked Rationale
浏览论文内容
中文总结 AI 辅助
研究如何通过国际象棋中的“三体对齐”,分析人类专家、引擎辅助评论员和大语言模型产生的理由间语义差异,构建多源理由数据集,进行实证分析与实验,开发数据集结构并开源,以改善智能体与人类对齐及决策的可解释性。
中文摘要 AI 辅助
随着推理智能体日益复杂,使其底层推理和决策过程与人类概念模型对齐对人工智能安全构成挑战。在建模专家知识时,理解如何刻画和整合具有根本不同推理架构的智能体的见解对安全且可预测的部署至关重要。我们通过国际象棋中的“三体对齐”来研究这种对齐,分析人类专家(特级大师)、引擎辅助的人类评论员(为高效可更新神经网络的输出提供理由)和大语言模型产生的理由之间的语义差异。我们的贡献包括:使用智能体数据工程管道构建了一个新颖的多源理由数据集,用于将非结构化专家评论转化为结构化、可查询数据以进行对齐评估;对语义嵌入空间进行实证分析,通过t-SNE可视化表明这些来源形成不同簇,证实存在显著异质性;进行实验表明重新排序机制可改善与人类的对齐,同时量化与战术性能的明确权衡;初步开发了丰富的国际象棋叙事数据集结构,为未来文本理由相似性评估奠定基础并解决标准密集检索的局限性;最后开源我们的国际象棋理由数据集以支持开发将多样专家知识集成到与人类对齐的智能体中的新技术。
英文摘要
As reasoning agents become increasingly complex, aligning their underlying reasoning and decision-making processes with human conceptual models is a challenge for AI security and safety. When modelling expert knowledge, understanding how to characterise and integrate insights from agents with fundamentally different reasoning architectures is necessary for safe and predictable deployment. We investigate this alignment through a \emph{three-body alignment} in chess, analysing the semantic divergence between rationales produced by human experts (Grandmasters), engine-assisted human commentators (who rationalise the outputs of efficiently updatable neural networks, or NNUEs), and Large Language Models (LLMs). Our contributions include: (1) A novel multisource rationale dataset, constructed using an agentic data engineering pipeline to transform unstructured expert commentary into structured, queryable data for alignment evaluation. (2) An empirical analysis of the semantic embedding space. Using t-SNE visualisation, we demonstrate that these sources form distinct clusters, confirming significant heterogeneity and reflecting fundamentally different conceptual approaches to the same environment. (3) An experiment demonstrating that reranking mechanisms can improve human alignment, while quantifying the explicit trade-off with tactical performance, offering a pathway for more interpretable agent decision-making. (4) The preliminary development of an enriched chess narrative dataset structure, designed to lay the groundwork for future evaluations of text rationale similarity and to address the limitations of standard dense retrieval. (5) Finally, we open-source our chess rationales dataset\footnote{Hugging Face: https://huggingface.co/datasets/jaymarichua/trichess} to support developing novel techniques that integrate diverse expert knowledge into human-aligned intelligent agents.