arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

以用户为中心的思维链推理提升大语言模型可解释性

Improving LLM Interpretability with User-Centric Chain-of-Thought Reasoning

Philipp Schröppel

arXiv 2608.26166首次发表:更新:

发表机构

University of Ulm(乌尔姆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出以人为本的LLM推理轨迹构建方法,采用类XML标签编码,在保持数学推理性能的同时提升可解释性,显著改善用户感知的有用性与易用性,助力高风险AI部署的人机协作。

AI 中文摘要

推理能力的提升使大语言模型(LLM)能够应对日益复杂的问题,而推理轨迹——即通往解决方案的中间步骤——通过让人类检查AI决策过程,开辟了高风险应用场景。然而,当前方法优先考虑模型性能而非人类可解释性,限制了人机协同的有效性。本研究设计并评估了一种以人为本的方法,该方法基于自包含、可验证的步骤构建推理轨迹,使用户能够独立评估和修正AI推理。该方法采用类XML标签对推理内容和元数据进行编码,便于针对性反馈。在数学推理任务上的评估显示,本方法在保持与标准思维链推理相当性能的同时,提升了可解释性;用户研究表明,其感知有用性和易用性显著提升。本研究增进了对LLM输出的以用户为中心的设计如何更好满足高风险AI部署中人类协作需求的理解。

英文摘要

Advancing reasoning capabilities allow large language models (LLMs) to tackle increasingly complex problems, while reasoning traces - intermediate steps toward solutions - open up high-stakes applications by enabling human inspection of AI decision-making. However, current approaches prioritize model performance over human interpretability, limiting effective human-AI collaboration. In this study, we design and evaluate a human-centered approach that structures reasoning traces based on self-contained, verifiable steps, enabling users to independently assess and correct AI reasoning. Our approach uses XML-like tags to encode reasoning content and metadata, facilitating targeted feedback. Evaluation on mathematical reasoning tasks shows our approach maintains equivalent performance to standard Chain-of-Thought reasoning while enhancing interpretability. User studies demonstrate significant improvements in perceived usefulness and ease of use. This work advances understanding of how user-centric design of LLM outputs can better serve human collaboration needs in high-stakes AI deployments.

Journal refProceedings of the 59th Hawaii International Conference on System Sciences (pp. 1435-1444). University of Hawai'i at Manoa, 2026

DOI:10.24251/HICSS.2026.171

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑