发表机构
University of Ulm(乌尔姆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出以人为本的LLM推理轨迹构建方法,采用类XML标签编码,在保持数学推理性能的同时提升可解释性,显著改善用户感知的有用性与易用性,助力高风险AI部署的人机协作。
AI 中文摘要
推理能力的提升使大语言模型(LLM)能够应对日益复杂的问题,而推理轨迹——即通往解决方案的中间步骤——通过让人类检查AI决策过程,开辟了高风险应用场景。然而,当前方法优先考虑模型性能而非人类可解释性,限制了人机协同的有效性。本研究设计并评估了一种以人为本的方法,该方法基于自包含、可验证的步骤构建推理轨迹,使用户能够独立评估和修正AI推理。该方法采用类XML标签对推理内容和元数据进行编码,便于针对性反馈。在数学推理任务上的评估显示,本方法在保持与标准思维链推理相当性能的同时,提升了可解释性;用户研究表明,其感知有用性和易用性显著提升。本研究增进了对LLM输出的以用户为中心的设计如何更好满足高风险AI部署中人类协作需求的理解。
英文摘要
Advancing reasoning capabilities allow large language models (LLMs) to tackle increasingly complex problems, while reasoning traces - intermediate steps toward solutions - open up high-stakes applications by enabling human inspection of AI decision-making. However, current approaches prioritize model performance over human interpretability, limiting effective human-AI collaboration. In this study, we design and evaluate a human-centered approach that structures reasoning traces based on self-contained, verifiable steps, enabling users to independently assess and correct AI reasoning. Our approach uses XML-like tags to encode reasoning content and metadata, facilitating targeted feedback. Evaluation on mathematical reasoning tasks shows our approach maintains equivalent performance to standard Chain-of-Thought reasoning while enhancing interpretability. User studies demonstrate significant improvements in perceived usefulness and ease of use. This work advances understanding of how user-centric design of LLM outputs can better serve human collaboration needs in high-stakes AI deployments.
Journal refProceedings of the 59th Hawaii International Conference on System Sciences (pp. 1435-1444). University of Hawai'i at Manoa, 2026