arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

推理表示是否有助于人类评估LLM输出?

Do Reasoning Representations Help Humans Evaluate LLM Outputs?

Jaewoo Lim, Sungbok Shin, Sanghyun Hong

arXiv 2609.09038首次发表:更新:

发表机构

Oregon State University; Sogang University(俄勒冈州立大学; 西江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过受控人类实验发现,推理表示虽受偏好,但简单思维链更利于人类评估LLM输出,且偏好表示存在信任校准风险。

AI 中文摘要

推理表示越来越多地被用作大型语言模型输出的解释。然而,它们通常以模型为中心的指标(如答案准确性和忠实度)进行评估,这使得它们是否有助于人们评估模型响应仍不明确。在这项工作中,我们将推理表示视为面向人类的界面,而非模型推理能力的代理。我们进行了一项受控的人类研究,涵盖六种推理格式,涉及不同复杂度的任务,并借助一个基于网络的框架,该框架随机化任务领域、问题实例和表示顺序。该研究收集了关于结构理解、错误检测与定位以及信任校准的细粒度判断。我们的研究显示,感知偏好与人类评估支持之间存在不匹配。参与者偏好基于规划和分解的表示,但更简单的思维链轨迹能更好地支持验证、信任和可解释性。偏好的表示还引入了校准风险,在正确轨迹上产生更多误报,且尽管验证意愿低,信任度却很高。

英文摘要

Reasoning representations are increasingly used as explanations for large language model outputs. Yet they are typically evaluated with model-centric criteria, such as answer accuracy and faithfulness, leaving it unclear whether they help people evaluate model responses. In this work, we study reasoning representations as human-facing interfaces rather than proxies for model reasoning ability. We conduct a controlled human study of six reasoning formats across tasks of varying complexity, supported by a web-based framework that randomizes task domains, problem instances, and representation order. The study collects fine-grained judgments of structural understanding, error detection and localization, and trust calibration. Our study shows a mismatch between perceived preference and support for human evaluation. Participants prefer planning- and decomposition-based representations, but simpler chain-of-thought traces better support verification, trust, and interpretability. Preferred representations also introduce calibration risks, with more false alarms on correct traces and high trust despite low willingness to verify.

Comments19 pages. EMNLP 2026 (Findings)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑