Graphionale:LLM推理依据的图形可视化如何影响人类决策
Graphionale: How Graph Visualizations of LLM Rationales Affect Human Decision Making
浏览论文内容
中文总结 AI 辅助
本研究开发了Graphionale作为测试平台,通过204人在线用户研究发现,图形化LLM推理依据对不同任务模态的决策影响不同,匹配推理格式与任务模态是AI解释设计的关键。
中文摘要 AI 辅助
大型语言模型(LLMs)正日益具备增强的推理能力,可生成支持人类决策的推理依据。然而,这些文本密集型的推理依据往往会带来沉重的认知负担。基于一项初步协同设计研究(该研究确定了用户对非线性推理表示的偏好),我们开发了Graphionale作为实证研究论证图式推理依据可视化的测试平台。该系统将线性的LLM推理依据转化为交互式多级图形,明确构建逻辑关系(如结论、前提、支持与反驳),同时进一步提取每个陈述中的实体与关系,以构建精简的节点-链接表示。我们开展了一项大规模在线用户研究(样本量N=204),检验在不同任务模态(言语推理vs视觉推理)、推理依据格式(文本型vs图形型)及问题难度(简单vs困难)下,图形化推理依据何时比文本型更有效。研究结果表明,图形化推理依据的效果并非统一:它们能提升言语推理的信任校准,但会带来更高的认知负担且满意度更低;对于视觉推理,它们会损害信任校准,但会更具吸引力且更有帮助。在每种模态中,更能支持校准决策的格式并非用户偏好的格式,这表明将推理依据格式与任务模态相匹配是有效AI解释设计的关键。我们的研究结果为图形化推理依据何时及如何支持人类决策提供了实证设计知识,并为下一代感知推理的AI界面提供了参考。
英文摘要
Large Language Models (LLMs) are increasingly equipped with augmented reasoning capabilities to generate rationales that support human decision-making. Yet these text-dense rationales often impose substantial cognitive burdens. Building on a formative co-design study that identified user preferences for non-linear reasoning representations, we developed Graphionale as a testbed for empirically studying argument-map-style rationale visualization. This system transforms linear LLM rationales into interactive, multi-level graphs. It explicitly structures logical relationships (e.g., conclusions, premises, support, and objections), while further extracting entities and relations within each statement to construct condensed node-link representations. We conduct a large-scale online user study (N = 204) to examine when graphical rationales are more effective than textual ones, across varying task modality (verbal vs. visual reasoning), rationale format (textual vs. graphical), and question difficulty (easy vs. hard). Our results show that graphical rationales do not help uniformly: they improve trust calibration for verbal reasoning yet feel more cognitively demanding and less satisfying; for visual reasoning, they impair calibration yet feel more engaging and helpful. In each modality, the format that better supports calibrated decisions is not the one users prefer, highlighting that matching rationale format to task modality is key to effective AI explanation design. Our findings contribute empirical design knowledge about when and how graphical rationales support human decision making, and inform the next-generation reasoning-aware AI interfaces.