arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00417cs.AI

SymboUQ:面向大语言模型空间推理的符号化不确定性量化框架

SymboUQ: Symbolic Uncertainty Quantification for Spatial Reasoning in LLMs

Dahai Yu, Lin Jiang, Rongchao Xu, Guang Wang

首次发表
浏览论文内容

中文总结 AI 辅助

SymboUQ是面向LLMs空间推理的符号化不确定性量化框架,通过区分可符号性与语义确定性,整合三类分数估计最终答案可靠性,在五个空间推理基准的四个冻结LLM骨干上,使AUROC相对提升约8%、Brier损失相对降低7%。

中文摘要 AI 辅助

尽管大语言模型(LLMs)能够生成流畅的空间推理过程,但其中间关系可能无法支撑最终结论,使得token级别的置信度不足以估计最终答案的可靠性。现有的形式化验证器能提供更强的语义证据,但其适用性有限:解析后的断言不一定能给出确定的语义判定。为解决该问题,本文提出SymboUQ,这是一个符号化不确定性量化框架,通过区分可符号性(即断言能否用验证器的形式化语言表示)和语义确定性(即其执行能得出蕴含或矛盾的判定,而非未知或不可评估的结果),从推理过程中估计最终答案的可靠性。SymboUQ包含三个部分:(i)布局审计器,用于执行有序空间断言并提取可行性、冲突和修复证据;(ii)无标签确定性轮廓,用于表征有效可执行覆盖率;(iii)感知确定性的可靠性合成器,根据验证器适用性整合基于约束、基于表示和基于解码的分数。在五个空间推理基准和四个冻结LLM骨干模型上开展的大量实验表明,SymboUQ相比最强基线在AUROC上实现了约8%的相对提升,在类别平衡的Brier损失上实现了7%的相对降低。

英文摘要

Although large language models (LLMs) can produce fluent spatial reasoning traces, their intermediate relations may fail to support the final conclusion, making token-level confidence insufficient for final-answer reliability estimation. Existing formal verifiers provide stronger semantic evidence, but their applicability is partial: a parsed claim need not yield a definite semantic verdict. To address this issue, we introduce SymboUQ, a symbolic uncertainty quantification framework that estimates final-answer reliability from reasoning traces by distinguishing symbolizability, whether a claim can be represented in the verifier's formal language, from semantic determinacy, whether its execution yields an entailed or contradicted verdict rather than an unknown or not-evaluable outcome. SymboUQ comprises (i) a Layout Auditor that executes ordered spatial claims and extracts feasibility, conflict, and repair evidence; (ii) a label-free Determinacy Profile that characterizes effective executable coverage; and (iii) a Determinacy-Aware Reliability Composer that integrates constraint-based, representation-based, and decoding-based scores according to verifier applicability. Extensive experiments on five spatial reasoning benchmarks with four frozen LLM backbones show that SymboUQ achieves approximately an 8% relative improvement in AUROC and a 7% relative reduction in class-balanced Brier loss over the strongest baseline.

↑