arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

这是你的最终答案吗?跨上下文一致性作为大语言模型可信度的度量

Is This Your Final Answer? Cross-Contextual Consistency as a Measure of LLM Credibility

Siyang Wu, Yibo Jiang, Bryon Aragam

arXiv 2608.10315首次发表:更新:

发表机构

University of Chicago(芝加哥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出用跨上下文一致性(C3)度量大语言模型可信度,通过对比26个模型在6个基准上的原始与扰动提示生成结果,发现C3可作为互补评估维度与基准有用性诊断工具。

AI 中文摘要

大语言模型(LLMs)是强大的黑箱系统,难以辨别其答案是否反映稳定的内部信念,还是表面的模式匹配。我们将跨上下文一致性确定为LLMs未被充分利用的行为属性:可信的答案在相同任务置于主题对齐、内容中立的上下文变化下应保持稳定。基于这一直觉,我们通过比较原始提示与扰动提示下的模型生成结果,将跨上下文一致性(C3)操作化。在涵盖推理、事实性和代码生成的26个模型与6个基准中,我们发现跨上下文偏移较小的答案更可能正确或符合事实。我们证明C3提供了互补的评估维度,可作为基准有用性诊断工具,即使聚合分数被广泛认为“饱和”,也能识别基准的哪些部分仍具信息性。

英文摘要

Large language models (LLMs) are powerful black-box systems, making it difficult to discern whether their answers reflect stable internal beliefs or superficial pattern matching. We identify cross-contextual consistency as an underutilized behavioral property of LLMs: a credible answer should remain stable when the same task is placed under topic-aligned, content-neutral contextual variation. Building on this intuition, we operationalize Cross-Contextual Consistency (C3) by comparing model generations under original and perturbed prompts. Across 26 models and six benchmarks spanning reasoning, factuality, and code generation, we find that answers with smaller cross-contextual shifts are more likely to be correct or factual. We demonstrate that C3 provides a complementary axis of evaluation and can serve as a benchmark usefulness diagnostic, identifying which portions of a benchmark remain informative even when aggregated scores are widely considered "saturate".

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑