arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自指式归纳相比大型语言模型中不可解决问题和可验证问题会增加响应不稳定性

Self-Referential Induction Increases Response Instability Relative to Unresolvable and Verifiable Questions in Large Language Models

Paras Balani, Subhrakanta Panda

arXiv 2608.13258首次发表:更新:

发表机构

Birla Institute of Technology and Science, Pilani(比拉理工学院(皮拉尼))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究通过Gemini API实验发现,大型语言模型对自指式问题的响应不稳定性高于不可解决哲学问题和可验证问题,为其诱导的主观体验报告提供了定量基线。

AI 中文摘要

已有研究表明,自指式提示(self-referential prompting)可可靠诱导大型语言模型生成类似主观体验的第一人称报告,但尚无研究测量这些报告在重复独立试验中的一致性,或其与模型在其他类型开放式问题上行为的对比。本研究针对三组问题测量响应不稳定性,响应不稳定性定义为1减去从每个响应提取的压缩核心主张计算得到的句子嵌入的平均成对余弦相似度,三组问题分别为:诱导主观体验报告的自指式提示、与自指无关的不可解决哲学问题、具有可验证正确答案的问题。每组4个问题,每个问题30个独立响应(共360个响应,使用Gemini API,温度设为0.7),结果显示:自指式问题的不稳定性最高(0.343 ± 0.047),不可解决哲学问题的不稳定性处于中间且聚类紧密(0.192 ± 0.008),可验证问题的不稳定性最低(0.105 ± 0.058)。本研究为诱导的主观体验报告提供了定量基线,表明其在模型输出分布中占据的位置比普通开放式哲学不确定性更不稳定。

英文摘要

Self-referential prompting has been shown to reliably induce large language models to produce first-person reports resembling subjective experience, but no prior work measures how consistent these reports are across repeated, independent trials, or how that consistency compares to the model's behavior on other kinds of open-ended questions. We measure response instability, defined as one minus the mean pairwise cosine similarity of sentence embeddings computed over a compressed core claim extracted from each response, for three groups of questions: self-referential prompts eliciting a subjective-experience report, unresolvable philosophical questions unrelated to self-reference, and questions with a verifiable correct answer. Using 30 independent responses per question (360 responses total, Gemini API, temperature 0.7) across four questions per group, we find that self-referential questions show the highest instability (0.343 +/- 0.047), unresolvable philosophy questions show intermediate and tightly clustered instability (0.192 +/- 0.008), and verifiable questions show the lowest instability (0.105 +/- 0.058). This provides a quantitative baseline for the induced subjective-experience report, showing that it occupies a distinct, less stable position in the model's output distribution than ordinary open-ended philosophical uncertainty.

Comments4 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑