arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

关系响应场:黑盒大语言模型响应一致性与恢复的通用理论

Relational Response Fields: A General Theory of Black-Box LLM Response Consistency and Recovery

Song Zichen

arXiv 2608.04552首次发表:更新:

发表机构

Sungkyunkwan University(成均馆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出关系响应场(RRF)理论,定义黑盒大语言模型响应恢复的内在难度γₖ(D,A),推导稀疏修复算法,通过实验验证其在响应一致性与恢复中的有效性。

AI 中文摘要

黑盒语言模型的可靠性通常通过采样、提示、投票、验证或迭代修正单个答案来实现。我们提出一个先验问题:是什么决定了一组黑盒响应是否可恢复?我们将对查询的类型变换的响应表示为关系响应场(Relational Response Field,RRF)。边传输编码了在 paraphrase(paraphrase:复述)、缩放、分解、重构或其他任务对称性下有效响应必须如何变化;锚点编码了独立可信的证据,如执行结果或验证器。对于关系算子 D、锚点算子 A,以及至多 k 个损坏的响应节点,我们将 γₖ(D,A) 定义为黑盒响应恢复的内在难度。当且仅当每一个 k 节点损坏都可被识别时,γₖ 为正;它给出与 1/γₖ 成比例的确定性稳定性界;匹配的两点极小极大下界表明,没有任何估计器能改善这种依赖关系。因此,一致性并非真实性:仅基于关系的方法对零空间方向(包括共享幻觉)是盲目的。我们推导了稀疏场修复算法,同时将信息论可识别性与凸优化所需的更强零空间条件区分开来。受控定理测试和黑盒数学/代码实验评估了四个由理论确定的结论:一致性-真实性分离、锚点相变、冗余饱和以及修复难度的跨模型、跨任务预测。结果支持 γₖ(D,A) 作为响应恢复实例的可测量属性,而非附加给某一修复启发式的分数。

英文摘要

Black-box language-model reliability is commonly pursued by sampling, prompting, voting, verifying, or iteratively revising individual answers. We ask a prior question: \emph{what determines whether a collection of black-box responses is recoverable at all?} We represent responses to typed transformations of a query as a \emph{relational response field} (RRF). Edge transports encode how valid responses must change under paraphrase, scaling, decomposition, refactoring, or other task symmetries; anchors encode independently trusted evidence such as execution or a verifier. For relation operator $D$, anchor operator $A$, and at most $k$ corrupted response nodes, we identify $γ_k(D,A)$ as the intrinsic difficulty of black-box response recovery. It is positive exactly when every $k$-node corruption is identifiable; it gives a deterministic stability bound proportional to $1/γ_k$; and a matching two-point minimax lower bound shows that no estimator can improve this dependence. Thus consistency is not truth: relation-only methods are blind to null directions, including shared hallucinations. We derive sparse field-repair algorithms while separating information-theoretic identifiability from the stronger null-space conditions required by convex optimization. Controlled theorem tests and black-box mathematics/code experiments evaluate four theory-fixed consequences: consistency--truth separation, anchor phase transitions, redundancy saturation, and cross-model, cross-task prediction of repair difficulty. The results support $γ_k(D,A)$ as a measurable property of a response-recovery instance, rather than a score attached to one repair heuristic.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑