NxN E-评估:通过共形CRT空假设进行假设验证
NxN E-valuation: Hypothesis Certification via a Conformal CRT Null
- Meta
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出NxN E-评估算法,利用大型训练集让样本互为空假设,实现共形CRT以验证LLM的假设,可作为LLM循环验证和保留数据测试的更优替代方案。
AI中文摘要:
我们提出了NxN E-评估,一种便捷的基于e值的假设验证算法,只要有足够大的数据集,该算法无需构建任何特定于案例的验证程序(例如构建专用的空假设)即可验证假设。该方法特别适用于基于大语言模型(LLM)的探索系统,在这类系统中,LLM非常擅长提出假设,但存在严重的幻觉问题;这种幻觉使我们无法直接利用LLM的输出,而现有的补救措施均存在不足。最常见的解决方案包括让LLM进行自我验证或修正(循环验证)以及保留测试(错误假设仍可能通过虚假关联通过),此外还有引言中详述的其他补救措施。为解决这一问题,NxN E-评估利用自然存在的大型训练集,让不同样本互为空假设。该设计直接实现了用于验证每个假设的条件随机测试(CRT)。只要LLM生成的是适用于每个单独样本的假设,该方法至少可以作为LLM循环验证和保留数据测试的更优通用替代方案。
英文摘要:
We propose NxN E-valuation, a handy, e-value-based hypothesis-certification algorithm that lets a hypothesis be verified without building any case-specific certification procedure---such as constructing a dedicated null hypothesis---as long as a large enough dataset is available. The method is especially suited to LLM-based exploration systems, where LLMs are remarkably good at proposing hypotheses but suffer badly from hallucination; this hallucination prevents us from harvesting LLM outputs directly, and existing remedies each fall short. The most common solutions include letting the LLM verify or correct itself circular verification and held-out testing (where false hypotheses can still pass via spurious correlations), among other remedies detailed in the introduction. To resolve this, NxN E-valuation exploits the naturally existing large training set and lets different samples serve as null hypotheses for one another. This design directly realizes a conditional randomization test (CRT) that certifies each hypothesis. The approach can be a universally better replacement for at least LLM circular verification and held-out-data testing, provided the LLM's generations are hypotheses that apply to each individual sample.