arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14509cs.AIcs.CLcs.LG

分工:将证据解释与决策聚合分离

Split the Labor: Separating Evidence Interpretation from Decision Aggregation

  • Atlassian(亚特兰西安公司)

机构由 AI 辅助整理,请以论文原文为准。

Zhelun Wu

AI总结:

该研究提出将证据解释与决策聚合分离,设计四字段证据元组解决计数尺度漂移问题,在纵向语料库上实现了优于基线的AUPRC性能。

AI中文摘要:

要求语言模型从多个来源得出结论的系统通常会将这些来源串联成一个提示。这混淆了两个具有不同要求的操作:解释一个来源需要能力和上下文,而组合解释则需要固定的运算、跨实例的可比性以及弃权(不执行)的选项。一旦将二者分离,设计问题就变成了它们之间的接口。我们提出了一个四字段证据元组(假设、可靠性桶、理由、来源),并表明固定该元组可确定两部分的设计。这种分离还揭示了此类系统组合时的一种失败模式,我们称之为计数尺度漂移:对未归一化权重之和设置阈值恰好是后验阈值处理,但该操作点会随所参考的来源数量而滑动,且滑动幅度随读者可靠性增大而增大。当来源可靠性存在差异时,投票规则与后验会对实例产生不同排序,且没有阈值可协调二者。汇集校准后的对数似然比可解决这两个问题,该解决方案属于运算层面而非架构层面,适用于语言模型之外的一类规则:得分求和的分诊引擎、通过统计阳性结果评分的诊断面板以及加法多信号检测器。随后,我们在一个纵向语料库上两次实例化该原则,一次在结果明确后,一次在结果明确前。相同的划分在两种场景下均适用,且粒度不同:前者针对阅读过程,后者针对学习能力。在后者场景中,一个基于简易辅助目标的小型序列编码器,加上携带删失生存损失的树集成,达到了0.921的AUPRC,而手工设计的基线仅为0.805。我们分离了可迁移的部分与必须按领域重新估计的部分,并给出了五项可证伪该框架的预测、三项负面结果以及哪些比较仍存在混淆。

英文摘要:

Systems that ask a language model to reach a conclusion from many sources usually concatenate them into one prompt. This conflates two operations with different requirements. Interpreting a source rewards capacity and context. Combining interpretations rewards fixed arithmetic, comparability across instances, and the option to return nothing. Once separated, the design problem becomes the interface between them. We propose a four-field evidence tuple (hypothesis, reliability bucket, rationale, provenance) and show that fixing it determines both halves. The separation also reveals a failure mode in how such systems combine, which we call count-scale drift. Thresholding a sum of unnormalized weights is exactly posterior thresholding, but at an operating point that slides with the number of sources consulted. The slide grows with reader reliability. When source reliabilities differ, the vote rule and the posterior order instances differently, and no threshold reconciles them. Pooling calibrated log-likelihood ratios addresses both problems. The fix is arithmetic rather than architectural, and applies to a class of rules beyond language models: score-summing triage engines, diagnostic panels scored by counting positives, and additive multi-signal detectors. We then instantiate the principle twice on one longitudinal corpus, once after outcomes resolve and once before. The same partition helps in both, at different granularities: over reading in the first, over learning capacity in the second. There, a small sequence encoder on an easy auxiliary objective plus a tree ensemble carrying the censored survival loss reaches 0.921 AUPRC against 0.805 for a hand-crafted baseline. We separate what transfers from what must be re-estimated per domain, and state five predictions that would falsify the framework, three negative results, and which comparisons remain confounded.

补充信息

↑