arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06422cs.LG

分片可防止大语言模型(LLM)的监督失效与对抗性利用

Sharding Prevents LLM Oversight Failures and Adversarial Exploitation

Victor Akinwande, J. Zico Kolter, Aran Nayebi

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对LLM监督中裁决数量增加导致一致性下降的问题,提出分片方法,可提升监督一致性、抵御对抗攻击,且分片后较弱裁判可超越更强的整体裁判。

中文摘要 AI 辅助

为大语言模型(LLM)裁判提供更多计算资源并不一定能使其检查更多要求。当一次调用必须返回多个裁决时,即便该调用获得的令牌或工具预算与一组独立调用相同,部分决策也会在证据上的依据变弱。在专家评分的研究复制、法律工作和临床试验评估中,与专家的一致性会随每次调用的裁决数量增加而下降。我们发现分片是缓解基于模型的监督中这一失效问题的干预手段:分片将要求划分为更小的组,为每组分配一次独立调用,并聚合裁决。与使用该组全部预算的单次调用相比,在保持模型、证据、总预算及每决策预算固定的情况下,分片可提升一致性。总体而言,我们发现分片后的较弱裁判可超越更强大的整体裁判,即便后者获得该组全部预算也能与之匹敌。此外,分片对对抗者具有鲁棒性:最佳N对抗者可在保持基础工作固定的情况下,仅改变其呈现方式,使过载裁判对未满足标准的接受度提高数倍;而分片在降低基线误差的同时,可消除这种对抗优势,即便对抗者的搜索范围扩大,也能将过度接受度保持在较低水平。分片无法应对那些针对每个标准分别说服裁判而非利用过载的攻击,在这种场景下,我们发现分片之上的辩论式对抗可抵御此类自适应重新优化。

英文摘要

Giving an LLM judge more compute does not necessarily make it check more requirements. When one call must return many verdicts, some decisions become weakly grounded in the evidence, even when that call receives the same token or tool budget as a panel of separate calls. Across expert-graded research replications, legal work, and clinical-trial assessments, agreement with experts falls as the number of verdicts per call grows. We identify sharding as the intervention that mitigates this failure in model-based oversight. Sharding partitions the requirements into smaller groups, assigns each group to a separate call, and aggregates the verdicts. Against a single call with the panel's full budget, sharding improves agreement while holding the model, evidence, total budget, and per-decision budget fixed. Overall, we find that a sharded weaker judge can outperform a more capable holistic judge and match that judge even when the latter receives the panel's full budget. Additionally, we find that sharding exhibits robustness against adversaries. A best-of-N adversary can hold the underlying work fixed, vary only its presentation, and increase an overloaded judge's acceptance of genuinely unmet criteria severalfold. Wherever sharding reduces baseline error, it removes this adversarial advantage, keeping over-acceptance low even as the adversary's search widens. Sharding does not address attacks that persuade the judge separately on each criterion rather than exploiting overload. In that setting, we find that debate-style opposition on top of sharding withstands such adaptive re-optimization.

↑