arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01559cs.AIcs.CLcs.LG

对抗自我博弈的竞争成分能否提升法律推理?一项受控负面结果

Does the Competitive Component of Adversarial Self-Play Improve Legal Reasoning? A Controlled Negative Result

Miseog Shawn Kim

首次发表
浏览论文内容

中文总结 AI 辅助

该研究通过四项测试及一项试点发现,对抗自我博弈的竞争成分未为法律推理模型带来可靠增益,其价值在于揭示相关陷阱并验证多教师课程的核心价值源于可验证环境而非竞争。

中文摘要 AI 辅助

对抗自我博弈是一种颇具吸引力的法律推理训练思路:让学生模型撰写论证,让对手模型攻击该论证,若学生模型的论证在攻击中存活则给予奖励。我们设计了这样的训练信号——一种可验证的“存活”奖励,其中学生模型引用的权威资料与对手模型的反权威资料均由引用验证器核查,因此存活判定基于可验证的依据而非修辞,伪造的引用会被自动中和。随后我们提出一个狭窄但重要的问题:竞争成分本身——即对手模型与存活奖励——是否在完全相同的非竞争训练运行基础上带来增益?我们通过四项独立测试及一项后续试点研究验证:四项测试包括自助法比较、双随机种子复现、按案例配对的对抗鲁棒性比较,以及对生成论证的盲法头对头评判;后续试点研究采用刻意强化的自我博弈对手。结果显示竞争成分未产生可靠增益:盲法评判中竞争模型胜率为49%(二项式p值约1.000);强化对手的试点研究中胜率为50%(32:32,p值约1.000);早期看似29%的优势经证实为小样本假象。我们报告这一真实负面结果,论文的价值在于可复现性及对具体陷阱的分享:一项最初有前景的指标在更多数据下反转,以及一项对抗鲁棒性指标在对手停止引用与标准答案相同的权威资料后,悄然崩溃为普通召回率。这一零结果与配套编码领域研究(Kim, 2026, arXiv:2607.08255)的结论一致,且在法律领域再次验证:多教师课程的价值源于构建可验证环境,而非竞争本身。

英文摘要

Adversarial self-play is an appealing recipe for legal reasoning: have a student model draft an argument, have an adversary attack it, and reward the student when its argument survives the attack. We designed exactly such a training signal -- a verifiable "survival" reward in which both the student's cited authorities and the adversary's counter-authorities are checked by a citation verifier, so that survival is decided on verified grounds rather than rhetoric, and fabricated citations are automatically neutralized. We then asked a narrow but important question: does the competitive component itself -- the adversary and the survival reward -- add anything on top of an otherwise identical non-competitive training run? Across four independent tests -- a bootstrap comparison, a two-seed replication, a paired per-case adversarial-robustness comparison, and a blinded head-to-head judgment of generated arguments, plus a follow-up pilot with a deliberately strengthened self-play adversary -- the competitive component produced no reliable benefit. The blinded judgment gave a 49% win rate (binomial p approx. 1.000); the strengthened-adversary pilot gave a 50% win rate (32:32, p approx. 1.000). An early apparent +29% advantage reversed and proved to be a small-sample artifact. We report this as an honest negative result. The value of the paper is reproducibility and the sharing of concrete pitfalls: an initially promising metric that inverted on more data, and an adversarial-robustness metric that silently collapsed to plain recall once the adversary stopped citing the same authorities as the gold answer. This null is consistent with, and reconfirms in the legal domain, the conclusion of the companion coding-domain study (Kim, 2026, arXiv:2607.08255) that the value of multi-teacher curricula arises from constructing a verifiable environment rather than from competition itself.

↑