arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

长时程法律推理中的测试时智能体演化

Test-Time Agent Evolution for Long-Horizon Legal Reasoning

Haotian Chen, Shuaicheng Niu, Haocong Rao, Kaisong Song, Jun Lin, Lizhen Cui, Zhiqi Shen, Yonghui Xu

arXiv 2610.08138首次发表:更新:

发表机构

Joint SDU–NTU Centre for Artificial Intelligence Research, Shandong University; Nanyang Technological University; Alibaba Group(山东大学; 南洋理工大学; 阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长时程法律推理中案件异质性和跨角色依赖的挑战,提出无需训练的测试时智能体自适应方法,通过测试时记忆演化和基于评分标准的协作,在多个基准上显著提升决策可靠性与效率。

AI 中文摘要

法律智能旨在支持跨长时程法律流程的可靠决策,这些流程涉及不断演变的案件状态和多个角色。然而,现实世界的法律部署在事实、证据和程序背景方面表现出显著的案件异质性,暴露了静态智能体策略的局限性。此外,法律推理在角色和程序阶段之间本质上相互依赖,使得全局可靠性根本不同于孤立的角色能力。为了应对这些挑战,我们研究了无需训练的自适应测试时智能体调整,其中智能体在不更新模型参数的情况下,持续利用来自先前案件和持续交互的部署时信号。我们提出了\method,引入了\emph{测试时记忆演化},以从先前案件中检索可复用的经验,将其适应当前的事实和程序背景,并为后续决策巩固积累的经验。此外,\emph{基于评分标准的协作}根据行为和程序要求验证和修订角色特定动作,实现跨角色和跨阶段的协调决策。在J1-EVAL和LegalWorld上使用五个骨干模型进行的广泛实验表明,与代表性的推理和智能体基线相比,该方法在合理的交互和计算成本下实现了持续改进。消融研究和案例研究进一步表明,这两个组件在经验适应和跨角色协调方面提供了互补的益处,提高了长时程法律推理的可靠性和效率。

英文摘要

Legal intelligence aims to support reliable decision-making across long-horizon legal processes involving evolving case states and multiple roles. However, real-world legal deployment exhibits substantial case heterogeneity in facts, evidence, and procedural contexts, exposing the limitations of static agent strategies. Moreover, legal reasoning is inherently interdependent across roles and procedural stages, making global reliability fundamentally different from isolated role competence. To address these challenges, we study training-free test-time agent adaptation, where agents continuously exploit deployment-time signals from preceding cases and ongoing interactions without updating model parameters. We propose \method, which introduces \emph{Test-Time Memory Evolution} to retrieve reusable experience from previous cases, adapt it to the current factual and procedural context, and consolidate accumulated experience for subsequent decision-making. Further, \emph{Rubric-Aligned Collaboration} verifies and revises role-specific actions according to behavioral and procedural requirements, enabling coordinated decision-making across roles and stages. Extensive experiments on J1-EVAL and LegalWorld across five backbone models demonstrate consistent improvements over representative reasoning and agent baselines with reasonable interaction and computational costs. Ablation and case studies further show that the two components provide complementary benefits in experience adaptation and cross-role coordination, improving the reliability and efficiency of long-horizon legal reasoning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑