arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2603.00540cs.AI

LOGIGEN: 逻辑驱动的可验证代理任务生成

LOGIGEN: Logic-Driven Generation of Verifiable Agentic Tasks

Yucheng Zeng, Weipeng Lu, Linyun Liu, Shupeng Li, Zitian Qu, Chenghao Zhu, Shaofei Li, Zhengdong Tan, Mengyue Liu, Haotian Zhao, Zhe Zhou, Jianmin Wu

首次发表 更新
浏览论文内容

中文总结 AI 辅助

LOGIGEN通过逻辑驱动和验证训练生成可验证的复杂任务,提升代理在复杂环境中的任务完成率。

中文摘要 AI 辅助

大语言模型(LLMs)从静态指令跟随者发展为自主代理,需要在复杂、具有状态的环境中操作以实现精确的状态转换目标。然而,这一范式受到数据稀缺的限制,因为现有的以工具为中心的逆向合成管道无法捕捉现实世界应用的严格逻辑。我们引入LOGIGEN,一个逻辑驱动的框架,该框架基于三个核心支柱合成可验证的训练数据:硬编译策略基础、逻辑驱动的正向合成和确定性状态验证。具体而言,采用三重代理协作:架构师将自然语言策略编译为数据库约束以强制硬性规则;集合设计师初始化边界相邻状态以触发关键策略冲突;探索者搜索此环境以发现因果解决方案路径。该框架产生了一个包含20,000个复杂任务的跨8个领域的数据集,其中有效性通过检查精确状态等价性严格保证。此外,我们提出了一种基于验证的训练协议,其中在可验证轨迹上进行监督微调(SFT)以确保符合硬编译策略,而由密集状态奖励引导的强化学习(RL)则优化长周期目标实现。在τ²-Bench上,LOGIGEN-32B(RL)实现了79.5%的成功率,显著优于基线模型(40.7%)。这些结果表明,结合逻辑驱动合成和基于验证的训练可以有效构建下一代代理所需的因果有效轨迹。

英文摘要

The evolution of Large Language Models (LLMs) from static instruction-followers to autonomous agents necessitates operating within complex, stateful environments to achieve precise state-transition objectives. However, this paradigm is bottlenecked by data scarcity, as existing tool-centric reverse-synthesis pipelines fail to capture the rigorous logic of real-world applications. We introduce \textbf{LOGIGEN}, a logic-driven framework that synthesizes verifiable training data based on three core pillars: \textbf{Hard-Compiled Policy Grounding}, \textbf{Logic-Driven Forward Synthesis}, and \textbf{Deterministic State Verification}. Specifically, a Triple-Agent Orchestration is employed: the \textbf{Architect} compiles natural-language policy into database constraints to enforce hard rules; the \textbf{Set Designer} initializes boundary-adjacent states to trigger critical policy conflicts; and the \textbf{Explorer} searches this environment to discover causal solution paths. This framework yields a dataset of 20,000 complex tasks across 8 domains, where validity is strictly guaranteed by checking exact state equivalence. Furthermore, we propose a verification-based training protocol where Supervised Fine-Tuning (SFT) on verifiable trajectories establishes compliance with hard-compiled policy, while Reinforcement Learning (RL) guided by dense state-rewards refines long-horizon goal achievement. On $τ^2$-Bench, LOGIGEN-32B(RL) achieves a \textbf{79.5\% success rate}, substantially outperforming the base model (40.7\%). These results demonstrate that logic-driven synthesis combined with verification-based training effectively constructs the causally valid trajectories needed for next-generation agents.

发表机构

  • Baidu Inc.(百度公司)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

↑