arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GRAIN:通过不变性奖励的智能体强化学习弥合真实世界图推理中的名称与叙事变化

GRAIN: Bridging Name and Narrative Shifts in Real-World Graph Reasoning through Invariance-Rewarded Agentic RL

Zike Yuan, Han Zhang, Jianzhi Yan, Le Liu, Cai Ke, Huozhi Zhou, Jian Xie, Jiran Yin, Yukun Cao, Yue Yu, Hui Wang, Ming Liu, Bing Qin

arXiv 2608.27142首次发表:更新:

发表机构

Harbin Institute of Technology, Shenzhen; Peng Cheng Laboratory; Xidian University(哈尔滨工业大学(深圳); 鹏城实验室; 西安电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对LLMs在真实世界图推理中对节点标识符与任务表述变化脆弱的问题,提出基于强化学习的单智能体框架GRAIN,结合结构不变性奖励提升鲁棒性,在准确率、延迟及分布外泛化性上优于多智能体基线。

AI 中文摘要

尽管大语言模型(LLMs)在标准化图任务中具有潜力,但它们对节点标识符和任务表述的真实世界变化仍表现脆弱。确定性图工具对这类变化具有不变性,但从噪声文本中提取拓扑结构对LLMs而言极为脆弱,它们常过拟合于表面模式。此外,通过多智能体系统缓解解析故障会产生过高的延迟。为解决这一问题,我们提出GRAIN,一种通过强化学习优化的单智能体框架。GRAIN将推理建模为语义解析与工具执行的流程,由结构不变性奖励引导。通过针对真实拓扑结构验证提取的中间图,该奖励迫使LLMs学习鲁棒的文本到结构映射,而非记忆语言人工制品。我们还引入GRIT,一个评估对这类语言变化敏感性的基准。GRAIN在准确率上较多智能体基线高出16.45%,延迟降低约24%。此外,它展现出更优的结构泛化性,将SFT模型的分布外(OOD)差距减半(从15.77%降至7.80%),并在训练分布之外的大规模图上保持鲁棒性。

英文摘要

Despite their potential in standardized graph tasks, Large Language Models (LLMs) remain brittle to real-world shifts in node identifiers and task formulation. While deterministic graph tools are invariant to such shifts, extracting topological structures from noisy text is highly fragile for LLMs, which often overfit to surface patterns. Moreover, mitigating these parsing failures via multi-agent systems incurs prohibitive latency. To address this, we propose GRAIN, a single-agent framework optimized via reinforcement learning. GRAIN models reasoning as a semantic parsing and tool-execution pipeline, guided by a Structure Invariance Reward. By validating extracted intermediate graphs against ground-truth topologies, this reward forces the LLM to learn robust text-to-structure mappings rather than memorizing linguistic artifacts. We also introduce GRIT, a benchmark evaluating sensitivity to such linguistic shifts. GRAIN outperforms multi-agent baselines by 16.45\% in accuracy with approximately 24\% lower latency. Furthermore, it demonstrates superior structural generalization, halving the out-of-distribution (OOD) gap of SFT models (from 15.77\% to 7.80\%) and maintaining robustness on large-scale graphs beyond the training distribution.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑