发表机构
Gaoling School of Artificial Intelligence; Renmin University of China(高瓴人工智能学院; 中国人民大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长周期数据分析智能体中过时分析产物被静默继承的问题,提出StateGuard框架,通过状态图外化分析进度,结合反事实监督与有效性引导策略优化,在三个基准上提升性能并减少错误传播。
AI 中文摘要
基于大语言模型的智能体在自动化数据分析方面展现出强大能力,并正日益向长周期、多阶段的分析工作流发展。然而,随着分析过程的推进,约束、变量和结论仍隐式地嵌入在交互历史中,这使得智能体难以在日益延长的周期和不断变化的依赖关系中追踪哪些分析产物仍然有效。因此,过时的产物可能被静默继承,从而将错误传播到下游阶段。为应对这一挑战,我们提出了StateGuard,一个面向长周期数据智能体的分析状态有效性管理框架。StateGuard将不断演化的分析进度外部化为一个包含约束、版本化变量、中间结论和跨状态关系的状态图,将每个状态视为一个可执行、可验证、可追踪的对象,而非仅仅是文本记忆。StateGuard通过基于证据的验证和分层干预来维护状态有效性。为使StateGuard具备这些能力,我们首先引入了面向管理器的反事实监督,通过反事实运行时合成构建了3K个以状态为中心的轨迹,以微调StateGuard用于状态维护、验证和修复。随后,我们应用有效性引导的策略优化,利用运行时有效性证据为协议正确性、状态基础性和干预质量提供细粒度的学习信号。在三个多样化的长周期数据分析基准上的实验表明,StateGuard持续提升了数据智能体的性能,同时减少了依赖引发的下游错误传播,展示了显式分析状态管理对于可靠长周期数据分析的优势。
英文摘要
LLM-based agents have shown strong capabilities in automated data analysis and are increasingly moving toward long-horizon, multi-stage analytical workflows. However, as the analytical process evolves, constraints, variables, and conclusions remain implicitly embedded in interaction histories, making it difficult for agents to track which analytical artifacts remain valid over increasingly long horizons and changing dependencies. Consequently, stale artifacts may be silently inherited, propagating errors to downstream stages. To address this challenge, we propose StateGuard, an analytical-state validity management framework for long-horizon data agents. StateGuard externalizes evolving analytical progress into a state graph containing constraints, versioned variables, intermediate conclusions, and cross-state relations, treating each state as an executable, verifiable, and traceable object rather than textual memory alone. StateGuard maintains state validity through evidence-grounded verification and hierarchical intervention. To equip StateGuard with these capabilities, we first introduce Manager-Oriented Counterfactual Supervision, which constructs 3K state-centric trajectories through counterfactual runtime synthesis to fine-tune StateGuard for state maintenance, verification, and repair. We then apply Validity-Guided Policy Optimization, using runtime validity evidence to provide fine-grained learning signals for protocol correctness, state grounding, and intervention quality. Experiments on three diverse long-horizon data-analysis benchmarks show that StateGuard consistently improves data-agent performance while reducing dependency-induced downstream error propagation, demonstrating the advantages of explicit analytical-state management for reliable long-horizon data analysis.