StateWise:在智能体行动前诊断并修复持久化操作状态
StateWise: Diagnosing and Repairing Persistent Operational State Before Agent Actions
浏览论文内容
中文总结 AI 辅助
StateWise通过反事实重规划识别决策关键记录,结合可靠性检查和类型化证据锚定修复失效的持久化状态,使LLM智能体在行动前恢复正确性,实验显示正确率从38.7%提升至93.3%。
中文摘要 AI 辅助
LLM智能体结合推理、工具使用和持久化记忆,通过将存储的操作记录作为后续行动的推理前提,支持跨任务的工作。然而,环境或需求的变化可能使这些记录失效,而现有的行动审查、来源追踪和澄清机制可能无法纠正底层持久化状态。我们对编码智能体轨迹的审计识别出候选故障链,其中无效记录被重用,导致任务失败和不安全的修改。我们提出StateWise,一个在行动执行前诊断和修复持久化操作状态的框架。StateWise使用记录级反事实重规划来识别决策关键记录,然后通过可靠性检查、机器可观察事实的只读验证以及针对开发者所有意图的定向澄清来确定其当前有效性。类型化证据锚定将证据绑定到特定记录和范围,从而实现带修复谱系的持久化纠正。随后,智能体从修复后的状态重新规划,并在执行前进行独立的状态-行动检查。我们在150个可执行编码智能体案例上评估StateWise,涵盖多样化的运行时环境、工作区配置和仓库设置,并辅以跨模型评估。在持久化状态被破坏的情况下,StateWise实现了93.3%的总体正确率,而基线智能体为38.7%,且没有不安全行动。组件消融、多任务实验和迁移评估进一步证明了有效的恢复、持久化纠正以及跨仓库和工具接口的可迁移性。
英文摘要
LLM agents combine reasoning, tool use, and persistent memory to support work across tasks by reusing stored operational records as premises for later actions. However, environmental or requirement changes can invalidate these records, while existing action review, provenance tracking, and clarification mechanisms may leave the underlying persistent state uncorrected. Our audit of coding-agent trajectories identifies candidate failure chains in which invalid records are reused, leading to task failures and unsafe modifications. We propose StateWise, a framework for diagnosing and repairing persistent operational state before action execution. StateWise uses record-level counterfactual replanning to identify decision-critical records, then establishes their current validity through reliability checks, read-only verification of machine-observable facts, and targeted clarification of developer-owned intent. Typed evidence grounding binds evidence to specific records and scopes, enabling persistent corrections with repair lineage. The agent then replans from the repaired state, followed by an independent state-action check before execution. We evaluate StateWise on 150 executable coding-agent cases across diverse runtime environments, workspace configurations, and repository settings, complemented by cross-model evaluations. Under corrupted persistent state, StateWise achieves 93.3% overall correctness, compared with 38.7% for the baseline agent, with no unsafe actions. Component ablations, multi-task experiments, and transfer evaluations further demonstrate effective recovery, persistent corrections, and transferability across repositories and tool interfaces.
发表机构
- Xi’an Jiaotong University(西安交通大学)
- Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。