AI 中文总结
研究智能体中TextGrad应用问题,通过区分遵循与学习策略能力来研究。发现两者有差距,人工编写策略能提升智能体表现,而从轨迹生成的策略无法可靠超越固定提示,主要挑战是从经验中生成和选择策略。
AI 中文摘要
TextGrad通过根据反馈修改文本改进语言模型系统。其核心观点是自然语言反馈可作为优化文本组件的梯度而不改变模型权重。将其应用于智能体更难,因为反馈在一系列行动后才出现,难以确定哪个决策导致失败。我们通过区分遵循有用策略的能力和从经验中学习该策略的能力来研究此问题。主要发现是这两种能力之间存在明显差距。人工编写的策略使两个冻结的7B智能体在TextWorldExpress上成功率提高5.0个百分点,表明存在有用的策略文本。然而,从智能体轨迹生成的策略即使有更丰富的痕迹、反事实证据或迭代GEPA搜索,也不能可靠地优于固定提示。因此,智能体级别的TextGrad的主要挑战不是执行文本策略更新,而是从经验中可靠地生成和选择它们。
英文摘要
TextGrad improves language-model systems by revising text from feedback. Its core thesis is that natural-language feedback can act as a gradient for optimizing text components without changing model weights. Applying it to agents is harder because feedback arrives only after a sequence of actions, making it difficult to identify which decision caused failure. We study this problem by separating the ability to follow a useful policy from the ability to learn that policy from experience. Our main finding is a clear gap between these two abilities. Human-written policies improve two frozen 7B agents on TextWorldExpress by 5.0 success points, showing that useful policy text exists. However, policies generated from agent trajectories do not reliably outperform fixed prompting, even with richer traces, counterfactual evidence, or iterative GEPA search. The main challenge for agent-level TextGrad is therefore not executing textual policy updates, but reliably generating and selecting them from experience.