AI 中文总结
Emotion2Skill框架提取LLM内部情绪向量,将其融入技能选择与演化,在WebShop、ALFWorld上用Qwen3模型实现显著性能提升,验证了情绪表征作为智能体决策信号的有效性。
AI 中文摘要
基于技能的大语言模型(LLM)智能体从外部库中选择可复用程序来解决复杂任务,但其路由决策完全依赖文本级信号,如任务描述、口头反思和经验推导的规则,而模型自身的内部表征状态仍未被利用。近期可解释性研究表明,LLM会维持线性情绪表征,且该表征会对行为产生因果影响;然而,这些表征仅被用于事后分析或直接输出引导,尚未被用于为智能体层面的决策提供依据。我们提出Emotion2Skill框架,该框架可提取LLM内部的情绪向量,并将其同时融入技能选择与技能演化过程。在每个决策步骤中,从残差流中提取27维情绪状态,并将其映射为置信度门控摘要,注入路由提示中。除在线选择外,还会分析情绪轨迹的突发内部状态变化,以精准定位有问题的技能调用,指导针对性的标准操作程序(SOP)重写,取代了先前方法中粗糙的二元结果信号。在WebShop和ALFWorld基准测试中,采用Qwen3-8B的Emotion2Skill比零样本基线的成功率提升26.9%,平均成功率提升25.5%,在两个基准测试中均优于所有基线,且在Qwen3-14B上也保持一致的提升。共激活分析进一步揭示了语义连贯的情绪-技能配对,证实路由改进反映的是有意义的内部状态信号,而非不透明的统计相关性。这些结果表明,LLM内部的情绪表征可作为编排智能体技能系统的有效决策级信号,将其效用从可解释性和输出引导拓展至决策层面。代码可在该https网址获取。
英文摘要
Skill-based LLM agents select reusable procedures from an external library to solve complex tasks, yet their routing decisions rely entirely on text-level signals such as task descriptions, verbal reflections, and experience-derived rules, while the model's own internal representational state remains unobserved. Recent interpretability work has shown that LLMs maintain linear emotion representations that causally influence behavior; however, these representations have been exploited only for post-hoc analysis or direct output steering, and have not been used to inform agent-level decision-making. We propose Emotion2Skill, a framework that extracts LLM-internal emotion vectors and incorporates them into both skill selection and skill evolution. At each decision step, a 27-dimensional emotion state is extracted from the residual stream and mapped to a confidence-gated summary injected into the routing prompt. Beyond online selection, emotion trajectories are analyzed for abrupt internal-state shifts to pinpoint problematic skill invocations, guiding targeted SOP rewriting that replaces the coarse binary outcome signal of prior methods. On WebShop and ALFWorld, Emotion2Skill with Qwen3-8B improves over the Zero-Shot baseline by +26.9% success rate and +25.5% average success respectively, outperforming all baselines on both benchmarks with consistent gains on Qwen3-14B. Co-activation analysis further reveals semantically coherent emotion--skill pairings, confirming that the routing improvements reflect meaningful internal-state signals rather than opaque statistical correlations. These results establish LLM-internal emotion representations as an effective decision-level signal for orchestrating agent skill systems, extending their utility beyond interpretability and output steering. The code is available at https://github.com/BoHan-LIN04/Emotion2Skill.