arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

InsightEmb:面向智能体洞察检索的动作-意图嵌入学习

InsightEmb: Learning Action-Intent Embeddings for Agentic Insight Retrieval

Tsz Ting Chung, Jiangnan Li, Jie Zhou, Mo Yu

arXiv 2608.04761首次发表:更新:

发表机构

The Hong Kong University of Science and Technology; WeChat AI, Tencent(香港科技大学; 腾讯微信人工智能)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对智能体洞察检索任务,提出对比嵌入框架InsightEmb,利用数学推理数据学习可迁移的检索几何结构,在无特定环境训练时优于现有模型,验证了匹配几何结构的跨域迁移性。

AI 中文摘要

自改进智能体从先前轨迹中积累可复用的洞察,这使得将积累的经验转化为可操作指导的检索工作愈发重要。在每个决策步骤,检索到恰当的洞察可帮助智能体朝着目标推进,我们将该设置称为智能体洞察检索。然而,现有检索方法主要对语义相似度进行建模,却忽略了检索到的洞察是否能解决智能体当前的决策瓶颈。我们提出InsightEmb,这是一种对比嵌入框架,仅利用数学推理数据学习可迁移的、面向进展的检索几何结构。InsightEmb联合学习将具体情境与抽象启发式规则对齐,并将具有相似进展结构的推理轨迹聚类。我们在动态智能体任务和静态技能检索基准上对InsightEmb进行评估。在未进行任何特定环境训练的情况下,InsightEmb在所有这些评估中均表现优于现有推理嵌入模型。这些结果表明,状态-洞察匹配的几何结构可跨域迁移,无需昂贵的特定环境监督即可利用公开推理数据进行有效训练。

英文摘要

Self-improving agents accumulate reusable insights from prior trajectories, making retrieval increasingly important for turning accumulated experience into actionable guidance. At each decision step, retrieving the right insight can help the agent progress toward its goal, a setting we refer to as agentic insight retrieval. However, existing retrieval methods primarily model semantic similarity, while overlooking whether a retrieved insight resolves the agent's current decision bottleneck. We propose InsightEmb, a contrastive embedding framework that learns transferable progress-oriented retrieval geometry using only mathematical reasoning data. InsightEmb jointly learns to align concrete situations with abstract heuristic rules and to cluster reasoning trajectories with similar progress structures. We evaluate InsightEmb on dynamic agent tasks and a static skill-retrieval benchmark. Without any environment-specific training, InsightEmb improves over all these evaluations, surpassing the performance of existing reasoning embedding models. These results suggest that the geometry of state-insight matching can transfer across domains, enabling effective training from publicly available reasoning data without expensive environment-specific supervision.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑