arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过自我发现的规范进行智能体上下文学习

Agentic Context Learning with Self-Discovered Specification

Jike Zhong, Ming Li, Yuxiang Lai, Ziyan Yang, Jingyu Xie, Jihyung Kil, Zheda Mai, Shao-Yuan Lo, Ren Xiang, Konstantinos Psounis, Yuanyuan Lei

arXiv 2607.09794首次发表:更新:

发表机构

University of Southern California; University of Florida; Emory University; Adobe Research; The Ohio State University; National Taiwan University(南加州大学; 佛罗里达大学; 埃默里大学; Adobe研究院; 俄亥俄州立大学; 国立台湾大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究上下文学习困难的原因,发现其不仅需内容获取还需规范获取。设计PSCI干预措施,提取并强化局部规范,在多个模型上取得显著提升,验证了规范获取的重要性,表明上下文学习依赖内容与规范获取。

AI 中文摘要

上下文学习是一种新兴的推理时任务,大型语言模型(LLMs)必须从预训练中不存在的复杂上下文中学习并应用新颖的、特定于任务的知识;即使是前沿模型的任务成功率也低于24%。本文进行了全面的实证研究以理解为何此设置仍很困难。一个自然假设是失败源于内容获取,但在CL - Bench这个广泛的上下文学习基准上的十二个检索、反思和验证基线中,我们发现相较于直接的全上下文提示,收益有限。进一步的失败分析揭示,与典型的长上下文任务不同,上下文学习不仅需要恢复局部内容,还需要获取局部规范,这些规范在查询中通常未明确指定但分布在上下文中。在所有31592个评分项目中,55.4%明确评估规范获取,而只有22.6%评估内容获取。此外,尽管76.7%的规范在用户查询中未指定,但95.5%可追溯到上下文,表明这些是可学习的义务而非隐藏要求。为验证此诊断,我们设计了一个故意简单的干预措施PSCI(私有规范 - 合同归纳),它提取局部规范并通过对抗性检查和修复来执行;在CL - Bench上,PSCI使用GPT - 5.1达到了28.14%的最新水平(绝对提高5.59个百分点,相对提高24.8%),在Qwen3.5 - 27B(提高5.28个百分点)和Gemini 3 Pro(提高6.17个百分点)上也得到了复制。十七个消融实验进一步分离了特定于任务的规范的作用。总体而言,我们的结果表明上下文学习不仅取决于内容获取,还取决于规范获取。

英文摘要

Context learning is an emerging inference-time task where LLMs must learn and apply novel, task-specific knowledge from intricate contexts absent from pre-training; even frontier models score under 24% task success. In this work, we conduct a comprehensive empirical study to understand why this setting remains difficult. A natural hypothesis is that failures stem from content access; yet across twelve retrieval, reflection, and verification baselines on CL-Bench, an extensive context learning benchmark, we find limited gains over direct full-context prompting. Further failure analysis reveals a key finding: unlike typical long-context tasks such as long document understanding, context learning requires not only recovering local content but also acquiring local specifications that are often unspecified in the query but distributed across the context: domain-specific formats, local rules, and completeness conditions. Across all 31,592 rubric items, we find that 55.4% clearly evaluate specification acquisition, while only 22.6% evaluate content acquisition. Moreover, despite 76.7% of specifications being unspecified in the user query, 95.5% are traceable to the context, indicating these are learnable obligations rather than hidden requirements. To validate this diagnosis, we design a deliberately simple intervention PSCI (private specification-contract induction) which extracts local specifications and enforces them through adversarial checking and repair; PSCI achieves state-of-the-art 28.14% with GPT-5.1 (+5.59 pp absolute and +24.8% relative) on CL-Bench, replicated on Qwen3.5-27B (+5.28 pp) and Gemini 3 Pro (+6.17 pp). Seventeen ablations further isolate the role of task-specific specifications. Overall, our results suggest context learning hinges on not only content acquisition but also specification acquisition.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑