发表机构
The Chinese University of Hong Kong; Macao Polytechnic University; Xi’an Jiaotong University; University of Science and Technology of China(香港中文大学; 澳门理工大学; 西安交通大学; 中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对大型语言模型上下文学习的脆弱性问题,提出上下文编译架构(CCA),在CL-bench上优于普通提示及两种基线方法,提升了Kimi K2.5的任务完成率。
AI 中文摘要
大型语言模型(LLMs)越来越多地处理上下文学习(ICL)任务,其中一段长而新颖的上下文为一系列问题定义了规则、知识和输出模式。在对上下文的每个细节进行评分的基准上,即使是强大的开源权重模型也只能通过12%-16%的任务:忽略单个规则就会导致整个响应失败。我们认为这种脆弱性是结构性的:主流的“读取-推理”范式要求模型在一次前向传播中完成提取、规划、生成和自验证。因此,我们探究显式上下文编译能否解决该问题、其与现有长上下文策略(要点检索、多智能体自博弈)的比较情况,以及所得框架的优势在任务结构和模型规模上的适用范围。我们提出上下文编译架构(CCA),其核心创新是带有固定槽位的类型化中间表示(IR),固定槽位包括规则(必须执行、禁止执行、条件)、输出规格、可用工具、数据概况,任何散文式上下文都可编译到这些槽位中;下游则包含可执行验证器和受违规门控的修正循环。在CL-bench(4个开源基础模型共1899项任务)上,CCA在每个基础模型上的表现均优于普通提示和两个长上下文基线(ReadAgent-P、Ctx2Skill),将Kimi K2.5的准确率从15.4%提升至21.4%,增益集中在规则密集的子类别。代码和缓存完成结果可在该https URL获取。
英文摘要
Large language models (LLMs) increasingly handle in-context learning (ICL) tasks where a long, novel context defines the rules, knowledge, and output schema for a series of questions. On benchmarks that grade against every detail of the context, even strong open-weights models pass only 12-16% of tasks: a single overlooked rule fails the whole response. We argue this brittleness is structural: the dominant "read-and-reason" paradigm asks the model to extract, plan, generate, and self-verify in one forward pass. We therefore ask whether explicit context compilation can fix it, how it compares to existing long-context strategies (gist retrieval, multi-agent self-play), and where the resulting harness benefit holds across task structure and model scale. We propose the Context Compilation Architecture (CCA), whose central novelty is a typed intermediate representation (IR) with fixed slots (rules.{must_do, must_not, conditional}, output_spec, available_tools, data_profile) into which any prose context is compiled once; executable verifiers and a violation-gated correction loop follow as downstream consequences. On CL-bench (1,899 tasks across 4 open base models), CCA outperforms vanilla prompting and two long-context baselines (ReadAgent-P, Ctx2Skill) on every base model, lifting Kimi K2.5 from 15.4% to 21.4% with gains concentrated on rule-dense sub-categories. Code and cached completions are available at https://github.com/TonyQJH/cca-emnlp2026.
CommentsAccepted to EMNLP 2026 (Findings). Code, data, and cached completions available at https://github.com/TonyQJH/cca-emnlp2026