DynaContext:面向异构参数提取的优化提示的自改进动态上下文化
DynaContext: Self-Improving Dynamic Contextualization of Optimized Prompts for Heterogeneous Parameter Extraction
浏览论文内容
中文总结 AI 辅助
DynaContext框架结合离线优化提取核心与推理时上下文适配及验证门控自改进,在异构参数提取任务中,较静态提示管道显著提升了准确率与F1值。
中文摘要 AI 辅助
自动化提示与技能优化通常会生成单个静态指令,在下次优化周期前会被重复用于所有推理实例,但当不同实例所需的上下文、约束和证据存在差异时,该方法无法适配。例如,从电子元件描述中提取参数就不符合这一假设:电阻器、电容器、晶体管和连接器需要不同的字段、单位约束和示例,且每个输入提供的证据状态各不相同。我们提出DynaContext,这一框架将通过GEPA或SkillOpt学习的离线优化提取核心,与推理时的上下文适配及验证门控自改进相结合。DynaContext将每个条目路由至内部、外部或后备证据路径,并从核心、模式、证据、未解决字段和已验证示例中组合出特定条目的提示。确定性验证和大语言模型(LLM)评判对所有输出进行门控,不确定的案例会被提交至人工审核,且仅有人工验证的修正会进入示例记忆。在单类别基准上,基础提示的平均准确率为86.6%,独立SkillOpt的准确率提升至96.9%,最佳DynaContext配置的准确率达98.6%;在850个异构黄金参数事实中,平均字段级F1值从无优化、无示例的对照组的51.8%,仅使用动态示例提升至59.2%,仅使用优化核心提升至66.9%,同时使用两者提升至71.0%;在模型固定的情况下,完整配置比已部署的静态提示管道平均高出17.3个F1点。
英文摘要
Automated prompt and skill optimization typically produces a single static instruction that is reused across inference instances until the next optimization cycle. However, this approach cannot adapt when the required context, constraints, and evidence vary from one instance to another. For instance, parameter extraction from electronic component descriptions breaks this assumption: resistors, capacitors, transistors, and connectors require different fields, unit constraints, and demonstrations, and each input provides a different evidence state. We introduce DynaContext, a framework that combines an offline-optimized extraction core, learned with GEPA or SkillOpt, with inference-time contextual adaptation and validation-gated self-improvement. DynaContext routes each item through internal, external, or fallback evidence paths and composes an item-specific prompt from the core, schema, evidence, unresolved fields, and validated demonstrations. Deterministic validation and an LLM judge gate every output, uncertain cases go to human review, and only human-verified corrections enter the demonstration memory. On a single-category benchmark, average accuracy increases from 86.6% for the base prompt to 96.9% for standalone SkillOpt and 98.6% for the best DynaContext configuration. Across 850 heterogeneous gold parameter facts, average field-level F1 increases from 51.8% for an unoptimized, demonstration-free control to 59.2% with dynamic demonstrations alone, 66.9% with the optimized core alone, and 71.0% with both. Holding the model fixed, the full configuration outperforms the deployed static-prompting pipeline by 17.3 F1 points on average.