EvolveScaler:通过可执行状态机与自然语言渲染合成信息演化上下文
EvolveScaler: Synthesizing Information-Evolution Contexts via Executable State Machines and Natural-Language Rendering
- Tencent(腾讯)
- The Chinese University of Hong Kong(香港中文大学)
- Renmin University of China(中国人民大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
EvolveScaler通过代码驱动方式定义信息演化并渲染为自然语言,生成多轮事件历史与参考答案,提供挑战性评估和可迁移训练监督。
AI中文摘要:
在持久交互中,长上下文可能编码的是一个演化过程而非固定记录:后续事件可以修订或撤销早期信息,从而改变哪些信息仍然有效以及可以得出哪些结论。我们将这种设置称为信息演化(IE)。解决IE问题需要识别有效记录、按顺序应用更新,并从事件历史中重建与查询相关的状态。现有的文本优先合成流水线使得此类数据难以验证,因为状态转换和答案逻辑仍然是隐式的。我们引入了EvolveScaler,一个代码驱动的框架,在将信息演化渲染为自然语言之前先对其进行定义。人工编写的操作规范定义了状态转换、记录有效性、难度控制和可执行的答案逻辑;然后,一个强大的LLM根据每个规范合成一个自包含的模拟器。执行经过验证的模拟器会产生自然语言的多轮事件历史,而确定性重放则计算参考答案和原子检查清单。我们使用117个任务原型和159个最终问题操作符实例化了EvolveScaler,涵盖五个难度级别,每个实例大约包含7到1,200个事件,产生了约35,100个训练示例和585个经过验证的评估实例。在very_long层级上,最强的模型达到了59.3%的avg@5,而六个模型的得分低于10%。在一个内部A3B模型上使用6,000个EvolveScaler示例进行训练,在所有八个独立构建的分布外基准上,其性能相对于基础检查点都有所提升,平均提升了5.25个百分点。这些结果表明,代码驱动的IE合成既提供了具有挑战性的评估,也提供了可迁移的训练监督。
英文摘要:
In persistent interactions, long contexts may encode an evolving process rather than a fixed record: later events can revise or revoke earlier information, changing what remains valid and what conclusions follow. We call this setting information evolution (IE). Solving IE requires identifying valid records, applying updates in order, and reconstructing the query-relevant state from the event history. Existing text-first synthesis pipelines make such data difficult to verify because state transitions and answer logic remain implicit. We introduce EvolveScaler, a code-driven framework that defines information evolution before rendering it as natural language. Human-authored operational specifications define state transitions, record validity, difficulty controls, and executable answer logic; a strong LLM then synthesizes a self-contained simulator from each specification. Executing validated simulators produces natural-language multi-turn event histories, while deterministic replay computes reference answers and atomic checklists. We instantiate EvolveScaler with 117 task prototypes and 159 final-question operators across five difficulty levels spanning approximately 7 to 1,200 events per instance, yielding about 35,100 training examples and 585 validated evaluation instances. On the very_long tier, the strongest model reaches 59.3% avg@5, while six models score below 10%. Training an internal A3B model on 6,000 EvolveScaler examples improves performance over its base checkpoint on all eight independently constructed out-of-distribution benchmarks, with a 5.25-point average gain. These results show that code-driven IE synthesis provides both challenging evaluation and transferable training supervision.