arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SerenAI:受基于文本的世界AI模型启发的状态转换系统

SerenAI: State-transition system inspired by text-based world AI models

Elvin Babayev, Artem Sinitsa, Arash Hajisharifi, Kabir Bakhshaei

arXiv 2609.06647首次发表:更新:

发表机构

Collision Technologies S.R.L.S. Societa Benefit(碰撞科技有限责任公司(福利型))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SerenAI提出一种受世界模型启发的状态转换系统,通过参数高效微调和基于验证器的强化学习,将结构化预测的JSON有效性从85%提升至93.2%,并定义了法律级验证协议。

AI 中文摘要

尽管专业工作流程广泛利用大语言模型,但当这种生成需要法律、运营或财务工作流程时,对无约束自由文本生成的审计解释通常是难以处理的。我们在此展示一个名为SerenAI的基于文本的系统——受世界模型启发,它是一个状态转换系统,输出可验证的预测而非仅仅是文本:给定环境、状态和动作的描述,生成的输出包含4个项目:对给定状态产生因果影响的因果增量、能够从给定状态和动作逻辑上跟随的下一状态、一个有效性奖励,以及一个终止信号。对于发布的原型模型,我们采用了2步适应训练,即参数高效微调,随后在跨越10个推理领域的12个环境中的50,000个示例因果关系上进行基于验证器的强化学习。与8B开放权重基线的初始内部评估相比,SerenAI将JSON有效性从85.0%提高到93.2%,模式有效性从55.0%提高到84.0%,精确结构化输出匹配从0.0%提高到41.5%,因果增量精确匹配从0.0%提高到41.5%,结果状态精确匹配从0.0%提高到42.0%,奖励精确匹配从1.0%提高到80.5%,终止精确匹配从38.0%提高到81.5%。这些支持了更窄的主张,即验证器兼容的适应可以改善结构化转换预测。它们尚未确立法律级可靠性。因此,论文还规定了一个验证协议,用于基于证据的法律工作流程、校准、人工监督和主权本地部署。

英文摘要

Although professional workflows leverage large language models widely, the interpretation for auditing unconstrained free-text generation is usually intractable if such generation demands legal, operational or financial workflow. We hereby demonstrate a text based system called SerenAI - inspired by world-models, it is a state transition system that outputs verifiable predictions rather than merely text: Provided with a description of the environment, state, and actions, the generated output contains 4 items: causal deltas that causally effect the given state, a next state that can logically follow from the given state and action, a validity reward, and a termination signal. For the released proto-model, we employ 2 steps of adaptation training, namely parameter efficient fine-tuning followed by verifier based RL over 50,000 exampled cause and effects in 12 environments spanning 10 reasoning domains. Compared to an initial internal evaluation of an 8B open-weight baseline, SerenAI increased JSON validity from 85.0% to 93.2%, schema validity from 55.0% to 84.0%, exact structured-output match from 0.0% to 41.5%, causal-delta exact match from 0.0% to 41.5%, resulting-state exact match from 0.0% to 42.0%, reward exact match from 1.0% to 80.5%, and termination exact match from 38.0% to 81.5%. These support the narrower claim that verifier-compatible adaptation can improve structured transition prediction. They do not yet establish legal-grade reliability. Accordingly, the paper also specifies a validation protocol for evidence-grounded legal workflows, calibration, human oversight, and sovereign on-premise deployment.

Comments8 pages, 4 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑