arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04940cs.LGcs.AIcs.SE

软件世界模型:从后果预测到决策价值

Software World Models: From Consequence Prediction to Decision Value

Tongli Su, Yuntong Hu, Liang Zhao, Bowen Zhu, JayaSai Somasundaram, Hasibul Haque

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出软件世界模型(SWM),通过探索、学习、行动三阶段预测代码变更的爆炸集,相比静态可达性显著提升F1并改善决策质量。

中文摘要 AI 辅助

编码智能体在修改一个代码仓库时,可能会静默地破坏依赖它的下游服务、库或数据存储。在智能体的每次操作后穷尽地运行集成测试是不切实际的,因此智能体必须在执行这些操作之前预测这些失败。现有的软件世界模型预测智能体自身的观察结果,而静态变更影响分析仅识别变更可能传播的位置。我们转而引入软件世界模型(SWM),该模型对受代码变更影响的更广泛系统进行建模,并预测其爆炸集:即变更将破坏的组件。SWM遵循三个阶段:探索、学习和行动。探索阶段从恢复的系统状态执行候选变更,优先考虑观察到的失败与依赖图相矛盾的区域。学习阶段在这些执行结果上微调语言模型,以预测下游破坏。行动阶段将采样的预测转换为每个消费者的破坏概率,用于变更排序、主动迁移以及决定何时再次执行是值得其成本的。在保留的合成系统上,SWM将爆炸集F1从静态可达性的0.431提高到$0.571\pm0.037$,将排序遗憾减少一半以上,并在所有九个评估检查点改善了迁移回报。这两种方法是互补的:可达性在图中表示的依赖上更强,而SWM恢复了图遗漏的耦合引起的失败。在保留的真实库上的实验进一步表明,预测结构化的失败结果(而不仅仅是标量风险)对于下游决策质量很重要。

英文摘要

A coding agent may safely modify one repository while silently breaking downstream services, libraries, or datastores that depend on it. Exhaustively running integration tests after every agent action is impractical, so the agent must predict these failures before executing them. Existing software world models predict the agent's own observations, while static change-impact analysis only identifies where a change may propagate. We instead introduce the Software World Model (SWM), which models the broader system affected by a code change and predicts its blast set: the components that the change will break. SWM follows three stages: explore, learn, and act. Explore executes candidate changes from restored system states, prioritizing regions where observed failures contradict the dependency graph. Learn fine-tunes a language model on these execution outcomes to predict downstream breakage. Act converts sampled predictions into per-consumer break probabilities for change ranking, proactive migration, and deciding when another execution is worth its cost. On held-out synthetic systems, SWM improves blast-set F1 from 0.431 for static reachability to $0.571\pm0.037$, more than halves ranking regret, and improves migration return at all nine evaluation checkpoints. The two methods are complementary: reachability is stronger on dependencies represented in the graph, while SWM recovers failures caused by couplings the graph misses. Experiments on held-out real libraries further show that predicting structured failure outcomes, rather than only scalar risk, is important for downstream decision quality.

发表机构

  • Emory University(埃默里大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑