发表机构
Vera Praxis Lab(维拉实践实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对网络智能体训练受环境成本等瓶颈制约的问题,提出含五个系统的数据框架,经验证生成16万条轨迹微调,模型性能大幅提升并登顶排行榜。
AI 中文摘要
训练有能力的网络智能体通常主要被视为模型规模的问题,然而,开放权重模型的后训练更直接地受限于可执行环境的成本、可靠的多轮监督以及访问强大教师模型的能力。我们提出一个以数据为中心的框架,通过五个互补系统解决这些瓶颈:Choulea分析隐藏的推理特征,SkyReal降低教师采样的成本,Hongzwang绕过教师执行时的API限制,PSBreakup恢复因模型合并而减弱的能力,Kreator将专家干预转化为可训练的推理。我们的数据引擎构建了可重置的编码、漏洞、CTF、内核历史、完整利用、固件以及设备支持的环境。候选轨迹仅在执行验证和证据审计后保留,从而产生164,269条用于长上下文监督微调的轨迹。三个检查点在其起始模型的基础上,在完整的CyberGym套件上平均提升了23.76%,在合并的CTF套件上平均提升了10.49%。截至2026年9月1日,Feyospace-s1实现了63.24%的验证成功率,并在官方CyberGym排行榜上排名第10,而所有三个检查点在可比参数规模的模型中均排名第1。据我们所知,这是首次端到端地证明一个七人独立团队能够训练出具有领先智能体网络能力的开放权重模型。
英文摘要
Training capable cyber agents is often treated primarily as a problem of model scale, yet open-weight post-training is constrained more directly by the cost of executable environments, reliable multi-turn supervision, and access to strong teachers. We present a data-centric framework that addresses these bottlenecks through five complementary systems: Choulea analyzes hidden reasoning signatures, SkyReal reduces teacher-sampling cost, Hongzwang bypasses API restrictions on teacher execution, PSBreakup restores capabilities weakened by model merging, and Kreator converts expert interventions into trainable reasoning. Our data engine constructs resettable coding, vulnerability, CTF, kernel-history, full-exploit, firmware, and device-backed environments. Candidate trajectories are retained only after execution verification and evidence auditing, yielding 164,269 trajectories for long-context supervised fine-tuning. The three checkpoints improve over their starting models by an average of 23.76% on the full CyberGym suite and 10.49% across the pooled CTF suites. As of September 1, 2026, Feyospace-s1 achieves a verified success rate of 63.24% and ranks 10th on the official CyberGym leaderboard, while all three checkpoints rank 1st among models at comparable parameter scales. To our knowledge, this is the first end-to-end demonstration that a seven-person independent team can train open-weight models with leading agentic cyber capability.