arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FACET:在终端任务合成中保留源意图与可执行状态

FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis

Kou Shi, Zun Wang, Qisheng Su, Shiting Huang, Ziao Zhang, Zhen Fang, Qingnan Ren, Jin Liu, Yu Zeng, Yiming Zhao, Lin Chen, Zehui Chen, Feng Zhao

arXiv 2608.18580首次发表:更新:

发表机构

Shanghai AI Laboratory; Fudan University(上海人工智能实验室; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FACET是保留源意图与可执行状态的终端任务合成框架,可生成高质量任务并提升Terminal-Bench 2.1上的智能体微调性能,为任务合成提供关键原则。

AI 中文摘要

训练终端智能体需要可扩展的可执行监督,但合成高质量终端任务仍具挑战性。每个任务包含指令、初始化环境、参考解决方案和可执行验证器;若这些构件由不一致的假设生成,所得任务可能无法解决或评估错误。同时,多阶段合成可能会丢弃原始源中编码的目标、依赖关系、状态转换和过程约束。我们提出FACET(Fine-grained Agentic Construction of Executable Tasks,细粒度可执行任务智能构建),这一框架可同时解决信息保留和跨构件一致性问题。FACET将相关智能体技能重构为连贯、信息丰富的场景,随后实现并修复执行环境,再生成最终任务构件。所得容器状态作为指令、解决方案和验证器的共享基础,而基于执行的验证和针对性修复可纠正构件特定的故障,无需不必要地重新生成有效组件。FACET生成具有密集可执行检查的复杂终端任务,从这些任务收集的成功轨迹提供了有效、数据高效的监督。在多个尺度上对模型进行微调始终能提升在Terminal-Bench 2.1上的性能,而对替代生成方案的分析支持基于环境的构建对任务有效性和解决方案-验证器对齐的重要性。这些结果确立了源意图保留和共享可执行状态基础是可扩展终端任务合成的关键原则。

英文摘要

Training terminal agents requires scalable executable supervision, yet synthesizing high-quality terminal tasks remains challenging. Each task couples an instruction, an initialized environment, a reference solution, and an executable verifier; if these artifacts are generated from inconsistent assumptions, the resulting task may be unsolvable or incorrectly evaluated. Meanwhile, multi-stage synthesis can discard the goals, dependencies, state transitions, and procedural constraints encoded in the original sources. We present FACET (Fine-grained Agentic Construction of Executable Tasks), a framework that addresses both information preservation and cross-artifact consistency. FACET reconstructs related agent skills into coherent, information-rich scenarios, then realizes and repairs the execution environment before generating the final task artifacts. The resulting container state serves as shared grounding for the instruction, solution, and verifier, while execution-based validation and targeted repair correct artifact-specific failures without unnecessarily regenerating valid components. FACET produces complex terminal tasks with dense executable checks, and successful trajectories collected from these tasks provide effective, data-efficient supervision. Fine-tuning models across multiple scales consistently improves performance on Terminal-Bench 2.1, while analyses of alternative generation schemes support the importance of environment-grounded construction for task validity and solution-verifier alignment. These results establish source-intent preservation and shared executable-state grounding as key principles for scalable terminal-task synthesis.

Commentshttps://stokou.github.io/FACET-Terminal/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑