发表机构
Tribhuvan University; Pulchowk Campus, IOE, Tribhuvan University; NAAMII; INSAIT, Sofia University “St. Kliment Ohridski”(特里布文大学; 特里布文大学工程学院普尔乔克校区; 尼泊尔先进数学与信息学研究所; 索菲亚大学圣克莱门特奥赫里德斯基分校INSAIT)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对现有3D场景生成器无法满足任务约束的问题,提出iARCS迭代智能体强化学习框架,通过两阶段策略优化约束保真度,生成的数据可改进基础生成器,具实用价值。
AI 中文摘要
合成3D场景生成正越来越多地被用作计算机视觉和具身AI的数据源,但现有生成器通常仅优化感知逼真度,无法可靠满足任务关键的功能约束。这种不匹配限制了合成数据在下游训练中的实用性,而可达性、可通行性和空间规则合规性往往是下游训练的关键需求。本文提出iARCS,一种迭代智能体强化学习框架,可使预训练场景生成器适配自然语言任务需求。iARCS采用两阶段策略:通用奖励预训练以提升物理合理性和布局质量,随后使用LLM生成的奖励程序进行任务特定微调,该奖励程序会根据训练反馈迭代优化。实验表明,iARCS在可通行性、可达性和净空聚焦任务上的约束保真度有所提升,能有效优化任务特定约束,且具备有竞争力的场景多样性;进一步研究显示,iARCS生成的数据可改进基础生成器,证明其作为实用合成数据生成工具的价值,而非仅作为可控场景编辑方法。
英文摘要
Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often optimize perceptual realism without reliably satisfying task-critical functional constraints. This mismatch limits the usefulness of synthetic data for downstream training, where accessibility, traversability, and spatial rule compliance are often essential. We present iARCS, an iterative agentic reinforcement learning framework that adapts a pretrained scene generator to naturallanguage task requirements. iARCS uses a two-phase strategy: universal-reward pretraining to improve physical plausibility and layout quality, followed by task-specific finetuning with LLM-generated reward programs that are iteratively refined from training feedback. Experiments show improved constraint fidelity on walkability, reachability, and clearance-focused tasks, effective task-specific constraint optimization, and competitive scene diversity. We further show that data generated by iARCS improves a base generator, supporting its value as a practical synthetic data generation tool rather than only a controllable scene editing method.
Comments20 pages, 13 figures, 9 tables. Includes appendix