DAC-Pose:用于姿态引导人体生成的双智能体协作框架
DAC-Pose: Dual-Agent Collaborative Framework for Pose-Guided Human Generation
浏览论文内容
中文总结 AI 辅助
针对姿态引导人体生成在剧烈视角变化下的视觉伪影问题,提出双智能体协作框架DAC-Pose,通过PSR与DAVE智能体的反馈循环实现高保真合成,在DeepFashion和Market-1501基准上验证了其优越性。
中文摘要 AI 辅助
AI智能体已成为生成式图像合成领域强大的新范式,使系统能执行复杂语义推理,而非被动的像素级映射。在姿态引导人体生成任务中,传统方法在剧烈视角变化下不可避免会产生严重视觉伪影,根本原因在于它们缺乏逻辑推断未可见区域的认知能力,以及对复杂空间形变进行建模的能力。为弥合这一差距,我们提出DAC-Pose,一种新颖的智能体驱动多模态框架,将单视角人体生成重新表述为协作双智能体系统。DAC-Pose集成了两个互补组件,即先验语义推理(PSR)智能体和感知差异视觉编码(DAVE)智能体。PSR作为认知引擎,利用协作推理推断未可见区域的细粒度属性;DAVE作为专用视觉感知智能体,量化并编码视角引发的空间错位,持续将鲁棒空间约束反馈到生成过程中。语义推理与视觉感知间的自主反馈循环确保了高保真细节合成。在DeepFashion和Market-1501基准上的大量实验验证了我们智能体驱动范式的优越性,值得注意的是,DAC-Pose在剧烈视角变化下能出色保留纹理对齐和身份一致性。代码可在该https URL获取。
英文摘要
AI agents have emerged as a powerful new paradigm in generative image synthesis, enabling systems to perform complex semantic reasoning rather than passive pixel-level mapping. In pose-guided human generation, conventional methods inevitably produce severe visual artifacts under drastic viewpoint shifts, fundamentally because they lack the cognitive capacity to logically deduce unseen regions and model complex spatial deformations. To bridge this gap, we propose DAC-Pose, a novel agent-driven multimodal framework that reformulates single-view human generation as a collaborative dual-agent system. DAC-Pose integrates two complementary components, namely, the Prior Semantic Reasoning (PSR) agent and the Discrepancy-Aware Visual Encoding (DAVE) agent. Functioning as a cognitive engine, PSR utilizes collaborative reasoning to deduce the fine-grained attributes of unseen regions. Concurrently, acting as a specialized visual perception agent, DAVE quantifies and encodes viewpoint-induced spatial misalignments, continuously feeding robust spatial constraints back into the generative process. This autonomous feedback loop between semantic deduction and visual perception ensures high-fidelity detail synthesis. Extensive experiments on the DeepFashion and Market-1501 benchmarks validate the superiority of our agent-driven paradigm. Notably, DAC-Pose excels in preserving texture alignment and identity consistency under drastic viewpoint changes. The code is available at https://github.com/AIVRC/DAC-Pose.
发表机构
- Faculty of Data Science, City University of Macau(澳门城市大学数据科学学院)
- Shenzhen University of Advanced Technology(深圳理工大学)
- Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
- School of Computing and Mathematical Sciences, University of Leicester(莱斯特大学计算与数学科学学院)
机构由 AI 辅助整理,请以论文原文为准。