arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08300cs.CV

以主体为中心的空间理解实现以人为本的图像描述

Human-Centric Image Captioning with Subject-Centered Spatial Understanding

  • Peking University(北京大学)
  • Chinese Academy of Sciences(中国科学院)
  • Kling Team(Kling团队)
  • Nanjing University(南京大学)
  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

Bozhou Li, Jiahang Zhang, Yue Ding, Yushuo Guan, Bohan Zeng, Yiyan Ji, Xinlong Chen, Yang Shi, Yifan Dai, Yuran Wang, Chengzhuo Tong, Pengfei Wan, Yuanxing Zhang, Wentao Zhang

AI总结:

针对多模态大模型在以人为本图像描述中的结构性空间幻觉,提出SPACE基准和专用数据构建与GRPO对齐流程,显著提升以主体为中心的空间推理质量。

AI中文摘要:

尽管多模态大语言模型(MLLMs)在通用图像描述任务上取得了显著性能,但在以人为本的场景中,它们经常出现结构性幻觉。准确建模人类主体对于关键的下游应用(如精确的虚拟形象/视频/图像生成和细粒度的人类动作理解)至关重要。然而,这些任务需要高度精确的以主体为中心的空间定位,例如区分以自我为中心的左/右偏侧性以及保持正确的解剖学-物体绑定。尽管这些局部空间反转对结构完整性是灾难性的,但在现有基准测试中,它们往往被整体描述性指标所掩盖。为了系统地揭示和量化这一瓶颈,我们引入了SPACE(主体中心姿态、外观和特征评估),这是一个旨在评估以主体为中心的空间理解的基准。在SPACE上,我们揭示出,尽管当前MLLMs具有强大的通用感知能力,但它们始终无法将描述锚定在主体的内在参考系中。为了弥合这一差距,我们提出了一种专门的数据构建和对齐流程。我们首先从细粒度的身体部位定位中提取结构化空间提示,以指导两阶段描述重写过程,从而生成高度空间保真的训练数据。此外,我们设计了一种基于评分标准的奖励,用于组相对策略优化(GRPO),在对齐过程中明确惩罚结构关键的空间错误。在SPACE上进行的大量实验表明,我们的框架显著提高了以人为本的描述质量,特别是在以主体为中心的空间推理方面,其性能可与强大的闭源模型相媲美。我们的基准和代码可在以下网址获取:此https URL。

英文摘要:

While multimodal large language models (MLLMs) achieve remarkable performance on generic image captioning, they frequently suffer from structural hallucinations in human-centric scenarios. Accurately modeling human subjects is foundational for critical downstream applications, such as accurate avatar/video/image generation and fine-grained human action understanding. However, these tasks require highly precise subject-centered spatial grounding, such as distinguishing egocentric left/right laterality and maintaining correct anatomical-object bindings. Although catastrophic for structural integrity, these localized spatial inversions are often overshadowed by overall descriptive metrics in existing benchmarks. To systematically expose and quantify this bottleneck, we introduce SPACE (Subject-centric Poses, Appearance, and Characteristics Evaluation), a benchmark designed to evaluate subject-centered spatial understanding. On SPACE, we reveal that despite strong generic perception, current MLLMs consistently fail to ground descriptions in the subject's intrinsic frame of reference. To bridge this gap, we propose a specialized data construction and alignment pipeline. We first extract structured spatial hints from fine-grained body-part localization to guide a two-stage caption rewriting process, yielding highly spatially-faithful training data. Furthermore, we design a rubric-based reward for Group Relative Policy Optimization (GRPO) that explicitly penalizes structurally critical spatial errors during alignment. Extensive experiments on SPACE demonstrate our framework significantly improves human-centric caption quality, particularly in subject-centered spatial reasoning, achieving performance competitive with strong closed-source models. Our benchmark and code are available at https://github.com/JHang2020/SPACE-Eval.

↑