arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25876cs.HCcs.CV

视觉-语言模型是否对形状的情感特质达成一致?面向生成式设计界面的跨模型审计

Do Vision-Language Models Agree on the Affective Qualities of Shape? A Cross-Model Audit for Generative Design Interfaces

  • Honda Research Institute Europe GmbH(本田欧洲研究院)
  • Julius-Maximilians-Universität Würzburg(维尔茨堡大学)

机构由 AI 辅助整理,请以论文原文为准。

Luca Bux, Thiago Rios, Ingo Scholtes, Stefan Menzel, Bernhard Sendhoff

AI总结:

本研究审计6个VLM对ShapeNet中10类3D物体的情感特质一致性,发现其一致性介于零基准与几何上限之间,据此实现UI原型指导生成式设计界面的语义控制选择。

AI中文摘要:

生成式设计界面越来越多地开放语义控制,允许用户通过“更优雅”或“更简约”等概念引导输出,这类控制通常由视觉-语言模型(VLM)编码。一个实际问题是,最先进的VLM是否在相同概念上对物体的表示保持一致。我们通过将无纹理的3D物体沿感性形容词对排序来审计6个VLM,其中感性(Kansei)指产品形态的情感印象,每个轴由其两个极点的文本表示之差定义。几何对作为正对照,无关形容词对建立经验零基准。在ShapeNet数据库的10个形状类别中,情感轴的一致性高于零基准(平均成对秩相关系数为0.36,而零基准为0.14),但低于几何上限(0.44)。模型间的一致性是部分的且高度不均:在所有类别共有的3个轴上,平均一致性从书架类的0.21到罐子类的0.51不等。一致性主要取决于类别的表示变化是否与所评估的语义方向对齐,而非单纯取决于物体整体形状的变化程度。跨模型一致性并不意味着与人类判断一致。基于我们的发现,我们实现了一个UI原型,展示该审计如何为给定物体类别提供应开放哪些感性描述符作为控制、哪些应保留的依据。

英文摘要:

Generative design interfaces increasingly expose semantic controls that let users steer output with concepts such as "more elegant" or "more minimalist," typically encoded by a vision-language model (VLM). A practical question is whether state-of-the-art VLMs represent objects consistently in terms of the same concept. We audit 6 VLMs by ranking untextured 3D objects along Kansei adjective pairs, where Kansei describes affective impressions of product form, with each axis defined as the difference between the text representations of its two poles. Geometric pairs serve as positive controls, and pairs of unrelated adjectives establish an empirical null. Across 10 categories of ShapeNet database, affective axes converge above the null (mean pairwise rank correlation 0.36 vs. 0.14) but below the geometric ceiling (0.44). The agreement between models is partial and highly uneven: on the three axes shared by all categories, mean convergence ranges from 0.21 for bookshelves to 0.51 for jars. Convergence depends primarily on whether a category's representational variation aligns with the semantic direction being evaluated, rather than simply on how much the objects vary in shape overall. Cross-model convergence does not imply agreement with human judgments. Based on our findings, we implement a UI prototype that shows how the audit can inform which Kansei descriptors to expose as controls for a given object class and which to withhold.

补充信息

↑