arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01383cs.AI

PRISM:一种用于测量和优化多模态类比的范畴论框架

PRISM: A Category-Theoretic Framework for Measuring and Refining Multimodal Analogies

Mirella Zeisler, Ojas Shirekar, Mircea Licǎ, Chirag Raman

首次发表
浏览论文内容

中文总结 AI 辅助

PRISM提出一种基于范畴论的框架,通过拉回分数测量并迭代优化多模态类比的关系对齐,在视觉隐喻生成中显著提升一致性,但优化可能偏向视觉拥挤而非深层关系。

中文摘要 AI 辅助

类比推理涉及跨领域识别和保持关系结构。然而,现有的AI驱动的多模态类比生成方法缺乏一种可解释的度量,来判断生成输出中是否理解并维持了这种结构。我们通过可解释结构映射的拉回优化(PRISM)来解决这一空白,这是一种与模态无关的框架,用于测量和改进多模态类比中的关系对齐,并在视觉隐喻生成上进行了评估。PRISM将类比表示为基于范畴论的显式关系映射,并使用视觉语言模型(VLM)跨模态实例化这些结构。其第一个组件,拉回分数,从生成的图表示中量化关系对齐。在AnaloBench基准上,仅通过拉回分数选择正确类比达到了82.5%的准确率,表明该分数捕获了有意义的关系信息。PRISM的第二个组件是一个迭代优化循环,使用拉回分数作为上下文反馈信号,迭代地修订生成的图像以增强关系深度。VLM作为评判者和人工评估表明,PRISM在隐喻一致性和类比恰当性上持续优于零样本生成,其中人类参与者在57.65%的成对比较中更倾向于优化后的输出。然而,定性分析显示,优化可能偏向于视觉上拥挤的构图,而非真正更深层的关系对应。

英文摘要

Analogical reasoning involves identifying and preserving relational structures across domains. However, existing approaches to AI-driven multimodal analogy generation lack an interpretable measure of whether this structure is understood and maintained in the generated output. We address this gap with Pullback Refinement via Interpretable Structural Mapping (PRISM), a modality-agnostic framework for measuring and improving relational alignment in multimodal analogies, evaluated on visual metaphor generation. PRISM represents analogies as explicit relational mappings grounded in category theory and uses VLMs to instantiate these structures across modalities. Its first component, the pullback score, quantifies relational alignment from the resulting graph representation. On the AnaloBench benchmark, selecting the correct analogy purely by pullback score achieves 82.5% accuracy, demonstrating that the score captures meaningful relational information. PRISM's second component is an iterative refinement loop that uses the pullback score as an in-context feedback signal to iteratively revise the generated image towards greater relational depth. VLM-as-a-judge and human evaluations show that PRISM consistently improves metaphor consistency and analogy appropriateness over zero- shot generation, with human participants preferring the refined output in 57.65% of pairwise comparisons. However, a qualitative analysis reveals that refinement can favour visually crowded compositions rather than genuinely deeper relational correspondences.

发表机构

  • Delft University of Technology(代尔夫特理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑