arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

感知、细化、推理:社交媒体战略视觉传播中指标测量的校准流水线

Perceive, Refine, Reason: A Calibrated Pipeline for Measuring Indicators in Strategic Visual Communication on Social Media

Weihong Qi, Chen Ling

arXiv 2609.14699首次发表:更新:

发表机构

Indiana University Bloomington(印第安纳大学伯明顿分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出PRR校准流水线,结合视觉-语言检测器、SAM空间细化和多模态LLM仲裁,实现社交媒体图像中战略视觉指标的精确测量,并在美国立法者图像分析中验证其有效性。

AI 中文摘要

视觉内容塑造了社交媒体上受众的感知和观点,计算社会科学日益依赖自动化工具来大规模分析图像。然而,测量差距依然存在:现有工具依赖预定义类别或仅产生粗略的图像级标签,而测量图像中出现哪些特定对象、其显著性如何以及在画面中的位置,在大规模情况下仍然困难。我们引入了感知、细化、推理(PRR),一个校准流水线,将灵活的视觉-语言检测器转变为可用于社会科学研究的可审计测量工具。PRR将自然语言类别提示与通过Segment Anything Model(SAM)进行的像素级空间细化以及多模态LLM仲裁层相结合,其推理链外化领域知识并降低了人类参与验证的专业门槛。一个互补的三层可审计性框架应用量化学习来刻画逐类别的可靠性、支持任务对齐配置,并统计校正流行率估计。在四个视觉-语言检测器和九个社会学类别中,该流水线相比零样本基线产生了显著的精度提升,包括最强骨干网络43.3个百分点的改进。将PRR应用于2024年选举周期中美国立法者的103,920张Facebook图像,并将检测结果与DW-NOMINATE意识形态得分相关联,我们发现更保守的立法者将美国国旗展示为更大的视觉元素,且倾向于较弱的边缘放置,这种空间模式在二元检测中不可见。PRR为计算社会科学家提供了一个模型无关的工具包,用于可访问、空间基础且可纠正的视觉测量。

英文摘要

Visual content shapes audience perception and opinion on social media, and computational social science increasingly relies on automated tools to analyze images at scale. Yet a measurement gap persists: existing tools rely on predefined categories or produce only coarse image-level labels, while measuring which specific objects appear in an image, how prominently, and where in the frame remains difficult at scale. We introduce Perceive, Refine, Reason (PRR), a calibrated pipeline that turns flexible vision-language detectors into auditable measurement instruments for social-scientific research. PRR combines natural-language category prompts with pixel-level spatial refinement via the Segment Anything Model (SAM) and a multimodal LLM arbitration layer whose reasoning chains externalize domain knowledge and lower the expertise threshold for human-in-the-loop validation. A complementary three-tier auditability framework applies quantification learning to profile per-category reliability, support task-aligned configuration, and statistically correct prevalence estimates. Across four vision-language detectors and nine sociological categories, the pipeline yields substantial precision gains over zero-shot baselines, including a 43.3-point improvement for the strongest backbone. Applying PRR to 103,920 Facebook images from U.S. legislators during the 2024 election cycle and linking detections to DW-NOMINATE ideology scores, we find that more conservative legislators display U.S. flags as larger visual elements, with a weaker tendency toward peripheral placement, a spatial pattern invisible to binary detection. PRR provides computational social scientists with a model-agnostic toolkit for accessible, spatially-grounded, and correctable visual measurement.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑