人行道时刻:更丰富的表示总是更符合人类认知吗?来自城市漫步视频的证据
Sidewalk Moments: Are Richer Representations Always More Human-Aligned? Evidence from City-Walk Videos
浏览论文内容
中文总结 AI 辅助
研究借助多种模态表示的城市漫步视频,探讨更丰富视觉表示与人类认知的关系。通过相关性分析等方法,发现不同模态在不同场景表现有别,挑战了丰富表示更符合人类认知的假设,提出时间压缩可替代全视频编码。
中文摘要 AI 辅助
我们使用来自YouTube的61个第一人称城市漫步视频,将其分割成超过50,000个十秒片段,并通过四种模态表示:时空视频特征、时间平均图像(TAI)、音频嵌入和基于文本的语义描述,来研究更丰富的视觉表示是否能产生更符合人类认知的城市参与度测量。Spearman相关性分析揭示了沿时间丰富度连续体的预期排序,视频特征显示出最强的连续一致性。然而,在高参与度与低参与度时刻的二元分类下(这是训练感知评分模型最常用的范式),这种排序被打破,在大多数分类器和分位数阈值下,TAI始终与视频匹配或优于视频。一项独立的亚马逊土耳其机器人二选一强制选择研究证实,这种平等反映了人类的判断:参与者从TAI和完整视频片段中识别出参与时刻的准确率相当,而文本表现则差得多,音频仍接近随机水平。差距分析揭示了一种功能分离:视频特征在具有动态内容的活动驱动场景中具有优势,而TAI在由稳定空间结构主导地构图驱动场景中与人类判断更一致。这些发现挑战了更丰富的表示本质上更符合人类认知的假设,并表明基于感知的时间压缩可以成为全视频编码的原则性替代方案。
英文摘要
We examine whether richer visual representations yield more human-aligned measures of urban engagement, using 61 first-person city-walk videos from YouTube segmented into over 50,000 ten-second clips and represented across four modalities: spatiotemporal video features, temporally averaged images (TAIs), audio embeddings, and text-based semantic descriptions. Spearman correlation analysis reveals the expected ordering along the temporal-richness continuum, with video features showing the strongest continuous alignment. However, this ordering breaks down under binary classification of high- versus low-engagement moments (the paradigm most commonly used to train perceptual scoring models), where TAIs consistently match or outperform video across most classifiers and quantile thresholds. An independent two-alternative forced-choice study on Amazon Mechanical Turk confirms that this parity reflects human judgment: participants identified engaging moments with comparable accuracy from TAIs and full video clips, while text performed substantially worse and audio remained near chance. Gap analysis reveals a functional dissociation: video features are advantaged in activity-driven scenes with dynamic content, whereas TAIs better align with human judgments in composition-driven scenes dominated by stable spatial structure. These findings challenge the assumption that richer representations are inherently more human-aligned, and suggest that perceptually grounded temporal compression can be a principled alternative to full video encoding.
发表机构
- Massachusetts Institute of Technology(麻省理工学院)
- City Form Lab, MIT(麻省理工学院城市形态实验室)
- Senseable City Lab, MIT(麻省理工学院可感知城市实验室)
机构由 AI 辅助整理,请以论文原文为准。