arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用视觉-语言模型进行增强现实中的感知质量评估与自主内容调整

Harnessing Vision-Language Models for Perceptual Quality Assessment and Autonomous Content Adjustment in Augmented Reality

Elias Rotondo, Lin Duan, Yanming Xiu, Sangjun Eom, Conrad Li, Maria Gorlatova

arXiv 2610.00677首次发表:更新:

发表机构

Duke University(杜克大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出基于视觉-语言模型的自动化框架,用于增强现实内容感知质量评估与调整,通过RateAR基准验证,预测与人类判断高度相关,用户研究显示超90%参与者认可系统改进。

AI 中文摘要

增强现实(AR)的进步持续催生创新解决方案,促进了教育系统、医疗服务和风险缓解协议中的新方法。然而,优化终端用户的沉浸感和舒适度仍然具有挑战性,因为AR头戴式显示器面临受限的场景几何、空间抖动和时间不稳定性。用户研究是评估AR视觉质量的标准方法,但其成本、可扩展性降低和灵活性不足在迭代应用设计过程中构成了瓶颈。为解决这一问题,我们提出了一种基于视觉-语言模型(VLMs)的AR内容评估与优化的自动化框架,用于评估和预测用户感知的AR场景视觉保真度。首先,我们引入了RateAR,一个涵盖多样场景和环境条件下收集的AR图像和视频的基准,其在包括物体放置、尺度和阴影一致性在内的感知因素上具有良好至优秀的可靠性(ICC(2,5) >=.90)。随后,我们在该基准上评估了十一个商业VLM。结果表明,基于VLM的质量预测与人类主观判断强相关,达到了最高0.8695的Spearman秩相关系数。一项消融研究进一步表明,与其他提示策略相比,我们的上下文提示在平衡引入的复杂性线索的同时,与人类评分的一致性更好。基于这些发现,我们构建了一个自动化的AR内容调整系统,并进行了21名参与者的用户研究。超过90%的参与者认为该系统改善了虚拟内容的放置和大小一致性。

英文摘要

Advancements in augmented reality (AR) continue to foster innovative solutions, facilitating novel methodologies within educational systems, healthcare delivery, and risk-mitigation protocols. However, optimizing for end-user immersion and comfort remains challenging, as AR head-mounted displays contend with constrained scene geometry, spatial jitter, and temporal instability. User studies are the standard AR evaluation method for visual quality, but their cost, diminishing scalability, and inflexibility pose bottlenecks during iterative application design. To address this problem, we present an automated framework for AR content evaluation and refinement, built on vision-language models (VLMs), to evaluate and predict the visual fidelity of AR scenes as perceived by users. First, we introduce RateAR, a benchmark of AR images and videos collected across diverse scenes and environmental conditions, with good-to-excellent reliability (ICC(2,5) >= .90) across perceptual factors, including object placement, scale, and shadow consistency. Subsequently, we evaluate eleven commercial VLMs on the crafted benchmark. Results support that VLM-based quality predictions strongly correlate with human subjective judgments, achieving Spearman's rank-order correlations of up to 0.8695. An ablation study further suggests that, compared to other prompting strategies, our contextual prompting yields better alignment with human ratings while balancing introduced complexity cues. Building on these findings, we construct an automated AR content adjustment system and conduct a 21-participant user study. More than 90% of participants found that the system improved placement and size coherence of virtual content.

CommentsTo be published in VRST 2026. Main Manuscript: 12 pages, 5 figures; Supplemental Materials: 7 pages, 10 figures. The accompanying public repository can be accessed by visiting https://github.com/Duke-I3T-Lab/RateAR

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑