arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当视觉质量误导时:渲染虚拟形象失真下的意图识别

When Visual Quality Misleads: Intent Recognition under Rendered Avatar Distortions

Ning-Hsuan Chang, Kai-Siang Ma, Yu-Chih Chen

arXiv 2609.27560首次发表:更新:

发表机构

National Chengchi University; National Yang Ming Chiao Tung University(国立政治大学; 国立阳明交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过受控行为实验发现,在虚拟形象流媒体中,视觉质量高并不保证动作识别准确,揭示了质量与准确率的分离,并提出意图质量分数(IQS)作为更合适的评估目标。

AI 中文摘要

虚拟形象流媒体系统通常使用图像和视频质量评估(IQA/VQA)指标进行评估,隐含地将视觉保真度视为通信成功的代理。我们通过一项受控行为研究来检验这一假设,该研究涉及在原始条件和十四种几何、光度、时间及组合失真下的渲染3D虚拟形象。五十九名参与者贡献了2,688个关于感知动作、响应置信度和视觉质量的判断。我们在该数据集中识别出“误导性质量”,即失真渲染保留了高于平均水平的感知质量,但产生了低于平均水平的动作识别准确率。我们还推导出一个意图质量分数(IQS),结合识别正确性和置信度作为客观指标的行为目标。在126个失真内容-条件组合中,31个(24.6%)表现出误导性质量;时间和几何失真的比率最高,分别为50.0%和31.1%。结果揭示了质量-准确率分离现象,其中失真族对外观和通信的影响不同。在24个直接评分的IQA/VQA指标和三个监督特征回归基线中,与IQS的对齐仍然有限;在λ=0.5时,最佳留一内容外基线达到PLCC=0.4435。在这种受控协议下,仅凭视觉保真度不足以实现虚拟形象通信,这促使需要意图感知的质量评估和流媒体目标。

英文摘要

Avatar-streaming systems are commonly evaluated with image and video quality assessment (IQA/VQA) metrics, implicitly treating visual fidelity as a proxy for communicative success. We test this assumption through a controlled behavioral study of rendered 3D avatars across a pristine condition and fourteen geometric, photometric, temporal, and combined distortions. Fifty-nine participants contributed 2,688 judgments of perceived action, response confidence, and visual quality. We identify Misleading Quality in this dataset as distorted renderings that retain above-average perceived quality but yield below-average action-recognition accuracy. We also derive an Intent Quality Score (IQS) combining recognition correctness and confidence as the behavioral target for objective metrics. Among 126 distorted content--condition cells, 31 (24.6%) exhibited Misleading Quality; temporal and geometric distortions showed the highest rates, at 50.0% and 31.1%, respectively. The results reveal a quality--accuracy dissociation where distortion families affect appearance and communication differently. Across 24 direct-scoring IQA/VQA metrics and three supervised feature-regression baselines, alignment with IQS remained limited; at $λ=0.5$, the best leave-one-content-out baseline reached PLCC $=0.4435$. Under this controlled protocol, visual fidelity alone is insufficient for avatar communication, motivating intent-aware quality assessment and streaming objectives.

CommentsAccepted to SIGGRAPH Asia 2026 Technical Communications. 6 pages

DOI:10.1145/3829339.3847861

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑