arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18627cs.CV

PCQA-R1:基于强化学习的通用三维点云质量评估方法

PCQA-R1: Advancing Generalized 3D Point Cloud Quality Assessment with Reinforcement Learning

Kangning Ye, Yunhao Li, Sijing Wu, Yucheng Zhu, Guangtao Zhai

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出首个用于三维点云质量评估的强化学习LMM PCQA-R1,基于GRPO策略构建PCQA-CoT数据集并引入高斯 proximity 奖励,在跨数据集泛化和域内准确率上表现优异。

中文摘要 AI 辅助

无参考点云质量评估(PCQA)是近年来的活跃研究方向,用于衡量和优化点云的视觉体验。然而,大型多模态模型(LMM)在该领域的应用鲜有探索。现有基于LMM的方法主要依赖监督微调直接预测数值质量分数,缺乏在具有不同平均意见得分(MOS)尺度和有限标注的数据集间的泛化能力。核心难点在于,绝对MOS回归在具有不同分数尺度和失真分布的数据集间易受干扰,而相对质量排序在这种分布偏移下更稳定。本文提出PCQA-R1,这是首个用于三维点云质量评估的强化学习LMM,可同时建模质量理解与评分。该方法基于组相对策略优化(GRPO)策略构建,首先通过反向推理策略构建思维链数据集PCQA-CoT,作为冷启动训练数据,指导LMM生成推理过程。进一步引入高斯 proximity 奖励,通过将分数预测锚定到源MOS范围,防止校准漂移。实验结果表明,PCQA-R1在五个基准测试中实现了最先进的跨数据集泛化性能,且在域内准确率上具有竞争力。消融研究验证了排序、高斯奖励和冷启动轨迹的作用。

英文摘要

No-reference point cloud quality assessment (PCQA) has been an active topic in recent years and is used to measure and optimize the visual experience of point clouds. However, large multimodal models (LMMs) have rarely been explored in this area. Previous LMM-based methods mainly rely on supervised fine-tuning to directly predict numerical quality scores, lacking the ability to generalize across datasets with heterogeneous MOS scales and limited annotations. A key difficulty is that absolute MOS regression can be brittle across datasets with different score scales and distortion distributions, whereas relative quality ranking is more stable under such shifts. In this paper, we present PCQA-R1, the first reinforcement learning LMM for 3D point cloud quality assessment to simultaneously model quality understanding and scoring. Built upon the group relative policy optimization (GRPO) strategy, PCQA-R1 first constructs a chain-of-thought dataset, PCQA-CoT, which serves as cold-start training data through a reverse reasoning strategy that teaches the LMM to generate its reasoning process. We further introduce a Gaussian proximity reward that prevents calibration drift by anchoring score predictions to the source MOS range. Experimental results demonstrate that PCQA-R1 achieves state-of-the-art cross-dataset generalization across five benchmarks and competitive in-domain accuracy. Ablation studies support the role of ranking, Gaussian reward, and cold-start traces.

↑