arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

科学实验场景建模:数据集与模型

Modeling Scientific Experiment Scenes: Dataset and Model

Minghao Zou, Qingtian Zeng, Shangkun Liu, Cong Liu, Paul L. Rosin, Guanghui Yue, Jun Liu, Wei Zhou

arXiv 2608.02892首次发表:更新:

发表机构

College of Computer Science and Engineering, Shandong University of Science and Technology; School of Computer Science and Informatics, Cardiff University; ABC FINTECH Company Limited; NOVA Information Management School, Universidade Nova de Lisboa; School of Biomedical Engineering, Shenzhen University(山东科技大学计算机科学与工程学院; 卡迪夫大学计算机科学与信息学院; ABC金融科技有限公司; 里斯本新大学NOVA信息管理学院; 深圳大学生物医学工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有SGG基准忽略科学实验场景的问题,构建首个物理实验场景SGG数据集PhysScene,并提出跨模态双路径生成器CM-DPG,在PhysScene和VG150上取得具竞争力的SGG性能。

AI 中文摘要

场景图生成(SGG)是结构化视觉理解的基础,但现有基准主要聚焦于日常生活图像,忽略了科学实验场景——这类场景包含专用仪器、特定任务的实验语义以及密集、细粒度的物理关系,且对自动化实验分析和智能教育愈发重要。为填补这一空白,我们推出首个面向物理实验场景的SGG数据集PhysScene,提供密集标注的场景图及多种监督与协议设置下的基准。PhysScene还揭示了SGG的两大关键算法挑战:长尾关系谓词分布显著,以及存在明显的视觉-文本语义鸿沟。为应对这些挑战,我们提出跨模态双路径生成器(CM-DPG),这是一种用于鲁棒开放词汇SGG的模型,该模型通过联合视觉-文本编码增强对象级语义表示,并利用互补的视觉与几何线索改进关系推理;我们还引入关系感知预训练、字幕衍生的伪监督以及自适应加权,以支持头尾谓词的平衡学习。在PhysScene和VG150上的大量实验表明,CM-DPG在多种评估设置下取得了具有竞争力的性能,消融研究验证了各组件的贡献。该数据集和代码可在指定URL公开获取。

英文摘要

Scene Graph Generation (SGG) is fundamental to structured visual understanding, yet existing benchmarks focus mainly on daily-life images and overlook scientific experiment scenes with specialized instruments, task-specific experimental semantics, and dense, fine-grained physical relations. Building upon PhysScene, our previously introduced SGG dataset for physics experiment scenes, we further identify two key challenges that such scientific environments pose to existing SGG models: a pronounced long-tail relational predicate distribution and a substantial visual-textual semantic gap. To address these challenges, we propose the Cross-Modal Dual-Path Generator (CM-DPG), a model for robust open-vocabulary SGG. The model enhances object-level semantic representations through joint visual-textual encoding and improves relational reasoning using complementary visual and geometric cues. We also incorporate relation-aware pre-training, caption-derived pseudo-supervision, and adaptive weighting to support balanced learning across head and tail predicates. Extensive experiments on PhysScene and VG150 show that CM-DPG achieves competitive performance across multiple evaluation settings, with ablation studies validating the contribution of each component. The dataset and code are publicly available at https://github.com/ZMH-SDUST/CM-DPG.

CommentsThe authors have identified issues that require substantial revision and have therefore decided to withdraw the current version

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑