arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Science and Technology of China(中国科学技术大学)

2025-12-03 至 2025-12-03 共收录 7
2512.02834 2025-12-03 cs.RO cs.AI

Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling Approach

引导视觉-语言-动作模型作为反探索:一种测试时间缩放方法

Siyuan Yang, Yang Zhang, Haoran He, Ling Pan, Xiu Li, Chenjia Bai, Xuelong Li

机构 * Institute of Artificial Intelligence, China Telecom(中国电信人工智能研究院) University of Science and Technology of China(中国科学技术大学) Tsinghua University(清华大学) The Hong Kong University of Science and Technology(香港科学与技术大学)

AI总结 本文提出TACO框架,通过测试时间缩放方法在视觉-语言-动作模型中引入反探索机制,提升推理稳定性和任务成功率。

Comments The first two authors contributed equally. Yang Zhang leads the whole project

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02812 2025-12-03 cs.AI

Enhancing Automated Paper Reproduction via Prompt-Free Collaborative Agents

通过无提示协作代理增强自动论文复现

Zijie Lin, Qilin Cai, Liang Shen, Mingjun Xiao

机构 * University of Science and Technology of China(中国科学技术大学) Meituan(美团)

AI总结 本文提出一种无提示协作代理框架,通过验证和细化代理提升论文到代码生成的准确性和完整性,实验显示性能提升约15%和13%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02624 2025-12-03 cs.CV

PPTBench: Towards Holistic Evaluation of Large Language Models for PowerPoint Layout and Design Understanding

PPTBench: 向大语言模型在PowerPoint布局和设计理解的全面评估迈进

Zheng Huang, Xukai Liu, Tianyu Hu, Kai Zhang, Ye Liu

机构 * University of Science and Technology of China(中国科学技术大学) State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室)

AI总结 PPTBench通过全面评估大语言模型在PowerPoint布局和设计理解上的能力,揭示了当前模型在视觉布局推理和生成中的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22052 2025-12-03 cs.CV

Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos

Ov3R:从RGB视频流中进行开放词汇语义3D重建

Ziren Gong, Xiaohan Li, Fabio Tosi, Jiawei Han, Stefano Mattoccia, Jianfei Cai, Matteo Poggi

机构 * University of Bologna(博洛尼亚大学) USTC(中科大) BIT(北京理工) Monash University(墨尔本大学)

AI总结 Ov3R通过结合CLIP语义和融合描述符,实现从RGB视频流中进行开放词汇语义3D重建,提升空间人工智能的实时性和语义感知能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21025 2025-12-03 cs.SD cs.AI cs.LG eess.AS

Text-Queried Audio Source Separation via Hierarchical Modeling

通过分层建模实现文本查询的音频源分离

Xinlei Yin, Xiulian Peng, Xue Jiang, Zhiwei Xiong, Yan Lu

机构 * University of Science and Technology of China(中国科学技术大学) School of Information and Communication Engineering, Communication University of China(中国通信大学信息与通信工程学院) Microsoft Research Asia(微软亚洲研究院)

AI总结 本文提出HSM-TSS框架,通过分层建模实现文本查询的音频源分离,结合双阶段语义分离与结构保持重建,提升复杂场景下的分离性能与语义一致性。

Comments Accepted by TASLP

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02536 2025-12-03 cs.CV

WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens

WeMMU: 通过噪声查询标记增强视觉-语言模型与扩散模型的桥梁

Jian Yang, Dacheng Yin, Xiaoxuan He, Yong Li, Fengyun Rao, Jing Lyu, Wei Zhai, Yang Cao, Zheng-Jun Zha

机构 * MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(脑启发智能感知与认知联合实验室,中国科学技术大学) ZheJiang University(浙江大学) The Hong Kong University of Science and Technology(香港科技大学)

AI总结 WeMMU通过噪声查询标记和VAE分支,提升视觉-语言模型与扩散模型的连接效率,缓解泛化崩溃问题,实现稳定持续学习。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02523 2025-12-03 cs.SD cs.LG

Generative Multi-modal Feedback for Singing Voice Synthesis Evaluation

生成式多模态反馈用于歌唱语音合成评估

Xueyan Li, Yuxin Wang, Mengjie Jiang, Qingzi Zhu, Jiang Zhang, Zoey Kim, Yazhe Niu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) University of Science and Technology of China(中国科学技术大学) Columbia University(哥伦比亚大学) Dalian University of Technology(大连理工大学) The Chinese University of Hong Kong(香港中文大学)

AI总结 本文提出生成式多模态反馈框架,通过音频-语言模型生成多维语言和音频反馈,提升歌唱语音合成评估的准确性和可解释性。

Comments 16 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏