arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Chinese Academy of Sciences(中国科学院大学)

2025-12-01 至 2025-12-01 共收录 11
2511.23034 2025-12-01 cs.RO

LatBot: Distilling Universal Latent Actions for Vision-Language-Action Models

LatBot: 从大规模物体操作视频中提炼通用潜在动作用于视觉-语言-动作模型

Zuolei Li, Xingyu Gao, Xiaofan Wang, Jianlong Fu

机构 * Institute of Microelectronics, Chinese Academy of Sciences(中国科学院微电子研究所) University of Chinese Academy of Sciences(中国科学院大学) Microsoft Research(微软研究院)

AI总结 LatBot通过整合动作预测和潜在动作分解,提升视觉-语言-动作模型在现实世界和模拟环境中的泛化与迁移能力。

Comments Project Page: https://mm-robot.github.io/distill_latent_action/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22892 2025-12-01 cs.CV cs.LG

ClearGCD: Mitigating Shortcut Learning For Robust Generalized Category Discovery

ClearGCD: 缓解捷径学习以实现稳健的通用类别发现

Kailin Lyu, Jianwei He, Long Xiao, Jianing Zeng, Liang Fan, Lin Shu, Jie Hao

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Loughborough University(洛桑大学)

AI总结 ClearGCD通过语义对齐和捷径抑制正则化缓解捷径学习,提升通用类别发现的稳健性和泛化能力。

Comments 5 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22434 2025-12-01 cs.CR cs.AI

FastFHE: Packing-Scalable and Depthwise-Separable CNN Inference Over FHE

FastFHE: 基于FHE的可扩展打包和深度可分离CNN推理

Wenbo Song, Xinxin Fan, Quanliang Jing, Shaoye Luo, Wenqi Wei, Chi Lin, Yunfeng Lu, Ling Liu

机构 * Institute of Computing Technology, CAS(中国科学院计算技术研究所) UCAS(中国科学院大学) Fordham University(福特汉姆大学) Dalian University of Technology(大连理工大学) Beihang University(北京航空航天大学) Georgia Institute of Technology(佐治亚理工学院)

AI总结 FastFHE通过可扩展加密打包、深度可分离卷积、BN点积融合和Legendre多项式近似,提升FHE下CNN推理效率与精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22171 2025-12-01 cs.CV cs.GR

BrepGPT: Autoregressive B-rep Generation with Voronoi Half-Patch

BrepGPT: 基于Voronoi半块的自回归B-rep生成

Pu Li, Wenhao Zhang, Weize Quan, Biao Zhang, Peter Wonka, Dong-Ming Yan

机构 * MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS) University of Chinese Academy of Sciences(中国科学院大学) King Abdullah University of Science and Technology(国王 Abdullah 科学与技术大学)

AI总结 BrepGPT通过Voronoi半块表示实现单阶段自回归B-rep生成,提升生成效率与模型紧凑性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22135 2025-12-01 cs.CV

EASL: Multi-Emotion Guided Semantic Disentanglement for Expressive Sign Language Generation

EASL: 多情绪引导的语义解耦用于表达性手语生成

Yanchao Zhao, Jihao Zhu, Yu Liu, Weizhuo Chen, Yuling Yang, Kun Peng

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) University of Chinese Academy of Sciences(中国科学院大学) University of Health and Rehabilitation Sciences(康复科学大学) The University of Aberdeen(阿伯丁大学)

AI总结 EASL通过多情绪引导的语义解耦架构,提升手语生成的表达性和情感表现力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21707 2025-12-01 cs.NI cs.AI

Sensing and Understanding the World over Air: A Large Multimodal Model for Mobile Networks

通过空气感知和理解世界:为移动网络设计的大型多模态模型

Zhuoran Duan, Yuhao Wei, Guoshun Nan, Zijun Wang, Yan Yan, Lihua Xiong, Yuhan Ran, Ji Zhang, Jian Li, Qimei Cui, Xiaofeng Tao, Tony Q. S. Quek

机构 * National Engineering Research Center for Mobile Network Technologies, Beijing University of Posts and Telecommunications (BUPT), Beijing(中国移动网络技术国家工程研究中心,北京邮电大学) Beiyou Shenzhen Institute(北邮深圳研究所) School of Cyber Security, University of Chinese Academy of Sciences (UCAS)(中国科学院大学网络安全学院) China Telecom Co., Ltd.(中国电信股份有限公司) Singapore University of Technology and Design (SUTD)(新加坡科技设计大学)

AI总结 本文提出了一种无线原生多模态模型,利用无线信号进行对比学习,验证了其在无线网络中的应用潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04182 2025-12-01 cs.CL cs.AI

COPO: Causal-Oriented Policy Optimization for Hallucinations of MLLMs

COPO:面向MLLM幻觉的因果导向策略优化

Peizheng Guo, Jingyao Wang, Wenwen Qiang, Jiahuan Zhou, Changwen Zheng, Gang Hua

机构 * University of Chinese Academy of Sciences(中国科学院大学) Institute of Software Chinese Academy of Sciences(中国科学院软件研究所) Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所) Amazon.com, Inc.(亚马逊公司)

AI总结 COPO通过引入因果完整性奖励和GRPO框架,解决多模态大语言模型的幻觉问题,通过令牌级因果约束确保输出的正确性和证据基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01873 2025-12-01 cs.CV

DiffusionFF: A Diffusion-based Framework for Joint Face Forgery Detection and Fine-Grained Artifact Localization

DiffusionFF: 一种基于扩散的联合人脸伪造检测与细粒度特征定位框架

Siran Peng, Haoyuan Zhang, Li Gao, Tianshuo Zhang, Xiangyu Zhu, Bao Li, Weisong Zhao, Zhen Lei

机构 * MAIS, CASIA(CASIA人工智能研究所) SAI, UCAS(UCAS智能信息学院) CMFT(计算机视觉与模式识别技术研究所) IIE, CAS(中国科学院信息工程研究所) SCS, UCAS(UCAS安全学院) CAIR, HKISI, CAS(中国科学院自动化研究所) SCSE, FIE, M.U.S.T(慕苏尔科技大学安全与电子工程系)

AI总结 DiffusionFF通过结合预训练的伪造检测器和去噪扩散模型,实现人脸伪造检测与细粒度特征定位,提升检测能力和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19482 2025-12-01 cs.CL

KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models

KSHSeek: 通过数据驱动方法缓解和检测生成模型中的知识捷径幻觉

Zhongxin Liu, Zhiwei Wang, Jun Niu, Ying Li, Hongyu Sun, Meng Xu, He Wang, Gaofei Wu, Yuqing Zhang

机构 * Xidian University(西安电子科技大学) University of Chinese Academy of Sciences(中国科学院大学) Hainan University(海南大学) University of Waterloo(滑铁卢大学)

AI总结 KSHSeek通过数据驱动方法缓解和检测生成模型中的知识捷径幻觉,提升模型的鲁棒性和可靠性。

Comments 16 pages, 34 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.05538 2025-12-01 cs.CV

A Survey on Personalized Content Synthesis with Diffusion Models

扩散模型在个性化内容合成中的综述

Xulu Zhang, Xiaoyong Wei, Wentao Hu, Jinlin Wu, Jiaxin Wu, Wengyu Zhang, Zhaoxiang Zhang, Zhen Lei, Qing Li

机构 * Department of Computing(计算系) The Hong Kong Polytechnic University(香港理工大学) Center for Artificial Intelligence and Robotics(人工智能与机器人中心) Hong Kong Institute of Science & Innovation(香港科学创新研究院) Chinese Academy of Sciences(中国科学院) State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室) Chinese Academy of Sciences Institute of Automation(中国科学院自动化研究所) School of Artificial Intelligence(人工智能学院) University of Chinese Academy of Sciences(中国科学院大学)

AI总结 本文综述了扩散模型在个性化内容合成中的应用,分析了测试时微调和预训练适应方法,探讨了个性化任务的挑战与创新,并提出了未来发展方向。

Journal ref Machine intelligence research, Oct. 2025, v. 22, no. 5, p. 817-848

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.08174 2025-12-01 cs.CV

OBSeg: Accurate and Fast Instance Segmentation Framework Using Segmentation Foundation Models with Oriented Bounding Box Prompts

OBSeg: 基于定向边界框提示的准确且快速的实例分割框架

Zhen Zhou, Junfeng Fan, Yunkai Ma, Sihan Zhao, Fengshui Jing, Min Tan

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学)

AI总结 OBSeg通过引入定向边界框提示,提升了实例分割的准确性和速度,同时优化了轻量级基础模型的性能。

Journal ref Machine Intelligence Research 2025

详情

展开后加载摘要…

URL PDF HTML 收藏