arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Chinese Academy of Sciences(中国科学院大学)

2026-01-08 至 2026-01-08 共收录 8
2601.04035 2026-01-08 cs.AI

MobileDreamer: Generative Sketch World Model for GUI Agent

MobileDreamer: 用于GUI代理的生成式草图世界模型

Yilin Cao, Yufeng Zhong, Zhixiong Zeng, Liming Zheng, Jing Huang, Haibo Qiu, Peng Shi, Wenji Mao, Wan Guanglu

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Meituan(美团)

AI总结 MobileDreamer通过生成式草图世界模型和rollout想象策略,提升GUI代理在长周期任务中的决策能力,任务成功率提升5.25%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04666 2026-01-08 cs.CV

PhysDepth: Plug-and-Play Physical Refinement for Monocular Depth Estimation in Challenging Environments

PhysDepth:用于在恶劣环境中单目深度估计的即插即用物理细化

Kebin Peng, Haotang Li, Zhenyu Qi, Huashan Chen, Zi Wang, Wei Zhang, Sen He, Huanrui Yang, Qing Guo

机构 * Department of Computer Science, East Carolina University(东卡罗来纳大学计算机科学系) Department of Electrical and Computer Engineering, The University of Arizona(亚利桑那大学电气与计算机工程系) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) School of Computer and Cyber Sciences, Augusta University(奥古斯塔大学计算机与网络安全学院) VCIP, CS, Nankai University(南开大学)

AI总结 PhysDepth通过引入物理先验信息,提升了在恶劣环境中单目深度估计的鲁棒性和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03783 2026-01-08 cs.CL

HearSay Benchmark: Do Audio LLMs Leak What They Hear?

HearSay基准:音频大语言模型是否泄露所听内容?

Jin Wang, Liang Lin, Kaiwen Luo, Weiliu Wang, Yitian Chen, Moayad Aloqaily, Xuehai Tang, Zhenhong Zhou, Kun Wang, Li Sun, Qingsong Wen

机构 * XDU(北华大学) NTU(国立台湾大学) NCEPU(南京工程大学) BUPT(北京邮电大学) SHU(上海大学) UAEU(阿联酋大学) UCAS-IIE(中国科学院大学国际学院) Squirrel AI

AI总结 HearSay基准研究发现音频大语言模型通过声纹泄露隐私,揭示其固有隐私风险及安全机制不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03713 2026-01-08 cs.CV

BREATH-VL: Vision-Language-Guided 6-DoF Bronchoscopy Localization via Semantic-Geometric Fusion

BREATH-VL:基于视觉-语言引导的6自由度支气管镜定位:通过语义-几何融合

Qingyao Tian, Bingyu Yang, Huai Liao, Xinyan Huang, Junyong Li, Dong Yi, Hongbin Liu

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Department of Pulmonary and Critical Care Medicine, The First Affiliated Hospital of Sun Yat-sen University(中山大学附属第一医院呼吸与危重症医学科) Centre of AI and Robotics, Hong Kong Institute of Science & Innovation, Chinese Academy of Sciences(香港科学院人工智能与机器人中心) School of Biomedical Engineering and Imaging Sciences, King’s College London(伦敦国王学院生物医学工程与影像科学学院)

AI总结 BREATH-VL通过融合语义和几何信息,实现高精度的6自由度支气管镜定位,减少定位误差并提升效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03577 2026-01-08 cs.LG

Variational Inference, Entropy, and Orthogonality: A Unified Theory of Mixture-of-Experts

变分推断、熵与正交性:混合专家模型的统一理论

Ye Su, Yong Liu

机构 * University of Chinese Academy of Sciences, Beijing, China(中国科学院大学)

AI总结 本文从贝叶斯和信息论视角构建混合专家模型的统一理论框架,推导路由机制为最优稀疏后验近似,并证明正交性可缩小全局最优与贪心近似间的差距。

Comments 27 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03301 2026-01-08 cs.MA cs.AI

PC2P: Multi-Agent Path Finding via Personalized-Enhanced Communication and Crowd Perception

PC2P:通过个性化增强通信与人群感知进行多智能体路径寻找

Guotao Li, Shaoyun Xu, Yuexing Hao, Yang Wang, Yuhui Sun

机构 * Institute of Microelectronics of the Chinese Academy of Sciences(中国科学院微电子研究所) University of Chinese Academy of Sciences(中国科学院大学)

AI总结 PC2P通过个性化增强通信与人群感知方法,提升多智能体路径寻找在复杂环境中的协同与扩展能力。

Comments 8 pages,7 figures,3 tables,Accepted to IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03261 2026-01-08 cs.CL cs.AI

DeepResearch-Slice: Bridging the Retrieval-Utilization Gap via Explicit Text Slicing

DeepResearch-Slice: 通过显式文本切片弥合检索-利用差距

Shuo Lu, Yinuo Xu, Jianjie Cheng, Lingxiao He, Meng Wang, Jian Liang

机构 * NLPR & MAIS CASIA(CASIA NLPR与MAIS) School of AI UCAS(UCAS人工智能学院) Meituan Inc.(美团公司)

AI总结 DeepResearch-Slice通过显式文本切片技术,有效解决检索与利用之间的差距问题,提升模型在嘈杂环境中的鲁棒性。

Comments Ongoing work

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14435 2026-01-08 cs.CV cs.LG

MoTE: Mixture of Ternary Experts for Memory-efficient Large Multimodal Models

MoTE:混合三元专家用于内存高效的大型多模态模型

Hongyu Wang, Jiayu Xu, Ruiping Wang, Yan Feng, Yitao Zhai, Peng Pei, Xunliang Cai, Xilin Chen

机构 * Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院人工智能安全重点实验室,计算技术研究所) University of Chinese Academy of Sciences(中国科学院大学)

AI总结 MoTE通过训练更多低精度三元专家,实现内存高效的大规模多模态模型训练,提升端任务性能并降低内存需求。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏