arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Science and Technology of China(中国科学技术大学)

共收录 2226
2608.17707 2026-08-19 cs.CV cs.MM 新提交

DynaForcing: Overcoming Dynamic Collapse in Self-Forcing Distillation for Streaming Avatar Generation

DynaForcing:克服流式化身生成中自强制蒸馏的动态崩溃问题

Yubo Huang, Sirui Zhao, Xinchen Yao, Zhengye Zhang, Jinyang Huang, Fengqi Cui, Shiwei Wu, Enhong Chen

机构 * University of Science and Technology of China(中国科学技术大学) Nanjing University(南京大学) Hefei University of Technology(合肥工业大学) Tsinghua University(清华大学)

AI总结 针对流式化身生成中自强制蒸馏的动态崩溃问题,提出DynaForcing框架,通过三种策略及优化技术恢复动态、提升视觉质量,解决了质量与动态的权衡问题。

Comments Accepted at ACM International Conference on Multimedia (MM '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17426 2026-08-19 cs.CV cs.AI 新提交

SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation

SemComp-Bench:视频生成中的语义任务完成基准

Keyu Tu, Zhuowei Chen, Mengqi Huang, Yuxin Wang, Jiahao Zhu, Zhendong Mao, Yongdong Zhang

机构 * University of Science and Technology of China(中国科学技术大学) Sun Yat-sen University(中山大学)

AI总结 该研究提出面向结果的视频生成任务语义任务完成,构建SemComp-Data数据集与SemComp-Bench基准,发现代表性视频生成模型在兼顾结果实现与参考图像语义基础上仍具挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17271 2026-08-19 cs.AI 新提交

ASI-Bench: At the Dawn of Artificial Superintelligence

ASI-Bench:人工智能超级智能的黎明

Junwei Zhou, Zhen Sun, Binyu Li, Jiangyu Zhou, Yuexi Pan, Hengyu Wang, Honghe Ren, Xiaohan Jia, Xueyang Zhou, Xiaoyu Cao, Yongchao Chen, Yuanning Feng, Junhao Wu, Cheng Zhang, Sijia Chen, Haoyu Xue, Chengsong You, Huan Wang, Koutian Wu, Peigan Gao, Jiakun Wu, Wenzhe Li, Ergan Shang, Qingyuan Zheng, Jingjing Zhou, Ruixuan Jia, Yan Xu, Hongrui Zhang, Xiao-Han Ma, Zhengxiang Cheng, Yuexing Hao, Liting Mai, Xianglin Ji, Wenjun Zhang, Zhuofan Chen, Yixiao Huang, Chi Wang, Wenyue Hua, Yilun Hao, Yuantao Zhai, Ziyan Zhao, Jingyan Xie

机构 * Tsinghua University(清华大学) Massachusetts Institute of Technology(麻省理工学院) Harvard University(哈佛大学) Carnegie Mellon University(卡内基梅隆大学) University of Michigan(密歇根大学) University of Illinois Urbana–Champaign(伊利诺伊大学厄巴纳-香槟分校) Boston University(波士顿大学) University of Queensland(昆士兰大学) University of Science and Technology of China(中国科学技术大学) Flatiron Institute(弗拉蒂伦研究所) Microsoft Research(微软研究院) AG2 AI(AG2 AI公司)

AI总结 研究人员推出首个联合评估AI创新探索与自主科学执行能力的基准ASI-Bench,逐步减少人类方法指导测试AI自主科研能力,发现当前AI仍严重依赖人类指导,该基准面向全球开放邀请贡献。

Comments 16 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13546 2026-08-19 cs.CV 版本更新

Alaya-EVOKE: From Linear-Scaling Supervision to Endless World

Alaya-EVOKE:从线性缩放监督到无尽世界

Yuanyang Yin, Gongxuan Wang, Yifan Zhan, Chuanhao Li, Kaipeng Zhang, Feng Zhao

机构 * MoE Key Lab of BIPC(BIPC教育部重点实验室) USTC(中国科学技术大学) Shanghai Innovation Institute(上海创新研究院) Alaya Lab(Alaya实验室)

AI总结 Alaya-EVOKE将持久世界状态外部化并重新设计教师模型,解决交互式世界模型的冲突需求,在WBench等基准上实现最优性能,支持开放式长时序生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07725 2026-08-19 cs.CL cs.AI 版本更新

SOD: Step-wise On-policy Distillation for Small Language Model Agents

SOD:分步式在线蒸馏用于小型语言模型代理

Qiyong Zhong, Mao Zheng, Mingyang Song, Xin Lin, Jie Sun, Houcheng Jiang, Xiang Wang, Junfeng Fang

机构 * Zhejiang University(浙江大学) Large Language Model Department, Tencent(腾讯大语言模型部门) University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学)

AI总结 针对小型语言模型中工具集成推理的稳定性问题,提出SOD分步式在线蒸馏框架,通过动态调整蒸馏强度缓解教师信号误导,提升推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14229 2026-08-19 cs.AI cs.CV cs.RO 版本更新

HA-VLN 2.0: An Open Benchmark and Leaderboard for Human-Aware Navigation in Discrete and Continuous Environments with Dynamic Multi-Human Interactions

HA-VLN 2.0:面向离散与连续环境中动态多人交互的人类感知导航开放基准与排行榜

Yifei Dong, Fengyi Wu, Qi He, Lingdong Kong, Heng Li, Minghan Li, Zebang Cheng, Yuxuan Zhou, Jingdong Sun, Qi Dai, Alexander G Hauptmann, Zhi-Qi Cheng

机构 * University of Science and Technology of China(中国科学技术大学)

AI总结 提出HA-VLN 2.0统一基准,通过标准化任务、HAPS 2.0数据集与模拟器、16844条社会指令基准测试及真实机器人实验,证明显式社会建模提升导航鲁棒性并减少碰撞。

Comments Accepted to IROS 2026. 35 pages, 20 figures, website: this https URL (https://uwmilab.github.io/HA-VLN-webpage/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06185 2026-08-19 cs.CY cs.AI cs.CL cs.HC 版本更新

Hidden Prompts in Manuscripts Exploit AI-Assisted Peer Review

手稿中的隐藏提示利用AI辅助同行评审

Zhicheng Lin

机构 * Department of Psychology, Yonsei University(延世大学心理学系) Department of Psychology, University of Science and Technology of China(中国科学技术大学心理学系)

AI总结 研究探讨了手稿中隐藏指令对AI同行评审的影响,分析了提示注入技术及学术出版物政策的不一致,指出需加强技术筛查与政策协调。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16647 2026-08-18 cs.CL 新提交

Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models

硬币皆有两面:关于大型语言模型的策略内蒸馏中泛化的双重性质

Zhaoyi Li, Deyang Kong, Yuan Wei, Evan Yang, Ranran Shen, Mahardika Krisna Ihsani, Ming Yang, Wei Zhang, Chuan Hao, Jian Yang, Ran Tao, Bryan Dai, Shikun Zhang, Wei Ye, Ying Wei, Defu Lian

机构 * University of Science and Technology of China(中国科学技术大学) Peking University(北京大学) IQuest Research(IQuest研究院) MBZUAI(穆罕默德·本·扎耶德人工智能大学) Zhejiang University(浙江大学)

AI总结 本研究探究大型语言模型策略内蒸馏(OPD)的泛化特性,发现其迁移教师推理行为而非答案,同源配对泛化性强,异源配对适配性有限,多教师组合存在能力跷跷板效应,为诊断多教师OPD提供了视角。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16477 2026-08-18 cs.LG 新提交

Pallas: A Proactive KV Cache Migration Framework for LLM Inference in AI-RAN

Pallas:面向AI-RAN中LLM推理的主动KV缓存迁移框架

Tianhang Ding, Jianchun Liu, Hongli Xu

机构 * University of Science and Technology of China(中国科学技术大学)

AI总结 Pallas是面向AI-RAN中LLM推理的主动KV缓存迁移框架,通过切换前在目标基站预准备推理状态,结合在线调度器优化,显著降低服务中断时间与令牌间延迟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16447 2026-08-18 cs.AI cs.RO 新提交

HaReCAP: Habitual-action Grounding for Recursive Large Language Model Agents

HaReCAP:面向递归大语言模型智能体的习惯性动作 grounding 方法

Shen Liu, Zhenguo Xu, Shaopu Wang, Yike Gao, Chunlei Wang

机构 * North China Institute of Computer System Engineering(华北计算机系统工程研究所) University of Science and Technology of China(中国科学技术大学) China Information Security Research Institute Co., Ltd.(中国信息安全研究院有限公司)

AI总结 HaReCAP 是针对 ReCAP 的低侵入性扩展,通过编译叶子反射规则减少长视距具身任务中 LLM 的重复调用,在 Robotouille 和 ALFWorld 上显著降低了 token 消耗。

Comments 15 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16417 2026-08-18 cs.CL 新提交

D2-ScaleAgent: Dual-Dimensional Scaling for Long Document Understanding

D2-ScaleAgent:面向长文档理解的双维度缩放方法

Hao Zhang, Longrong Yang, Lunhao Duan, Ziyang Wang, Qing-Guo Chen, Shanshan Zhao

机构 * Zhejiang University(浙江大学) Alibaba Group(阿里巴巴集团) University of Science and Technology of China(中国科学技术大学)

AI总结 针对现有多模态RAG长文档理解方法缺乏动态计算缩放能力的问题,提出D2-ScaleAgent双维度缩放智能体框架,在相关基准上取得良好效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16210 2026-08-18 cs.LG stat.ML 新提交

Conditional Evaluation of Language Models with Cheap Auxiliary Signals

基于廉价辅助信号的语言模型条件评估

Zhi Zhang, Lingfeng Lyu, Yue Kang, Doudou Zhou

机构 * University of California, Los Angeles(加利福尼亚大学洛杉矶分校) University of Science and Technology of China(中国科学技术大学) Microsoft(微软公司) National University of Singapore(新加坡国立大学)

AI总结 针对现有语言模型条件评估中廉价辅助信号存在偏差的问题,提出半监督估计器LACE,结合局部中心化与岭控制变量,在多个基准数据集上完成实证评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16195 2026-08-18 cs.RO 新提交

RoboStriker: Latent-Space Strategic Games for Autonomous Humanoid Boxing

RoboStriker:用于自主人形机器人拳击的隐空间策略博弈

Kangning Yin, Kaige Liu, Zhe Cao, Wentao Dong, Weishuai Zeng, Tianyi Zhang, Qiang Zhang, Jingbo Wang, Jiangmiao Pang, Yang Li, Ming Zhou, Weinan Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Innovation Institute(上海创新研究院) Huazhong University of Science and Technology(华中科技大学) University of Science and Technology of China(中国科学技术大学)

AI总结 本研究提出RoboStriker框架,将人形拳击任务形式化为隐空间零和马尔可夫博弈,通过分层结构解耦推理与执行,实现了优于原始动作空间方法的战术性能并成功部署到现实人形机器人。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16082 2026-08-18 cs.LG 新提交

Towards Reasonable Molecular Structure Elucidation from Infrared Spectroscopy with Chemical Feedback

基于化学反馈的红外光谱合理分子结构解析研究

Yusen Tan, Hongyu Zhan, Hai-tao Yu, Changxi Chi, Wenjie Du, Jun Xia

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Westlake University(西湖大学) University of Science and Technology of China(中国科学技术大学)

AI总结 针对现有分子结构解析模型预测结果不合理的问题,提出化学反馈驱动的FIRMPO框架,在三类红外数据集上显著提升了解析准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15509 2026-08-18 cs.RO cs.FL cs.LG 新提交

Temporal Logic Guided Universal Task Representations for Reinforcement Learning

用于强化学习的时序逻辑引导通用任务表示

Hao Zhang, Zhangli Zhou, Zhen Kan

机构 * University of Science and Technology of China(中国科学技术大学) Anhui Provincial Key Laboratory of Humanoid Robots(安徽省人形机器人重点实验室) Chinese Academy of Sciences(中国科学院)

AI总结 本文提出受时序逻辑启发的通用任务表示框架LOTUS,将其集成至强化学习算法,通过LTL编码器建模语义、双模拟度量保障稳定性,在多场景下的学习效率、泛化能力等指标优于现有方法。

Comments Accepted by IEEE Transactions on Neural Networks and Learning Systems (Early Access). Project page: https://lotus-website.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15224 2026-08-18 cs.LG 新提交

Structuring Semantic Embeddings for Principle Evaluation: A Prototype-Guided Contrastive Learning Approach

用于原则评估的语义嵌入结构化:一种原型引导的对比学习方法

Che Shen, Junwei Su, Lingpeng Kong, Chuan Wu

机构 * The University of Hong Kong(香港大学) University of Science and Technology of China(中国科学技术大学)

AI总结 本文针对通用文本嵌入在任务区分上的不足,提出PGCL方法,通过多流结合的正则化模块生成任务适配表示,在三个代理任务数据集上均优于原始冻结嵌入,明确了方法适用边界并修订了理论分析。

Comments Accepted for publication in Transactions on Machine Learning Research (TMLR). 27 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15211 2026-08-18 cs.CV cs.DC 新提交

TERRA: A Hierarchical Parallel Training and Memory Orchestration Framework for High-Resolution AI-based Earth Modeling

TERRA:面向高分辨率AI地球建模的分层并行训练与内存编排框架

Ruohan Wu, Ziqi Zhu, Yang Zhao, Jiarui Tang, Yingzhe Cui, Junshi Chen, Zhao Jing, Jun Shi, Hong An

机构 * School of Computer Science and Technology, University of Science and Technology of China(中国科学技术大学计算机科学与技术学院) School of Artificial Intelligence and Data Science, University of Science and Technology of China(中国科学技术大学人工智能与数据科学学院) Laoshan Laboratory(崂山实验室) Ocean University of China(中国海洋大学)

AI总结 TERRA是面向高分辨率AI地球预报的分层并行训练与内存编排框架,通过SAWSTP和MO技术,在96个H200 GPU上支持114亿参数模型,实现高算力与内存优化,提升预报精度。

Comments 15 pages, 16 figures, 6 tables, and 2 algorithms. Submitted to IEEE Transactions on Parallel and Distributed Systems (TPDS). Code is available at https://github.com/ruohan12345/TERRA

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14783 2026-08-18 cs.CV cs.GR 新提交

MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling

MegaParts:通过令牌高效自回归建模将部件感知三维物体生成扩展至300个部件

Manwen Liao, Xinyu Lian, Jian Mao, Kaixu Chen, Li Luo, Jinghao Yan, Wanshui Gan, Qiao Yu, Weitian Zhang, Chunhua Shen, Guang Chen, Bo Dai, Xudong Xu, Zhaoyang Lyu

机构 * The University of Hong Kong(香港大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Fudan University(复旦大学) Tongji University(同济大学) University of Science and Technology of China(中国科学技术大学) Shanghai Jiao Tong University(上海交通大学)

AI总结 MegaParts提出结合结构化序列建模与令牌高效矢量量化形状令牌化器的自回归框架,将部件感知3D生成扩展至300个部件,性能优于基线模型,为大规模部件感知3D生成提供新方案。

Comments 12 pages, 6 pages appendix, 13 figures, technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15124 2026-08-18 cs.LG math.OC 新提交

Decision-Driven Regularization: A Blended Model for Learning and Optimization

决策驱动正则化:一种用于学习与优化的混合模型

Gar Goei Loke, Qinshen Tang, Yangge Xiao, Xun Zhang

机构 * Durham University Business School(杜伦大学商学院) Nanyang Business School, Nanyang Technological University(南洋理工大学南洋商学院) Faculty of Business and Economics, The University of Melbourne(墨尔本大学商学院与经济学院) International Institute of Finance, School of Management, University of Science and Technology of China(中国科学技术大学管理学院国际金融研究所)

AI总结 针对上下文优化中分离式学习与优化易致决策有效性下降的问题,提出决策驱动正则化框架,平衡预测精度与代价最小化,可推广SPO+,在合成研究中优于多个基准模型。

Comments 42 pages (including appendix), 7 figures in main, journal paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12888 2026-08-18 cs.CL 版本更新

When Your Agent Opens the Chat App: Agent-Controlled Search over Raw Chat Logs Rivals Structured Memory

当你的智能体打开聊天应用时:智能体控制的原始聊天日志搜索可与结构化记忆相媲美

Ruizhe Li, Licheng Zhang, Benfeng Xu, Mingxuan Du, Zheren Fu, Weidong Chen

机构 * University of Science and Technology of China(中国科学技术大学) MetaStone Technology(MetaStone科技公司)

AI总结 该研究提出无语义结构的智能体控制搜索界面ReFind,在MemoryAgentBench等任务上,其检索准确率优于HippoRAG 2等结构化记忆系统,证明可控词汇检索可替代复杂记忆结构实现高效会话记忆任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03025 2026-08-18 cs.AI 版本更新

DiffImaginE: Imagine to Verify Entity Types with Diffusion

DiffImaginE:通过想象验证实体类型

Feng Zhang, Feiyu Han, Rongxin Yang, Yang Liu, Yancheng Chen, Rui Wang, Yingguang Yang, Tian Xueyun, Chongyang Zhang, Hao Zheng, Xu Kefu, Congjing Ran, Fuhai Chen, Bin Chong

机构 * Fuzhou University(福州大学) Chinese Academy of Sciences(中国科学院) Peking University(北京大学) Alibaba Group(阿里巴巴集团) University of Science and Technology of China(中国科学技术大学) Fullive Innovation (Beijing) AI Technology Co., Ltd.(福莱创新(北京)人工智能科技有限公司) Baidu(百度) Wuhan University(武汉大学)

AI总结 DiffImaginE将多模态命名实体识别的类型验证建模为条件潜在扩散推理,在Twitter-2015和Twitter-2017数据集上,相比确定性对照模型取得了一致的性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20753 2026-08-18 physics.chem-ph cs.AI 版本更新

Empowering Polymeric Materials Discovery by Artificial Intelligence

人工智能赋能高分子材料发现

Chenyao Ma, Linda Zhang, Yuheng Chen, Wei Du, Shangwen Fang, Zihao Jiang, Chuanyu Liu, Xinyu Ma, Rui Su, Gang Wang, Muyao Yu, Dong Zhong, Jie Zhu, Weibo Gong, Huan Gu, Limin Li, Chen Shen, Rui Wu, Zhenghao Wu, Kan Xu, Min Zhou, Donglin He, Xiayun Huang, Shan Jiang, Pengfei Ou, Jiayu Peng, Yuwei Zhang, Jie Zhao, Di Zhang, Piao Ma, Zhenghao Li, Hao Li

机构 * Suzhou MatSource Technology Co., Ltd.(苏州MatSource科技有限公司) Gusu Laboratory of Materials(材料Gusu实验室) Advanced Institute for Materials Research (WPI-AIMR)(先进材料研究所(WPI-AIMR)) Frontier Research Institute for Interdisciplinary Sciences (FRIS)(交叉学科前沿研究所(FRIS)) State Key Laboratory of Advanced Environmental Technology, Department of Environmental Science and Engineering, University of Science and Technology of China(先进技术国家实验室,环境科学与工程系,中国科学技术大学) Jiangsu Key Laboratory of New Power Batteries, Jiangsu Collaborative Innovation Centre of Biomedical Functional Materials, School of Chemistry and Materials Science, Nanjing Normal University(新型动力电池江苏省重点实验室,生物医学功能材料协同创新中心,化学与材料科学学院,南京师范大学) Department of Chemistry, National University of Singapore(新加坡国立大学化学系) Thrust of Sustainable Energy and Environment, The Hong Kong University of Science and Technology (Guangzhou)(可持续能源与环境方向,香港科技大学(广州)) Department of Materials Design and Innovation, University at Buffalo(材料设计与创新系,布法罗大学) College of Smart Materials and Future Energy, State Key Laboratory of Molecular Engineering of Polymers, Fudan University(智能材料与未来能源学院,聚合物分子工程国家重点实验室,复旦大学) School of Physical Science and Technology, Shanghai tech University(物理科学与技术学院,上海科技大学) The State Key Laboratory of Molecular Engineering of Polymers and Department of Macromolecular Science, Fudan University(聚合物分子工程国家重点实验室和大分子科学系,复旦大学) Department of Chemistry and Materials Science, Xi'an J Liverpool University(化学与材料科学系,西安J Liverpool大学)

AI总结 本文综述了数据基础设施、机器学习、大模型和实验室自动化如何融合成自主发现生态系统,通过自改进反馈循环实现高分子材料的预测性、可重复和可扩展创新。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06628 2026-08-18 cs.AI 版本更新

Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability

重新思考推理SFT中的泛化:对优化、数据和模型能力的条件分析

Qihan Ren, Peng Wang, Ruikun Cai, Shuai Shao, Dadi Guo, Yuejin Xie, Yafu Li, Quanshi Zhang, Xia Hu, Jing Shao, Dongrui Liu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) University of Science and Technology of China(中国科学技术大学)

AI总结 本文研究推理SFT的泛化问题,发现泛化受优化动态、训练数据和模型能力共同影响,指出短训练检查点可能低估泛化能力,数据质量和结构及模型能力均影响泛化效果,且推理提升与安全下降存在不对称性。

Comments Accepted by COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27959 2026-08-18 cs.CV 版本更新

MathGen: Revealing the Illusion of Mathematical Competence through Text-to-Image Generation

MathGen:通过文本到图像生成揭示数学能力的幻觉

Ruiyao Liu, Hui Shen, Ping Zhang, Yunta Hsieh, Yifan Zhang, Jing Xu, Qi Han, Junchen Li, Jiawei Lu, Jianing Ma, Jiaqi Mo, Sicheng Chen, Zhen Zhang, Zhongwei Wan, Jing Xiong, Xin Wang, Ziyuan Liu, Hangrui Cao, Ngai Wong

机构 * University of Pennsylvania(宾夕法尼亚大学) University of Michigan(密歇根大学) The Ohio State University(俄亥俄州立大学) USTC(中国科学技术大学) City University of Hong Kong(香港城市大学) University of Wisconsin(威斯康星大学) UCSB(加州大学圣塔芭芭拉分校) University of Hong Kong(香港大学) Peking University(北京大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 MathGen通过900道跨七个核心领域的数学问题,评估生成模型在视觉化数学解答中的准确性,发现现有模型在数学 fidelity 上存在显著瓶颈。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18094 2026-08-18 cs.CV cs.AI cs.DB 版本更新

OODBench: Out-of-Distribution Benchmark for Large Vision-Language Models

OODBench: 用于大型视觉-语言模型的分布外基准

Ling Lin, Yang Bai, Heng Su, Congcong Zhu, Yaoxing Wang, Yang Zhou, Huazhu Fu, Jingrun Chen

机构 * University of Science and Technology of China Suzhou Institute for Advanced Research, USTC Key Laboratory of the Ministry of Education for Mathematical Foundations Unmanned System Research Institute, Northwestern Polytechnical University

AI总结 OODBench提出了一种自动化方法,用于构建评估大型视觉-语言模型处理分布外数据能力的基准,并展示了当前模型在面对此类数据时的性能下降。

Comments 54 pages, 21 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17360 2026-08-18 cs.LG cs.AI cs.CR 版本更新

Robust Privacy: Inference-Stage Privacy through Certified Robustness

鲁棒隐私:通过认证鲁棒性实现推理阶段隐私

Jiankai Jin, Xiangzheng Zhang, Zhao Liu, Wenzhuo Xu, Dongdong Yang, Deyue Zhang, Quanchen Zou

机构 * University of Science and Technology of China(中国科学技术大学)

AI总结 提出鲁棒隐私(RP)概念,基于认证鲁棒性确保预测在输入邻域内不变,从而限制推理阶段隐私泄露;实验表明RP在属性推断和模型反演攻击中有效提升隐私-效用权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24900 2026-08-18 cs.CV cs.AI 版本更新

OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing

OpenGPT-4o-Image:用于高级图像生成与编辑的综合数据集

Zhihong Chen, Xuehai Bai, Yang Shi, Chaoyou Fu, Huanyu Zhang, Haotian Wang, Xiaoyan Sun, Zhang Zhang, Liang Wang, Yuanxing Zhang, Pengfei Wan, Yi-Fan Zhang

机构 * USTC(中国科学技术大学) Kling Team(Kling团队) HDU(华中科技大学) PKU(北京大学) NJU(南京大学) CASIA(中国科学院自动化研究所) THU(清华大学)

AI总结 该研究构建了涵盖11个领域51个子任务的OpenGPT-4o-Image数据集,经其微调主流模型后,图像编辑任务性能最高提升18%、生成任务最高提升13%,证实系统化数据构建对多模态AI的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14452 2026-08-17 cs.AI 新提交

SheetCompass: Hierarchical Relation Graphs for Agentic Spreadsheet Reasoning

SheetCompass:面向智能体电子表格推理的分层关系图

Panjing He, Mingyue Cheng, Yucong Luo, Li Li, Xiaohan Zhang

机构 * University of Science and Technology of China(中国科学技术大学)

AI总结 针对LLM处理电子表格时丢失多维结构语义的问题,提出SheetCompass框架,通过显式建模表内外结构关系并结合内存机制,提升智能体对复杂工作簿的推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14385 2026-08-17 cs.LG cs.AI 新提交

DeaMoE: Efficient MoE Structure for Fast Small-Batch Decoding

DeaMoE:用于快速小批量解码的高效MoE结构

Zewen Jin, Shen Fu, Zeping Duan, Shannon Wang, Weihao Wu, Chengjie Tang, Congkun Ai, Ping Gong, Zijian Dai, Youhui Bai, Cheng Li

机构 * University of Science and Technology of China(中国科学技术大学) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院) Shanxi University(山西大学)

AI总结 针对小批量解码下MoE模型专家权重加载的瓶颈,提出DeaMoE架构,通过专家分组共享参数与两阶段路由策略提升效率,在多款模型及显卡上实现显著速度提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14125 2026-08-17 cs.AI 新提交

Traj-LeWM: Path-Aware World-Model Planning via Latent Trajectory Cost

Traj-LeWM:通过潜在轨迹代价实现路径感知的世界模型规划

Xiaodi Huang, Ziyi Ding, Jingtian Wan, Yuchen Liu, Yuan Zhang, Xiao-Ping Zhang, Jiayu Chen, Zhang Zhang, Tao Huang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Shanghai Jiao Tong University(上海交通大学) Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院) The University of Hong Kong(香港大学) INFIFORCE University of Science and Technology of China(中国科学技术大学) Peking University(北京大学)

AI总结 本文针对LeWM的局限提出Traj-LeWM,通过引入潜在轨迹代价结合终点距离的联合评分,在多个机器人与导航任务上实现性能提升,验证了轨迹级信息的互补作用。

详情

展开后加载摘要…

URL PDF HTML 收藏