arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Tsinghua University(清华大学)

2026-08-27 至 2026-08-27 共收录 11
2608.26101 2026-08-27 cs.CV 新提交

RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing

RefVideo-6M:一个用于教学视频编辑的可靠基于参考的数据集

Bojia Zi, Xiaoyan Yang, Yu Zhou, Ruijie Sun, Lihan Zhang, Bin Liang, Kam-Fai Wong, Haibin Huang, Chi Zhang, Xuelong Li

机构 * Institute of Artificial Intelligence, China Telecom (TeleAI)(中国电信人工智能研究院(TeleAI)) The Chinese University of Hong Kong (CUHK)(香港中文大学) Sun Yat-sen University(中山大学) Fudan University(复旦大学) Tsinghua University(清华大学)

AI总结 本文针对现有视频编辑数据集的局限,构建含600万样本的RefVideo-6M参考引导编辑数据集,训练Ref-MoT模型,实验证明其监督更可靠,可提升编辑模型的视觉质量、可控性与参考一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25990 2026-08-27 cs.LG 新提交

Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon

频谱分配:为何Muon优于Adam,以及如何改进Muon

Xiaodong Wu, Wenyi Yu, Chao Zhang, Philip Woodland

机构 * University of Cambridge(剑桥大学) Tsinghua University(清华大学)

AI总结 本文通过频谱分析揭示Muon优于Adam的机制,提出SAMuon及其简化版SAMuon-lite,在多规模modded-nanogpt模型上,二者均优于AdamW和Muon,且SAMuon可减少13.3%-24.0%训练令牌。

Comments 34 pages, 13 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25872 2026-08-27 cs.RO 新提交

VISTA: Visually Inferred Spatial ConTact Attention for Contact-Rich Manipulation

VISTA:用于接触丰富操作的视觉推断空间接触注意力机制

Jiayi Chen, Wenlong Dong, Yan Huang, Xianglin Chen, Zijian Lin, Jiaqi Yin, Yushan Liu, Wenbo Ding

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Southern University of Science and Technology(南方科技大学) School of Future Technology, Harbin Institute of Technology(哈尔滨工业大学未来技术学院)

AI总结 针对接触丰富操作的视觉反馈不足问题,提出基于VDF的VISTA-Policy,通过整合物理感知编码引擎等模块,在多任务上优于纯视觉与触觉基准,具分布外泛化及抗干扰能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25823 2026-08-27 cs.LG cs.MM 新提交

Learning Continuous Regional Temperature Fields with Lead-Time and Resolution Queries

基于提前期与分辨率查询的连续区域温度场学习

Chunlei Shi, Jiong Wang, Yi-Lin Wei, Junming Hou, Jinjin Liu, Yecheng Zhang, Dan Niu

机构 * Southeast University(东南大学) Fudan University(复旦大学) Sun Yat-sen University(中山大学) Tsinghua University(清华大学)

AI总结 该研究针对现有区域温度预报器输出固定的问题,提出CSTF模型,通过引入提前期、分辨率等查询实现灵活的温度预报,在基准测试中降低了17.0%的偏差,表现出更优性能。

Comments 16 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25659 2026-08-27 cs.RO 新提交

GaussianDream++: Efficient 3D Gaussian World Modeling for Robotic Manipulation

Yuqing Jiang, Zijian Zhang, Weitao Zhou, Jiawei Wang, Junjie He, Lei Yang, Haifang Qing, Si Liu, Ding Zhao, Ping Luo, Haibao Yu

机构 * Tuojing Intelligence(拓境智能) University of Chinese Academy of Sciences(中国科学院大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Tsinghua University(清华大学) University of Science and Technology of China(中国科学技术大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Nanyang Technological University(南洋理工大学) Beihang University(北京航空航天大学) Carnegie Mellon University(卡内基梅隆大学) The University of Hong Kong(香港大学)

Comments 17 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25653 2026-08-27 cs.CV 新提交

Towards Purified Multi-Label Test-Time Adaptation of Vision-Language Models

面向视觉-语言模型的纯化多标签测试时自适应方法

Yiwen Liang, Hui Chen, Yizhe Xiong, Mengyao Lyu, Yuhan Cao, Zijia Lin, Shuaicheng Niu, Sicheng Zhao, Jungong Han, Guiguang Ding

机构 * Tsinghua University(清华大学) Nanyang Technological University(南洋理工大学) University of Washington(华盛顿大学)

AI总结 针对视觉-语言模型多标签测试时自适应的缓存方法PuRF,通过区域与缓存纯化解决现有方法的偏差与校准问题,在五项数据集上使ViT-B/32的mAP提升4.05%,性能优于现有最优方法。

Comments Accepted by ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25618 2026-08-27 cs.CL 新提交

AWM: Answerable Working Memory for Long-Document VQA Agents

AWM:面向长文档VLM智能体的可回答工作记忆

Dongzhuoran Zhou, Yuqicheng Zhu, Yule Liu, Zhen Yang, Rui Lu, Yuxiao Dong, Jie Tang, Evgeny Kharlamov

机构 * University of Oslo(奥斯陆大学) Bosch Center for AI(博世人工智能中心) University of Stuttgart(斯图加特大学) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Tsinghua University(清华大学)

AI总结 针对长文档VQA智能体的记忆质量盲区,提出AWM方法,通过融入仅记忆可回答性的GRPO奖励机制,提升了最终答案准确率并降低记忆缺失正确的比例。

Comments EMNLP 2026 Findings. 16 pages, 4 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25572 2026-08-27 cs.RO cs.AI 新提交

ConfAL-WM: Confidence-Guided Active Learning for Action-Conditioned World Models

ConfAL-WM:用于动作条件世界模型的置信度引导主动学习

Xiang Liu, Sen Cui, Changshui Zhang

机构 * Tsinghua University(清华大学)

AI总结 针对动作条件世界模型在新场景下的局部错误问题,提出ConfAL-WM框架,基于EVAC和UNet实现置信度引导的高效数据选择与局部训练增强,在RoboTwin2.0上验证了其提升训练效率与预测质量的效果。

Comments Project page: this https URL (https://ConfAL-WM.github.io)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25417 2026-08-27 cs.AI 新提交

Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents

绘制你所见:多模态智能体的灵巧视觉工具使用基准测试

Shudong Liu, Dongyang Chen, Enci Zhang, Jinwei Liang, Zheng Ma, Lewei Lu

机构 * Peking University(北京大学) Tsinghua University(清华大学) SenseTime Research(商汤科技研究院)

AI总结 本文提出名为EASEL的基准测试,结合44万样本的EASEL-Data数据集与EASEL-9B模型,评估多模态智能体的灵巧视觉工具使用能力,发现现有25个模型在该任务上表现不佳,EASEL-9B相对基础模型提升6.3%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24945 2026-08-27 cs.LG cs.AI cs.DC 新提交

FAMPWQ: Fisher Information-based Adaptive Mixed Precision Weight Quantization for Effective LLM Inference

FAMPWQ:基于费舍尔信息的自适应混合精度权重量化方法,用于高效的大语言模型推理

Gongwei Lee, Ji Liu, Juncheng Jia, Ji Wu

机构 * School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院) Hithink Research(海思睿研究中心) Electronic Engineering, Tsinghua University(清华大学电子工程系)

AI总结 针对LLM部署的资源瓶颈,该研究提出FAMPWQ方法,通过费舍尔信息度量层敏感度结合强化学习分配位宽,在7个模型和5个基准上优于7种基线方法,提升了量化性能。

Comments 21 pages, to appear in EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24886 2026-08-27 cs.AI 新提交

VLM-based automatic multi-granularity graph representation of building layouts for design informatics

基于视觉语言模型的建筑布局多粒度图自动表示方法,用于设计信息学

Song Guo, Zhuoshi Chen, Maosu Li, Weimin Zhuang

机构 * Massachusetts Institute of Technology(麻省理工学院) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Tsinghua University(清华大学)

AI总结 本研究提出基于VLM的多粒度图自动构建方法,以147份高校图书馆平面图为案例,验证其生成的图与人工标注图一致性良好,可用于设计信息学相关任务,提升建筑全生命周期设计信息利用率。

详情

展开后加载摘要…

URL PDF HTML 收藏