arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Tsinghua University(清华大学)

2025-12-08 至 2025-12-08 共收录 5
2512.05965 2025-12-08 cs.CV

EditThinker: Unlocking Iterative Reasoning for Any Image Editor

EditThinker: 解锁任何图像编辑器的迭代推理

Hongyu Li, Manyuan Zhang, Dian Zheng, Ziyu Guo, Yimeng Jia, Kaituo Feng, Hao Yu, Yexin Liu, Yan Feng, Peng Pei, Xunliang Cai, Linjiang Huang, Hongsheng Li, Si Liu

机构 * Beihang University(北航大学) Meituan(美团) CUHK MMLab CUHK IMIXR Tsinghua University(清华大学)

AI总结 EditThinker通过迭代推理框架提升图像编辑指令遵循能力,利用强化学习优化编辑过程,显著提高模型性能。

Comments Project page: https://appletea233.github.io/think-while-edit

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05515 2025-12-08 cs.CV cs.LG

DashFusion: Dual-stream Alignment with Hierarchical Bottleneck Fusion for Multimodal Sentiment Analysis

DashFusion: 基于分层瓶颈融合的双流对齐多模态情感分析

Yuhua Wen, Qifei Li, Yingying Zhou, Yingming Gao, Zhengqi Wen, Jianhua Tao, Ya Li

机构 * School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院) Beijing National Research Center for Information Science and Technology, Tsinghua University(清华大学信息科学与技术国家研究中心) Department of Automation, Tsinghua University(清华大学自动化系)

AI总结 DashFusion通过双流对齐与分层瓶颈融合技术,提升多模态情感分析的性能与效率。

Comments Accepted to IEEE Transactions on Neural Networks and Learning Systems (TNNLS), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16334 2025-12-08 cs.AI cs.CL

OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe

OpenMMReasoner: 推动多模态推理的前沿研究:一个开放且通用的配方

Kaichen Zhang, Keming Wu, Zuhao Yang, Bo Li, Kairui Hu, Bin Wang, Ziwei Liu, Xingxuan Li, Lidong Bing

机构 * MiroMind AI Nanyang Technological University(南洋理工大学) Tsinghua University(清华大学) LMMs-Lab Team(多模态实验室团队)

AI总结 OpenMMReasoner提出了一种开放且通用的多模态推理训练配方,通过两阶段方法提升推理性能,实现在多个基准测试中超越现有基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15436 2025-12-08 cs.CV

Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs

基于动态视觉搜索和缩放的自适应聚焦推理方法用于高效VLMs

Xintong Zhang, Zhi Gao, Bofei Zhang, Pengxiang Li, Xiaowen Zhang, Yang Liu, Tao Yuan, Yuwei Wu, Yunde Jia, Song-Chun Zhu, Qing Li

机构 * organization= School of Computer Science \& Technology, Beijing Institute of Technology , city= Beijing , country= China organization= State Key Laboratory of General Artificial Intelligence, BIGAI , city= Beijing , country= China organization= School of Intelligence Science Technology, Peking University , city= Beijing , country= China organization= Guangdong Laboratory of Machine Perception Intelligent Computing, Shenzhen MSU--BIT University , city= Shenzhen , country= China organization= Department of Automation, Tsinghua University , city= Beijing , country= China

AI总结 本文提出基于动态视觉搜索和缩放的自适应聚焦推理方法,提升VLMs的多模态推理效率和实际应用效果。

Comments https://github.com/xtong-zhang/Chain-of-Focus

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01203 2025-12-08 cs.LG

Hypergraph Foundation Model

超图基础模型

Yue Gao, Yifan Feng, Shiquan Liu, Xiangmin Han, Shaoyi Du, Zongze Wu, Han Hu

机构 * School of Software, BNRist, THUIBCS, BLBCI, Tsinghua University(软件学院,BNRist,THUIBCS,BLBCI,清华大学) Institute of Artificial Intelligence and Robotics, College of Artificial Intelligence, Xi’an Jiaotong University(人工智能与机器人研究院,人工智能学院,西安交通大学) College of Mechatronics and Control Engineering, Shenzhen University(机械与控制工程学院,深圳大学) Beijing Institute of Technology(北京理工大学)

AI总结 Hyper-FM是一种多领域知识提取的超图基础模型,通过层次化高阶邻居引导的顶点知识嵌入和结构知识提取,提升超图建模能力,并提出超图基础模型的扩展定律。

详情

展开后加载摘要…

URL PDF HTML 收藏