arXivDaily arXiv每日学术速递 周一至周五更新

大厂专区

Tencent(腾讯)

2026-03-18 至 2026-03-18 共收录 3
2603.16206 2026-03-18 cs.LG cs.CL

Offline Exploration-Aware Fine-Tuning for Long-Chain Mathematical Reasoning

离线探索感知微调用于长链数学推理

Yongyu Mu, Jiali Zeng, Fandong Meng, JingBo Zhu, Tong Xiao

机构 * NLP Lab, School of Computer Science and Engineering, Northeastern University, Shenyang, China(东北大学计算机科学与工程学院自然语言处理实验室,中国沈阳) Pattern Recognition Center, WeChat AI, Tencent Inc, China(腾讯公司微信人工智能部门模式识别中心,中国)

AI总结 本文提出OXA微调方法,通过优化两个目标提升数学推理能力,实验显示在六个基准测试中OXA相比传统SFT在Pass@1和Pass@$k$上平均提升6和5分。

Comments Working in process

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15228 2026-03-18 cs.CV

HYDRA: Unifying Multi-modal Generation and Understanding via Representation-Harmonized Tokenization

HYDRA:通过表征协调的标记化统一多模态生成与理解

Xuerui Qiu, Yutao Cui, Guozhen Zhang, Junzhe Li, JiaKui Hu, Xiao Zhang, Yang Li, Songtao Liu, Miles Yang, Yu Shi, Zhao Zhong, Liefeng Bo

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Tencent Hunyuan(腾讯文英) Zhongguancun Academy(中关村学院) Nanjing University(南京大学) Peking University(北京大学)

AI总结 HYDRA通过表征协调的标记化统一多模态生成与理解,提出了一种新的统一框架,在单一参数空间内整合感知与生成,实验表明其在视觉重建和生成任务中均取得新突破。

Comments Work in progress: We are actively scaling up the models. More updates coming soon

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20107 2026-03-18 cs.LG cs.CV

Refining Few-Step Text-to-Multiview Diffusion via Reinforcement Learning

通过强化学习细化少步文本到多视图扩散模型

Ziyi Zhang, Li Shen, Deheng Ye, Yong Luo, Huangxuan Zhao, Meng Liu, Wei Yu, Lefei Zhang

机构 * School of Computer Science, National Engineering Research Center for Multimedia Software and Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University(计算机学院、多媒体软件国家工程研究中心和湖北多媒体与网络通信工程重点实验室、武汉大学) School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University(中山大学信息科学与技术学院深圳校区) Tencent Inc.(腾讯公司) Xiaomi Inc., China(小米公司,中国)

AI总结 本文提出MVC-ZigAL框架,通过联合视图奖励模型和自适应优化策略,提升少步多视图扩散模型的生成质量和一致性。

Comments Accepted to CVPR 2026

Journal ref IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏