arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-11-25 至 2025-11-25 共收录 7 信号源:cs.CL, cs.AI, cs.LG

1. 视觉空间推理 7 篇

2511.18450 2025-11-25 cs.AI 79%

ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical Constraints

ORIGAMISPACE:基于数学约束的多模态大语言模型多步空间推理基准测试

Rui Xu, Dakuan Lu, Zicheng Zhao, Xiaoyu Tan, Xintao Wang, Siyu Yuan, Jiangjie Chen, Yinghui Xu

机构 * Fudan University(复旦大学) SII INF Technology(INF技术)

专题命中 视觉空间推理 :reasoning(title,abstract);分类 cs.AI

AI总结 ORIGAMISPACE通过折纸任务评估多模态大语言模型在多步空间推理和数学约束处理中的能力,提出四个评估任务并探索强化学习训练方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03596 2025-11-25 cs.CV 78%

ControlThinker: Unveiling Latent Semantics for Controllable Image Generation through Visual Reasoning

ControlThinker: 通过视觉推理揭示潜在语义以实现可控图像生成

Feng Han, Yang Jiao, Shaoxiang Chen, Junhao Xu, Jingjing Chen, Yu-Gang Jiang

机构 * Shanghai Key Lab of Intell. Info. Processing, School of CS, Fudan University(上海智能信息处理关键实验室,复旦大学计算机学院) Shanghai Collaborative Innovation Center on Intelligent Visual Computing(上海智能视觉计算协同创新中心) MiniMax

专题命中 视觉空间推理 :reasoning(title,abstract)

AI总结 ControlThinker通过视觉推理挖掘潜在语义,提升可控图像生成的语义一致性和视觉质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18874 2025-11-25 cs.AI cs.CV cs.LG cs.MA cs.RO cs.SI 62%

GContextFormer: A global context-aware hybrid multi-head attention approach with scaled additive aggregation for multimodal trajectory prediction

GContextFormer: 一种基于全局上下文的混合多头注意力方法,通过缩放加法聚合实现多模态轨迹预测

Yuzhi Chen, Yuanchang Xie, Lei Zhao, Pan Liu, Yajie Zou, Chen Wang

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 GContextFormer通过全局上下文感知的混合多头注意力和缩放加法聚合,实现无地图依赖的多模态轨迹预测,提升鲁棒性和预测精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12609 2025-11-25 cs.CL cs.AI cs.CV 62%

Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data

Uni-MoE-2.0-Omni: 通过先进MoE、训练和数据扩展语言导向的多模态大模型

Yunxin Li, Xinyu Chen, Shenyuan Jiang, Haoyuan Shi, Zhenyu Liu, Xuanyu Zhang, Nanhao Deng, Zhenran Xu, Yicheng Ma, Meishan Zhang, Baotian Hu, Min Zhang

机构 * Research Institute of Computing and Intelligence(计算与智能研究 institute) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 Uni-MoE-2.0-Omni通过先进MoE、训练和数据技术,实现了语言导向的多模态大模型,展现出在多模态理解、推理和生成任务中的卓越性能。

Comments 47 pages,10 Figures, Project Website: https://idealistxy.github.io/Uni-MoE-v2.github.io/ Codes: https://github.com/HITsz-TMG/Uni-MoE

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18422 2025-11-25 cs.CV cs.LG 57%

NeuroVascU-Net: A Unified Multi-Scale and Cross-Domain Adaptive Feature Fusion U-Net for Precise 3D Segmentation of Brain Vessels in Contrast-Enhanced T1 MRI

NeuroVascU-Net:一种统一的多尺度和跨域自适应特征融合U-Net,用于精确的对比增强T1 MRI脑血管三维分割

Mohammad Jafari Vayeghan, Niloufar Delfan, Mehdi Tale Masouleh, Mansour Parvaresh Rizi, Behzad Moshiri

机构 * School of Electrical and Computer Engineering, College of Engineering, University of Tehran(电信工程学院,工程学院,德黑兰大学) Department of EECS, Lassonde School of Engineering, York University(电子工程系,拉索nde工程学院,约克大学) Department of Neurosurgery, School of Medicine, Iran University of Medical Sciences(神经外科系,医学院,伊朗医学科学大学) Department of Electrical and Computer Engineering, University of Waterloo(电子工程系,滑铁库大学)

专题命中 视觉空间推理 :planning(abstract);分类 cs.LG

AI总结 NeuroVascU-Net通过多尺度和跨域自适应特征融合模块,实现高精度脑血管分割,适用于临床标准T1CE MRI,提升神经外科手术规划的准确性与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17664 2025-11-25 cs.LG cs.CV cs.CY 57%

CubeletWorld: A New Abstraction for Scalable 3D Modeling

CubeletWorld: 一种可扩展3D建模的新抽象

Azlaan Mustafa Samad, Hoang H. Nguyen, Lukas Berg, Henrik Müller, Yuan Xue, Daniel Kudenko, Zahra Ahmadi

机构 * L3S Research Center(L3S研究所以) Leibniz University Hannover(汉诺威莱布尼茨大学) University of Tennessee at Chattanooga(田纳西大学查塔努加分校) PLRI Medical Informatics Institute(医学信息学研究所) Hannover Medical School(汉诺威医学院)

专题命中 视觉空间推理 :planning(abstract);分类 cs.LG

AI总结 CubeletWorld通过离散3D网格空间单元实现可扩展的城市建模,提升隐私保护和跨区域通用性。

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02162 2025-11-25 cs.RO cs.AI cs.HC 57%

Text to Robotic Assembly of Multi Component Objects using 3D Generative AI and Vision Language Models

通过3D生成AI和视觉语言模型实现多组件物体的文本到机器人组装

Alexander Htet Kyaw, Richa Gupta, Dhruv Shah, Anoop Sinha, Kory Mathewson, Stefanie Pender, Sachin Chitta, Yotto Koga, Faez Ahmed, Lawrence Sass, Randall Davis

机构 * Massachusetts Institute of Technology (MIT)(麻省理工学院) MIT(麻省理工学院) Google DeepMind(谷歌DeepMind) Google, Paradigms of Intelligence(谷歌、范式智能) Autodesk Research(Autodesk研究) MIT Mechanical Engineering(麻省理工学院机械工程系) MIT Architecture(麻省理工学院建筑系) MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

AI总结 本文提出利用3D生成AI和视觉语言模型实现多组件物体的文本到机器人组装,通过多模态推理分解生成网格并优化组件分配。

Comments Accepted to NeurIPS 2025, Conference on Neural Information Processing Systems, Creative AI Track

详情

展开后加载摘要…

URL PDF HTML 收藏