arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

共收录 3146 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 具身推理 3146 篇

2512.03058 2025-12-04 cs.LG 57%

Dynamical Properties of Tokens in Self-Attention and Effects of Positional Encoding

自注意力机制中令牌的动力学特性及位置编码的影响

Duy-Tung Pham, An The Nguyen, Viet-Hoang Tran, Nhan-Phu Chung, Xin T. Tong, Tan M. Nguyen, Thieu N. Vo

机构 * FPT Software AI Center(FPT软件AI中心) National University of Singapore(国立新加坡大学) Ho Chi Minh University of Economics(胡志明经济大学)

专题命中 具身推理 :world model(abstract);分类 cs.LG

AI总结 本文研究了Transformer中令牌的动力学特性,揭示了位置编码对模型性能的影响,并提出了改进方法以缓解收敛问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02457 2025-12-04 cs.CV 57%

Does Hearing Help Seeing? Investigating Audio-Video Joint Denoising for Video Generation

听觉有助于视觉吗?研究音频视频联合去噪用于视频生成

Jianzong Wu, Hao Lian, Dachao Hao, Ye Tian, Qingyu Shi, Biaolong Chen, Hao Jiang, Yunhai Tong

机构 * Peking University(北京大学) Alibaba Group(阿里巴巴集团)

专题命中 具身推理 :world model(abstract);分类 cs.CV

AI总结 本文提出通过音频视频联合去噪提升视频生成质量,验证了跨模态训练对构建更物理基础的世界模型的潜力。

Comments Project page at https://jianzongwu.github.io/projects/does-hearing-help-seeing/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02016 2025-12-02 cs.CV 57%

Objects in Generated Videos Are Slower Than They Appear: Models Suffer Sub-Earth Gravity and Don't Know Galileo's Principle...for now

生成视频中的物体比看起来更慢:模型遭受次地球重力并目前还不知道伽利略原理...

Varun Varma Thozhiyoor, Shivam Tripathi, Venkatesh Babu Radhakrishnan, Anand Bhattad

机构 * Indian Institute of Science(印度科学研究院) Johns Hopkins University(约翰霍普金斯大学)

专题命中 具身推理 :world model(abstract);分类 cs.CV

AI总结 研究发现视频生成器在重力表示上存在偏差,通过针对性适配可提升其重力模拟精度。

Comments https://gravity-eval.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01878 2025-12-02 cs.AI 57%

Graph Distance as Surprise: Free Energy Minimization in Knowledge Graph Reasoning

图距离作为惊喜:知识图谱推理中的自由能最小化

Gaganpreet Jhajj, Fuhua Lin

机构 * School of Computing(计算学院) Information Systems(信息系统) Athabasca University(亚伯达大学)

专题命中 具身推理 :world model(abstract);分类 cs.AI

AI总结 本文提出利用图距离最小化惊喜来改进知识图谱推理,通过连接自由能原理与KG系统,探索图距离在生成模型中的应用及其对语法结构的影响。

Comments Accepted to NORA Workshop at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10194 2025-12-02 cs.CV 57%

B2N3D: Progressive Learning from Binary to N-ary Relationships for 3D Object Grounding

B2N3D: 从二元关系到N元关系的3D物体接地的渐进式学习

Feng Xiao, Hongbin Xu, Hai Ci, Wenxiong Kang

机构 * School of Automation Science and Engineering, South China University of Technology(自动化科学与工程学院,华南理工大学) ByteDance Seed(字节跳动种子) Show Lab, National University of Singapore(Show Lab,新加坡国立大学)

专题命中 具身推理 :robotic(abstract);分类 cs.CV

AI总结 B2N3D通过引入N元关系学习提升3D物体接地的准确性,利用分组监督损失和混合注意力机制实现更精确的多模态关系建模。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11673 2025-12-01 cs.CV 57%

DINO-Foresight: Looking into the Future with DINO

DINO-Foresight: 用DINO展望未来

Efstathios Karypidis, Ioannis Kakogeorgiou, Spyros Gidaris, Nikos Komodakis

机构 * Archimedes, Athena Research Center(阿基米德研究所,雅典研究学院) National Technical University of Athens(雅典国立技术大学) University of Crete(克里特大学) IACM-Forth(福伊恩研究所)

专题命中 具身推理 :robotics(abstract);分类 cs.CV

AI总结 DINO-Foresight通过在预训练视觉基础模型的语义特征空间中训练,实现了对未来动态的预测,提升了自动驾驶和机器人领域的环境理解能力。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21925 2025-12-01 cs.RO 57%

OpenTwinMap: An Open-Source Digital Twin Generator for Urban Autonomous Driving

OpenTwinMap: 一种用于城市自动驾驶的开源数字孪生生成器

Alex Richardson, Jonathan Sprinkle

机构 * Vanderbilt University(范德比大学)

专题命中 具身推理 :world model(abstract);分类 cs.RO

AI总结 OpenTwinMap是一种基于Python的开源框架,用于生成高保真的3D城市数字孪生,通过整合LiDAR和OSM数据,支持自动驾驶模拟和扩展性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12718 2025-11-27 cs.CV 57%

EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVer

EvoEmpirBench: 基于智能体经验的动态空间推理

Pukun Zhao, Longxiang Wang, Miaowei Wang, Chen Chen, Fanqing Zhou, Haojian Huang

机构 * Guangdong University of Finance and Economics(广东金融学院) Chongqing University(重庆大学) University of Edinburgh(爱丁堡大学) The University of Hong Kong(香港大学)

专题命中 具身推理 :navigation(abstract);分类 cs.CV

AI总结 EvoEmpirBench通过动态空间基准评估模型在部分可观测和动态环境下的空间推理与记忆能力,揭示主流模型的限制并提供未来研究平台。

Comments Accepted by AAAI 2026, 29 pages, 3 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18922 2025-11-25 cs.CV 57%

One4D: Unified 4D Generation and Reconstruction via Decoupled LoRA Control

One4D:通过解耦LoRA控制实现统一的4D生成与重建

Zhenxing Mi, Yuxin Wang, Dan Xu

机构 * The Hong Kong University of Science and Technology(香港科技大学)

专题命中 具身推理 :world model(abstract);分类 cs.CV

AI总结 One4D通过解耦LoRA控制实现4D生成与重建的统一框架,提升视频扩散模型在几何重建中的性能

Comments Project page: https://mizhenxing.github.io/One4D

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16567 2025-11-24 cs.CV 57%

POMA-3D: The Point Map Way to 3D Scene Understanding

POMA-3D:点地图方式的3D场景理解

Ye Mao, Weixun Luo, Ranran Huang, Junpeng Jing, Krystian Mikolajczyk

机构 * Imperial College London(伦敦帝国学院)

专题命中 具身推理 :navigation(abstract);分类 cs.CV

AI总结 POMA-3D通过点地图实现3D场景理解,结合自监督学习和多视角对齐策略,提升3D任务表现。

Comments 11 pages, 6 tables, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16432 2025-11-21 cs.AI q-bio.NC 57%

From generative AI to the brain: five takeaways

从生成式AI到大脑:五点启示

Claudius Gros

专题命中 具身推理 :world model(abstract);分类 cs.AI

AI总结 本文探讨生成式AI与大脑之间的潜在联系,通过五个例子展示神经科学可能从机器学习研究中获得的启示。

Comments Frontiers in Computational Neuroscience, in press

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08978 2025-11-13 cs.MM cs.CV 57%

Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding

Jingtian Ma, Jingyuan Wang, Wayne Xin Zhao, Guoping Liu, Xiang Wen

机构 * School of Computer Science and Engineering, and the MOE Engineering Research Center of Advanced Computer Application Technology, Beihang University(计算机科学与工程学院,以及教育部先进计算机应用技术工程研究中心,北京航空航天大学) School of Computer Science and Engineering, the School of Economics and Management, and the MIIT Key Laboratory of Data Intelligence and Management, Beihang University(计算机科学与工程学院,经济管理学院,以及工信部数据智能与管理重点实验室,北京航空航天大学) Gaoling School of Artificial Intelligence, Renmin University of China(中关村人工智能学院,中国人民大学) DiDi Global Inc.(滴滴出行公司)

专题命中 具身推理 :navigation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17685 2025-11-12 cs.CV 57%

FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving

Shuang Zeng, Xinyuan Chang, Mengwei Xie, Xinran Liu, Yifan Bai, Zheng Pan, Mu Xu, Xing Wei, Ning Guo

机构 * Xi’an Jiaotong University(西安交通大学) Amap, Alibaba Group(阿里巴巴集团) DAMO Academy, Alibaba Group(阿里巴巴达摩院)

专题命中 具身推理 :world model(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025 as Spotlight Presentation. Code: https://github.com/MIV-XJTU/FSDrive

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04670 2025-11-07 cs.CV 57%

Cambrian-S: Towards Spatial Supersensing in Video

Shusheng Yang, Jihan Yang, Pinzhi Huang, Ellis Brown, Zihao Yang, Yue Yu, Shengbang Tong, Zihan Zheng, Yifan Xu, Muhan Wang, Daohan Lu, Rob Fergus, Yann LeCun, Li Fei-Fei, Saining Xie

机构 * New York University(纽约大学) Stanford University(斯坦福大学)

专题命中 具身推理 :world model(abstract);分类 cs.CV

Comments Website: https://cambrian-mllm.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02824 2025-11-06 cs.AI 57%

Kosmos: An AI Scientist for Autonomous Discovery

Ludovico Mitchener, Angela Yiu, Benjamin Chang, Mathieu Bourdenx, Tyler Nadolski, Arvis Sulovari, Eric C. Landsness, Daniel L. Barabasi, Siddharth Narayanan, Nicky Evans, Shriya Reddy, Martha Foiani, Aizad Kamal, Leah P. Shriver, Fang Cao, Asmamaw T. Wassie, Jon M. Laurent, Edwin Melville-Green, Mayk Caldas, Albert Bou, Kaleigh F. Roberts, Sladjana Zagorac, Timothy C. Orr, Miranda E. Orr, Kevin J. Zwezdaryk, Ali E. Ghareeb, Laurie McCoy, Bruna Gomes, Euan A. Ashley, Karen E. Duff, Tonio Buonassisi, Tom Rainforth, Randall J. Bateman, Michael Skarlinski, Samuel G. Rodriques, Michaela M. Hinks, Andrew D. White

专题命中 具身推理 :world model(abstract);分类 cs.AI

Comments Revision: figure layout changes and minor text edits

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07491 2025-11-06 cs.CV 57%

SpatialLM: Training Large Language Models for Structured Indoor Modeling

Yongsen Mao, Junhao Zhong, Chuan Fang, Jia Zheng, Rui Tang, Hao Zhu, Ping Tan, Zihan Zhou

机构 * Manycore Tech Inc.(Manycore科技公司) Hong Kong University of Science and Technology(香港科技大学)

专题命中 具身推理 :robotics(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01618 2025-11-04 cs.CV cs.CL 57%

Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models

Xiaoyu Zhan, Wenxuan Huang, Hao Sun, Xinyu Fu, Changfeng Ma, Shaosheng Cao, Bohan Jia, Shaohui Lin, Zhenfei Yin, Lei Bai, Wanli Ouyang, Yuanqi Li, Jie Guo, Yanwen Guo

机构 * Nanjing University(南京大学) Xiaohongshu Inc.(小红书公司) East China Normal University(华东师范大学) The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学) University of Oxford(牛津大学)

专题命中 具身推理 :robotics(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00706 2025-10-31 cs.CR cs.CL cs.LG 57%

Model Provenance Testing for Large Language Models

Ivica Nikolic, Teodora Baluta, Prateek Saxena

机构 * National University of Singapore(新加坡国立大学) Georgia Institute of Technology(佐治亚理工学院)

专题命中 具身推理 :world model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25129 2025-10-30 cs.CV 57%

AtlasGS: Atlanta-world Guided Surface Reconstruction with Implicit Structured Gaussians

Xiyu Zhang, Chong Bao, Yipeng Chen, Hongjia Zhai, Yitong Dong, Hujun Bao, Zhaopeng Cui, Guofeng Zhang

机构 * State Key Lab of CAD & CG, Zhejiang University(计算机辅助设计与图形学国家重点实验室,浙江大学)

专题命中 具身推理 :world model(abstract);分类 cs.CV

Comments 18 pages, 11 figures. NeurIPS 2025; Project page: https://zju3dv.github.io/AtlasGS/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24459 2025-10-29 cs.AI cs.MA cs.SE 57%

Affordance Representation and Recognition for Autonomous Agents

Habtom Kahsay Gidey, Niklas Huber, Alexander Lenz, Alois Knoll

机构 * Technische Universität München(慕尼黑技术大学) Jessy Works(杰西工作)

专题命中 具身推理 :world model(abstract);分类 cs.AI

Journal ref The Second International Workshop on Hypermedia Multi-Agent Systems (HyperAgents 2025), in conjunction with the 28th European Conference on Artificial Intelligence (ECAI 2025); October 26, 2025, Bologna, Italy

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22200 2025-10-29 cs.CV 57%

LongCat-Video Technical Report

Meituan LongCat Team, Xunliang Cai, Qilong Huang, Zhuoliang Kang, Hongyu Li, Shijun Liang, Liya Ma, Siyu Ren, Xiaoming Wei, Rixu Xie, Tong Zhang

机构 * Meituan LongCat Team(美团LongCat团队)

专题命中 具身推理 :world model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21682 2025-10-27 cs.CV cs.GR 57%

WorldGrow: Generating Infinite 3D World

Sikuang Li, Chen Yang, Jiemin Fang, Taoran Yi, Jia Lu, Jiazhong Cen, Lingxi Xie, Wei Shen, Qi Tian

机构 * MoE Key Lab of Artificial Intelligence, School of Computer Science, SJTU(人工智能前沿实验室,计算机科学学院,上海交通大学) Huawei Inc.(华为公司) Huazhong University of Science and Technology(华中科技大学)

专题命中 具身推理 :world model(abstract);分类 cs.CV

Comments Project page: https://world-grow.github.io/ Code: https://github.com/world-grow/WorldGrow

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17393 2025-10-22 cs.AI cs.CL 57%

Program Synthesis via Test-Time Transduction

Kang-il Lee, Jahyun Koo, Seunghyun Yoon, Minbeom Kim, Hyukhun Koh, Dongryeol Lee, Kyomin Jung

机构 * Dept. of ECE, Seoul National University(电子工程系,首尔国立大学) IPAI, Seoul National University(IPAI,首尔国立大学) Adobe Research(Adobe研究院)

专题命中 具身推理 :world model(abstract);分类 cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15422 2025-10-20 stat.ML cs.LG 57%

Information Theory in Open-world Machine Learning Foundations, Frameworks, and Future Direction

Lin Wang

机构 * Shenzhen Key Laboratory of Neuropsychiatric Modulation(深圳心理行为调控重点实验室) Shenzhen-Hong Kong Institute of Brain Science(深圳-香港脑科学研究院) Shenzhen Institutes of Advanced Technology(深圳先进技术研究院) Chinese Academy of Sciences(中国科学院)

专题命中 具身推理 :world model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07944 2025-10-17 cs.CV 57%

CVD-STORM: Cross-View Video Diffusion with Spatial-Temporal Reconstruction Model for Autonomous Driving

Tianrui Zhang, Yichen Liu, Zilin Guo, Yuxin Guo, Jingcheng Ni, Chenjing Ding, Dan Xu, Lewei Lu, Zehuan Wu

机构 * Sensetime Research(商汤科技研究院) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 具身推理 :world model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11649 2025-10-14 cs.CV 57%

PhySIC: Physically Plausible 3D Human-Scene Interaction and Contact from a Single Image

Pradyumna Yalandur Muralidhar, Yuxuan Xue, Xianghui Xie, Margaret Kostyrko, Gerard Pons-Moll

机构 * University of Tübingen, Zuse School ELIZA(图宾根大学Zuse学校ELIZA) University of Tübingen, Tübingen AI Center(图宾根大学图宾根人工智能中心) University of Tübingen, Tübingen AI Center, MPI for Informatics, SIC(图宾根大学图宾根人工智能中心、马克斯·普朗克信息研究所、SIC)

专题命中 具身推理 :robotics(abstract);分类 cs.CV

Comments Accepted to ACM SIGGraphAsia 2025. Project website: https://yuxuan-xue.com/physic

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07417 2025-10-10 cs.RO 57%

FLEET: Formal Language-Grounded Scheduling for Heterogeneous Robot Teams

Corban Rivera, Grayson Byrd, Meghan Booker, Bethany Kemp, Allison Gaines, Emma Holmes, James Uplinger, Celso M de Melo, David Handelman

机构 * JHU APL(约翰霍普金斯大学应用物理实验室) JHU(约翰霍普金斯大学) DEVCOM ARL(国防高级研究计划局)

专题命中 具身推理 :world model(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25282 2025-10-09 cs.AI cs.HC cs.SE 57%

Toward Causal-Visual Programming: Enhancing Agentic Reasoning in Low-Code Environments

Jiexi Xu, Jiaqi Liu, Lanruo Wang, Su Liu

机构 * School of Information \& Computer Science University of California, Irvine Irvine, CA, USA Independent Researcher University of Texas at Dallas Dallas, TX, USA Georgia Institute of Technology Atlanta, GA, USA

专题命中 具身推理 :world model(abstract);分类 cs.AI

Comments 5 pages, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02287 2025-10-03 cs.CV 57%

MultiModal Action Conditioned Video Generation

Yichen Li, Antonio Torralba

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)

专题命中 具身推理 :world model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02110 2025-10-03 cs.SD cs.LG eess.AS 57%

SoundReactor: Frame-level Online Video-to-Audio Generation

Koichi Saito, Julian Tanke, Christian Simon, Masato Ishii, Kazuki Shimada, Zachary Novack, Zhi Zhong, Akio Hayakawa, Takashi Shibuya, Yuki Mitsufuji

专题命中 具身推理 :world model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏