arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

共收录 4334 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 4334 篇

2605.04899 2026-05-12 cs.LG 69%

A geometric relation of the error introduced by sampling a language model's output distribution to its internal state

语言模型输出分布采样引入的误差与内部状态的几何关系

Albert F. Modenbach

机构 * Department of Mathematics, King's College London, United Kingdom(伦敦国王学院数学系)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.LG

AI总结 研究揭示语言模型输出分布采样误差与内部状态几何结构的关系,通过几何方法分析token嵌入空间,发现其曲率在棋类任务中反映模型对问题的内部表示。

Comments 12 Pages, 10 Figures, 2 Appendices. To appear in Proceedings of ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08007 2026-05-11 cs.LG 69%

Interpreting Reinforcement Learning Agents with Susceptibilities

用脆弱性解释强化学习智能体

Chris Elliott, Einar Urdshals, David Quarel, Daniel Murfet

机构 * Timaeus

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.LG

AI总结 本文将神经网络可解释性技术扩展到深度强化学习的后悔设定,探讨脆弱性在简单网格世界模型中的应用,揭示模型在参数空间中的内部发展特征。

Comments 55 pages, comments welcome

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06334 2026-05-08 cs.CL cs.LG cs.LO 69%

MANTRA: Synthesizing SMT-Validated Compliance Benchmarks for Tool-Using LLM Agents

MANTRA: 为使用工具的LLM代理合成经过SMT验证的合规性基准

Ashwani Anand, Ivi Chatzi, Ritam Raha, Anne-Kathrin Schmuck

机构 * Max Planck Institute for Software Systems(马克斯·普朗克软件系统研究所)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.LG

AI总结 MANTRA通过自动合成可机检验的合规性基准,解决LLM代理在复杂手册下的合规性验证问题,支持任意领域和长流程手册,并提供可调节的任务复杂度,生成285个任务的基准套件。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28980 2026-05-04 cs.CV 69%

Stepper: Stepwise Immersive Scene Generation with Multiview Panoramas

Stepper:基于多视角全景图的分步沉浸式场景生成

Felix Wimbauer, Fabian Manhardt, Michael Oechsle, Nikolai Kalischek, Christian Rupprecht, Daniel Cremers, Federico Tombari

机构 * Google(谷歌) University of Oxford(牛津大学) MCML Technical University of Munich(慕尼黑技术大学)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.CV

AI总结 Stepper通过分步扩展多视角全景图,解决传统方法在视觉保真度与可探索性之间的权衡问题,实现高质量沉浸式3D场景生成。

Comments Accepted at CVPR 2026 Findings; Find our project page under https://fwmb.github.io/stepper/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27209 2026-05-04 cs.SE cs.AI 69%

Theory Under Construction: Orchestrating Language Models for Research Software Where the Specification Evolves

构建中的理论:协调语言模型用于研究软件,其中规范在演变

Halley Young, Nikolaj Björner

机构 * Microsoft Research(微软研究院)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.AI

AI总结 本文提出Comet-H框架,通过迭代提示自动机协调语言模型在研究软件中的应用,解决规范演变中的 hallucination accumulation 和 desynchronization 问题,并通过实验证明其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04212 2026-05-04 cs.CL cs.AI 69%

Language Models Struggle to Use Representations Learned In-Context

语言模型在上下文学习的表示使用上面临困难

Michael A. Lepori, Tal Linzen, Ann Yuan, Katja Filippova

机构 * Brown University(布朗大学) Google Research(谷歌研究) New York University(纽约大学) Google DeepMind(谷歌DeepMind)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.AI

AI总结 研究探讨了语言模型是否能利用上下文学习的表示完成简单下游任务,发现开放权重模型在处理新语义表示时存在困难,即使它们在潜在表示中编码了这些语义。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27899 2026-05-01 cs.AI 69%

Simulating clinical interventions with a generative multimodal model of human physiology

用生成式多模态模型模拟临床干预

Guy Lutsker, Gal Sapir, Jordi Merino, Smadar Shilo, Anastasia Godneva, Eli Meirom, Shie Mannor, Hagai Rossman, Gal Chechik, Eran Segal

机构 * Department of Computer Science and Applied Mathematics, Weizmann Institute of Science(魏茨曼科学研究所计算机科学与应用数学系) Department of Molecular Cell Biology, Weizmann Institute of Science(魏茨曼科学研究所分子细胞生物学系) NVIDIA Novo Nordisk Foundation Center for Basic Metabolic Research, University of Copenhagen(诺沃维克基金会基础代谢研究中心,哥本哈根大学) Faculty of Medical and Health Sciences, Tel Aviv University(特拉维夫大学医学与健康科学学院) The Jesse Z and Sara Lea Shafer Institute for Endocrinology and Diabetes, National Center for Childhood Diabetes, Schneider Children’s Medical Center of Israel(杰西Z和索菲亚·李·沙弗内分泌学与糖尿病研究所,以色列儿童糖尿病国家中心,施耐德儿童医学中心) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.AI

AI总结 本文提出HealthFormer模型,通过训练人类表型项目数据,生成人类生理轨迹,实现对个体生理变化的预测和干预模拟,提升临床风险评分和疾病预测能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23278 2026-04-28 cs.AI 69%

Active Inference: A method for Phenotyping Agency in AI systems?

主动推断:一种用于AI系统中表型代理的方法?

Philip Wilson, Axel Constant, Mahault Albarracin, Nicolás Hinrichs, Jasmine Moore, Daniel Polani, Karl Friston

机构 * Independent Researcher Laboratoire d'Analyse Cognitive de l'Information, Universit\' e du Qu\' e bec \` a Montr\' e al, Qu\' e bec, Canada Centre of Excellence for AI Robotics, Sheffield Hallam University, Sheffield, UK Methods Statistical Computing, Max Planck Institute for Human Cognitive Brain Sciences, Leipzig, Germany Department of Computer Science, School of Physics, Engineering Computer Science, University of Hertfordshire, Hatfield, UK Wellcome Centre for Human Neuroimaging, University College London, London, UK Department of Engineering Informatics, University of Sussex, Falmer, Brighton, BN1 9RH, UK

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.AI

AI总结 本文提出基于主动推断的表型代理方法,通过变分框架将信念、偏好和自由能最小化结合,以区分不同代理表型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19089 2026-04-28 cs.CL cs.AI 69%

Language Models Might Not Understand You: Evaluating Theory of Mind via Story Prompting

语言模型可能并不理解你:通过故事提示评估理论自我意识

Nathaniel Getachew, Abulhair Saparov

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.AI

AI总结 通过故事提示框架评估大语言模型的理论自我意识和世界建模能力,发现模型在世界建模任务中表现优于理论自我意识任务,且对人物推理更准确。

Comments 21 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20460 2026-04-23 cs.CV 69%

CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs

CCTVBench:用于多模态大语言模型的对比一致性交通视频问答基准

Xingcheng Zhou, Hao Guo, Rui Song, Walter Zimmer, Mingyu Liu, André Schamschurko, Hu Cao, Alois Knoll

机构 * Technical University of Munich(慕尼黑技术大学) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 通用世界模型 :world-model(abstract);world-model(abstract);分类 cs.CV

AI总结 CCTVBench通过真实事故视频与世界模型生成的反事实场景配对,提出对比一致性评估方法,揭示视频问答中对比一致性与标准指标间的显著差距,并引入C-TCD方法提升问答与一致性性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17960 2026-04-21 q-bio.NC cs.LG 69%

The Umwelt Representation Hypothesis: Rethinking Universality

环境表征假设:重新思考普遍性

Victoria Bosch, Rowan Sommers, Adrien Doerig, Tim C Kietzmann

机构 * Institute of Cognitive Science, University of Osnabrück(奥尔登堡大学认知科学研究所) Department of Psychology and Education, Freie Universität Berlin(柏林自由大学心理学与教育系) Bernstein Center for Computational Neuroscience, Berlin(柏林计算神经科学中心)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.LG

AI总结 本文提出环境表征假设,认为系统表征的相似性源于生态约束下的重叠而非单一最优解,挑战了普遍性观点。

Comments preprint v1

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17883 2026-04-21 cs.SE cs.HC cs.LG 69%

Scaling Human-AI Coding Collaboration Requires a Governable Consensus Layer

规模化的人工智能编码协作需要一个可治理的共识层

Tianfu Wang, Zhezheng Hao, Yin Wu, Wei Wu, Qiang Lin, Hande Dong, Nicholas Jing Yuan, Hui Xiong

机构 * Tencent(腾讯)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.LG

AI总结 本文提出Agentic Consensus共识层,通过类型属性图替代代码作为主要工程产物,解决AI辅助开发中代码与聊天历史维度塌陷导致的系统不透明问题,强调通过共识熵衡量协作流程的对齐度和干预距离。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14965 2026-04-17 cs.RO 69%

POMDP-based Object Search with Growing State Space and Hybrid Action Domain

基于增长状态空间和混合动作域的POMDP目标搜索

Yongbo Chen, Hesheng Wang, Shoudong Huang, Hanna Kurniawati

机构 * School of Automation and Intelligent Sensing, Shanghai Jiao Tong University, Shanghai, 200240, People's Republic of China(自动化与智能感知学院,上海交通大学,上海,200240,中华人民共和国) School of Computing, Australian National University (ANU), Canberra, ACT, 2601, Australia(计算学院,澳大利亚国立大学(ANU),堪培拉,ACT,2601,澳大利亚) Key Laboratory of System Control and Information Processing, Ministry of Education of China, Shanghai, 200240, People's Republic of China(系统控制与信息处理重点实验室,中华人民共和国教育部,上海,200240,中华人民共和国) Robotics Institute, University of Technology Sydney, Australia(机器人研究所,悉尼技术大学,澳大利亚)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.RO

AI总结 本文提出GNPF-kCT算法,通过结合神经过程网络和k中心聚类,解决复杂室内环境中移动机器人高效定位目标物体的问题,实验证明其在目标搜索任务中优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09928 2026-04-10 cs.RO 69%

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models

HiF-VLA:通过运动表示实现视觉-语言-动作模型的回顾、洞察与前瞻性

Minghui Lin, Pengxiang Ding, Shu Wang, Zifeng Zhuang, Yang Liu, Xinyang Tong, Wenxuan Song, Shangke Lyu, Siteng Huang, Donglin Wang

机构 * Westlake University(西湖大学) Zhejiang University(浙江大学) HKUST(GZ)(香港科技大学(广州)) Nanjing University(南京大学) Westlake Robotics(西湖机器人)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.RO

AI总结 HiF-VLA通过运动表示提升视觉-语言-动作模型的时序推理能力,采用回顾、洞察和前瞻性机制,在长时域任务中表现优异,且推理延迟低。

Comments CVPR 2026, Project page: https://hifvla.github.io, Github: https://github.com/OpenHelix-Team/HiF-VLA

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10723 2026-04-10 cs.LG 69%

Generalized Spherical Neural Operators: Green's Function Formulation

广义球面神经算子:格林函数公式

Hao Tang, Hao Chen, Chao Li

机构 * University of Dundee, United Kingdom(英国邓迪大学) University of Cambridge, United Kingdom(英国剑桥大学)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.LG

AI总结 本文提出基于可设计球面格林函数及其谐波展开的广义算子设计框架,提出适应非等变系统的GSNO算子,结合多尺度谱模型与球面上下采样构建SHNet,实验证明其在扩散磁共振成像、浅水动力学和全球天气预测中优于现有方法。

Comments ICLR 2026 (International Conference on Learning Representations)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04723 2026-04-07 cs.CL cs.AI 69%

Individual and Combined Effects of English as a Second Language and Typos on LLM Performance

英语作为第二语言与拼写错误的个体与综合影响

Serena Liu, Yutong Yang, Prisha Sheth, Weixuan Dong, Mingjiao Diao, Xinru Zhu, Nikhil Banga, Oscar Melendez, Arnav Sharma, Minda Zhao, Marina Lin, Mengyu Wang

机构 * Harvard University(哈佛大学)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.AI

AI总结 研究探讨了英语作为第二语言和拼写错误对大语言模型性能的影响,发现两者结合时性能下降更显著,且在封闭式任务中表现更一致,而开放式任务结果更混杂。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13920 2026-04-07 cs.LG 69%

Causal Process Models: Reframing Dynamic Causal Graph Discovery as a Reinforcement Learning Problem

因果过程模型:将动态因果图发现重新表述为强化学习问题

Turan Orujlu, Christian Gumbsch, Martin V. Butz, Charley M Wu

机构 * University of Tübingen(图宾根大学) MPI for Biological Cybernetics(马克斯·普朗克生物控制论研究所) University of Amsterdam(阿姆斯特丹大学) TU Darmstadt(达姆施塔特工业大学)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.LG

AI总结 本文提出因果过程模型,通过将动态因果图构建视为多智能体强化学习问题,实现了从视觉观测中学习稀疏时间变化因果图,提升可解释性和计算效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03524 2026-04-07 cs.AI 69%

Structural Rigidity and the 57-Token Predictive Window: A Physical Framework for Inference-Layer Governability in Large Language Models

结构刚性和57个标记的预测窗口:一种用于大语言模型推理层可控性的物理框架

Gregory M. Ruddell

专题命中 通用世界模型 :world-model(abstract);world-model(abstract);分类 cs.AI

AI总结 本文提出一种物理框架,通过分析大语言模型的推理行为,揭示了预提交信号的存在条件,并展示了结构刚性在不同几何区域中的统一度量。

Comments Extends arXiv:2603.21415. 30 pages. Also available on Zenodo (10.5281/zenodo.19393882)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01359 2026-04-03 cs.AI 69%

Semantic Modeling for World-Centered Architectures

面向世界中心架构的语义建模

Andrei Mantsivoda, Darya Gavrilina

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.AI

AI总结 本文提出世界中心多智能体系统(WMAS)作为传统智能体中心架构的替代方案,通过共享世界模型实现语义一致性、可解释性和长期稳定性,并介绍了Ontobox平台作为其实现。

Comments 15 pages, 1 figure, MathAI conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18887 2026-04-01 cs.CV 69%

SafeDrive: Fine-Grained Safety Reasoning for End-to-End Driving in a Sparse World

SafeDrive: 为稀疏世界中的端到端驾驶进行细粒度安全推理

Jungho Kim, Jiyong Oh, Seunghoon Yu, Hongjae Shin, Donghyuk Kwak, Jun Won Choi

机构 * Seoul National University(首尔大学)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.CV

AI总结 SafeDrive提出了一种端到端规划框架,通过轨迹条件化的稀疏世界模型进行显式且可解释的安全推理,实现了在开放循环和闭合循环基准上的最佳性能,有效降低了碰撞率。

Comments Accepted to CVPR 2026, 19 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01342 2026-03-31 cs.CV 69%

InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision

InternVideo-Next:无需视频-文本监督的通用视频基础模型

Chenting Wang, Yuhan Zhu, Yicheng Xu, Jiange Yang, Lang Lin, Ziang Yan, Yali Wang, Yi Wang, Limin Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室) Shanghai Innovation Institute(上海创新研究院) Shenzhen Institutes of Advanced Technology, China(中国科学院深圳先进技术研究院) Nanjing University(南京大学)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.CV

AI总结 本文提出InternVideo-Next,通过解耦传统编码器-解码器设计为Encoder-Predictor-Decoder框架,解决像素重建与语义冲突问题,构建语义一致且细节保留的潜在空间,提升视频基础模型的通用性。

Journal ref CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24835 2026-03-27 cs.CV 69%

DCARL: A Divide-and-Conquer Framework for Autoregressive Long-Trajectory Video Generation

DCARL:一种用于自回归长轨迹视频生成的分而治之框架

Junyi Ouyang, Wenbin Teng, Gonglin Chen, Yajie Zhao, Haiwei Chen

机构 * Institute for Creative Technologies(创意技术研究所) University of Southern California(南加州大学)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.CV

AI总结 本文提出DCARL框架,通过分而治之策略结合VDMs的高保真生成能力,解决长轨迹视频生成中的视觉漂移和可控性问题,实现高质量且稳定的生成。

Comments 29 pages, 11 figures. Project page: https://junyiouy.github.io/projects/dcarl

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16177 2026-03-24 cs.LG 69%

The Finetuner's Fallacy: When to Pretrain with Your Finetuning Data

微调者的谬误:何时应使用微调数据进行预训练

Christina Baek, Ricardo Pio Monti, David Schwab, Amro Abbas, Rishabh Adiga, Cody Blakeney, Maximilian Böther, Paul Burstein, Aldo Gael Carranza, Alvin Deng, Parth Doshi, Vineeth Dorna, Alex Fang, Tony Jiang, Siddharth Joshi, Brett W. Larsen, Jason Chan Lee, Katherine L. Mentzer, Luke Merrick, Haakon Mongstad, Fan Pan, Anshuman Suri, Darren Teh, Jason Telanoff, Jack Urbanek, Zhengping Wang, Josh Wills, Haoli Yin, Aditi Raghunathan, J. Zico Kolter, Bogdan Gaza, Ari Morcos, Matthew Leavitt, Pratyush Maini

机构 * DatologyAI Team(DatologyAI团队)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.LG

AI总结 本文研究了一种简单策略,即专门预训练(SPT),通过在预训练阶段重复使用小领域数据集,以提升领域性能并保留通用能力。实验显示,SPT能减少预训练token数量,提升领域表现,同时降低过拟合风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01641 2026-03-24 cs.CV 69%

FideDiff: Efficient Diffusion Model for High-Fidelity Image Motion Deblurring

FideDiff:高效的高保真图像运动去模糊扩散模型

Xiaoyang Liu, Zhengyan Zhou, Zihang Xu, Jiezhang Cao, Zheng Chen, Yulun Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Harvard University(哈佛大学)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.CV

AI总结 本文提出FideDiff,一种高效的单步扩散模型,用于高保真图像运动去模糊。通过将运动去模糊转化为扩散过程,结合Kernel ControlNet和自适应时间步预测,提升了去模糊性能。

Comments Accepted to ICLR 2026. Code is available at https://github.com/xyLiu339/FideDiff

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16306 2026-03-18 cs.CV 69%

DriveFix: Spatio-Temporally Coherent Driving Scene Restoration

DriveFix:时空一致的驾驶场景修复

Heyu Si, Brandon James Denis, Muyang Sun, Dragos Datcu, Yaoru Li, Xin Jin, Ruiju Fu, Yuliia Tatarinova, Federico Landi, Jie Song, Mingli Song, Qi Guo

机构 * Zhejiang University(浙江大学) Huawei(华为)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.CV

AI总结 DriveFix提出一种多视角修复框架,通过交错扩散变换器架构和几何感知损失,实现驾驶场景时空一致性,提升4D世界建模的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14087 2026-03-17 cs.LG cs.CL 69%

Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors

理解下一token预测器中看似无用特征的出现

Mark Rofin, Jalal Naghiyev, Michael Hahn

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.LG

AI总结 研究探讨了训练中的Transformer模型如何计算看似冗余的抽象特征,并提出方法分析梯度信号对特定特征形成的影响,通过玩具任务和小型模型验证,揭示了预训练语言模型中形式推理领域特征的产生机制。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26796 2026-03-13 cs.CV cs.GR 69%

See4D: Pose-Free 4D Generation via Auto-Regressive Video Inpainting

See4D: 通过自回归视频修复实现无姿态4D生成

Dongyue Lu, Ao Liang, Tianxin Huang, Xiao Fu, Yuyang Zhao, Baorui Ma, Liang Pan, Wei Yin, Lingdong Kong, Wei Tsang Ooi, Ziwei Liu

机构 * NUS(新加坡国立大学) HKU(香港大学) CUHK(香港中文大学) THU(清华大学) Shanghai AI Lab(上海人工智能实验室) Horizon Robotics(地平线机器人) NTU(国立科技大学)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.CV

AI总结 See4D通过自回归视频修复实现无姿态4D生成,提升从随意视频到4D世界建模的实用性。

Comments Eurographics2026; 26 pages; 21 figures; 3 tables; project page: https://see-4d.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09285 2026-03-11 cs.CV 69%

Learning Convex Decomposition via Feature Fields

通过特征场学习凸分解

Yuezhi Yang, Qixing Huang, Mikaela Angelina Uy, Nicholas Sharp

机构 * NVIDIA The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.CV

AI总结 本文提出通过特征场学习实现高质量的开放世界凸分解,解决了长期存在的凸分解问题,并实现了首个自监督学习的开放世界模型。

Comments 14 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23488 2026-03-10 cs.AI cs.CL 69%

Mapping Overlaps in Benchmarks through Perplexity in the Wild

通过在野 perplexity 映射重叠关系

Siyang Wu, Honglin Bao, Sida Li, Ari Holtzman, James A. Evans

机构 * Data Science Institute, University of Chicago(芝加哥大学数据科学研究所)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.AI

AI总结 通过分析LLM基准测试的perplexity,揭示了不同任务间的重叠结构及LLM能力差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03515 2026-03-05 cs.CY cs.AI 69%

The Controllability Trap: A Governance Framework for Military AI Agents

可控性陷阱:军事AI代理的治理框架

Subramanyam Sahoo

机构 * MARS (Mentorship for Alignment Researchers) 4.0 Fellow(MARS(对齐研究导师计划)4.0 Fellow) Cambridge AI Safety Hub (CAISH) University of Cambridge(剑桥AI安全中心(CAISH)剑桥大学)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.AI

AI总结 本文提出AMAGF框架,通过预防、检测和纠正三个支柱,解决军事AI代理中的控制失效问题,通过控制质量评分实现持续控制管理。

Comments Accepted at ICLR 2026 Workshop on Agents in the Wild. 20 Pages and 3 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏