arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-12-05 至 2025-12-05 共收录 12 信号源:cs.CL, cs.AI, cs.LG

1. 规划推理 12 篇

2512.05112 2025-12-05 cs.CV cs.AI cs.CL cs.LG 88%

DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation

DraCo: 文本到图像预览与稀有概念生成的草稿作为CoT

Dongzhi Jiang, Renrui Zhang, Haodong Li, Zhuofan Zong, Ziyu Guo, Jun He, Claire Guo, Junyan Ye, Rongyao Fang, Weijia Li, Rui Liu, Hongsheng Li

机构 * CUHK MMLab(香港中文大学 MMLab) CUHK IMIXR(香港中文大学 IMIXR) Sun Yat-Sen University(中山大学) SCUT(华南理工大学) CUHK (Shenzhen)(香港中文大学(深圳))

专题命中 规划推理 :CoT(title,abstract);reasoning(abstract);chain-of-thought(abstract);planning(abstract)

AI总结 DraCo通过结合文本和视觉内容的交错推理,提升文本到图像生成的精度和稀有概念生成能力。

Comments Project Page: https://github.com/CaraJ7/DraCo

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01926 2025-12-05 cs.AI cs.CL cs.LG 88%

Large language models can learn and generalize steganographic chain-of-thought under process supervision

大语言模型可以在过程监督下学习和泛化隐写链式思维

Joey Skaf, Luis Ibanez-Lissen, Robert McCarthy, Connor Watts, Vasil Georgiv, Hannes Whittingham, Lorena Gonzalez-Manzano, David Lindner, Cameron Tice, Edward James Young, Puria Radmard

机构 * Mentorship for Alignment Research Students (MARS)(对齐研究 mentorship 项目) University College London(伦敦大学学院) Queen Mary University of London(伦敦女王学院) ML Alignment & Theory Scholars (MATS)(对齐与理论学者) Meridian Impact, Cambridge(剑桥 Meridian Impact) Universidad Carlos III de Madrid(马德里卡洛斯三世大学) Geodesic Research and University of Cambridge(Geodesic Research 和剑桥大学)

专题命中 规划推理 :chain-of-thought(title,abstract);reasoning(abstract);CoT(abstract);planning(abstract)

AI总结 大语言模型在过程监督下能够学习并泛化隐写链式思维,通过替换特定字符串实现推理编码,提升监控可靠性。

Comments 10 pages main text, 3 figures main text, 17 pages supplementary material, 1 figure supplementary material, accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04459 2025-12-05 cs.CV 82%

dVLM-AD: Enhance Diffusion Vision-Language-Model for Driving via Controllable Reasoning

dVLM-AD:通过可控推理增强扩散视觉语言模型以实现驾驶

Yingzi Ma, Yulong Cao, Wenhao Ding, Shuibai Zhang, Yan Wang, Boris Ivanovic, Ming Jiang, Marco Pavone, Chaowei Xiao

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) NVIDIA(英伟达) Stanford University(斯坦福大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 规划推理 :reasoning(title,abstract);planning(abstract)

AI总结 dVLM-AD通过可控推理提升扩散视觉语言模型,实现更一致的驾驶推理与规划性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13983 2025-12-05 cs.MA cs.RO cs.SY eess.SY 82%

Learning Two-agent Motion Planning Strategies from Generalized Nash Equilibrium for Model Predictive Control

从广义纳什均衡学习双智能体运动规划策略用于模型预测控制

Hansung Kim, Edward L. Zhu, Chang Seok Lim, Francesco Borrelli

机构 * University of California Berkeley(加州大学伯克利分校) PlusAI Inc(PlusAI公司)

专题命中 规划推理 :planning(title,abstract);reasoning(abstract)

AI总结 本文提出了一种基于广义纳什均衡的学习方法,用于多智能体运动规划中的模型预测控制,通过训练神经网络预测奖励结果来实现博弈论互动的隐式处理。

Comments Accepted Proceeding at 2025 Learning for Dynamics and Control Conference (L4DC)

Journal ref 283:112-123, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04405 2025-12-05 eess.SP cs.AI 70%

Towards 6G Native-AI Edge Networks: A Semantic-Aware and Agentic Intelligence Paradigm

迈向6G原生AI边缘网络:一种语义感知和代理智能范式

Chenyuan Feng, Anbang Zhang, Geyong Min, Yongming Huang, Tony Q. S. Quek, Xiaohu You

机构 * College of Computer Science, University of Exeter, Exeter, U.K.(埃克塞特大学计算机科学学院) Sony China Research Laboratory, Beijing, China(索尼中国研究实验室) School of Information Science and Engineering, Southeast University, Nanjing, China(东南大学信息科学与工程学院) Purple Mountain Laboratories, Nanjing, China(紫金山实验室) Information System Technology and Design Pillar, Singapore University of Technology and Design, Singapore(新加坡科技设计大学信息系统技术与设计支柱)

专题命中 规划推理 :reasoning(abstract);planning(abstract);分类 cs.AI

AI总结 本文提出语义感知和代理智能范式,旨在通过统一分类法和促进技术推动6G原生AI边缘网络的发展。

Comments submitted to Digital Communications and Networks

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05066 2025-12-05 cs.LG cs.AI cs.CL 67%

Multi-LLM Collaboration for Medication Recommendation

多LLM协作用于药物推荐

Huascar Sanchez, Briland Hitaj, Jules Bergmann, Linda Briesemeister

机构 * Computer Science Laboratory, SRI International(SRI国际计算机科学实验室) University of Maryland St. Joseph Medical Center(马里兰大学圣约瑟夫医疗中心)

专题命中 规划推理 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出基于LLM化学的多模型协作方法,通过增强互补性、稳定性和校准性,提高药物推荐的可靠性与可信度。

Comments 8 pages, 5 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04451 2025-12-05 cs.CV 67%

StreamEQA: Towards Streaming Video Understanding for Embodied Scenarios

StreamEQA:迈向具身场景的流式视频理解

Yifei Wang, Zhenkai Li, Tianwen Qian, Huanran Zheng, Zheng Wang, Yuqian Fu, Xiaoling Wang

机构 * School of Computer Science and Technology, East China Normal University(计算机科学与技术学院,华东师范大学) Zhejiang University of Technology(浙江工业大学)

专题命中 规划推理 :reasoning(abstract);planning(abstract)

AI总结 StreamEQA是一个针对具身场景中流式视频问答的基准测试,通过评估模型在感知、交互和规划方面的能力,推动流式视频理解在具身应用中的研究。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04785 2025-12-05 cs.AI cs.CR 57%

ASTRIDE: A Security Threat Modeling Platform for Agentic-AI Applications

ASTRIDE:面向智能体AI应用的安全威胁建模平台

Eranga Bandara, Amin Hass, Ross Gore, Sachin Shetty, Ravi Mukkamala, Safdar H. Bouk, Xueping Liang, Ng Wee Keong, Kasun De Zoysa, Aruna Withanage, Nilaan Loganathan

机构 * Old Dominion University(老奥布里恩大学) Accenture Technology Labs(埃森哲技术实验室) Florida International University(佛罗里达国际大学) Nanyang Technological University(南洋理工大学) University of Colombo(科伦坡大学)

专题命中 规划推理 :reasoning(abstract);分类 cs.AI

AI总结 ASTRIDE是首个专为AI智能体应用设计的自动化威胁建模平台,通过扩展STRIDE框架并结合微调的视觉语言模型与推理LLM,实现基于图的威胁建模自动化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09080 2025-12-05 cs.AI 57%

BioAnalyst: A Foundation Model for Biodiversity

BioAnalyst: 一种用于生物多样性的基础模型

Athanasios Trantas, Martino Mensio, Stylianos Stasinos, Sebastian Gribincea, Taimur Khan, Damian Podareanu, Aliene van der Veen

机构 * Nederlandse Organisatie voor Toegepast Natuurwetenschappelijk Onderzoek – TNO(荷兰应用自然科学研究院 – TNO) Eindhoven University of Technology(埃因霍温理工大学) Amazon(亚马逊) University of Groningen(格罗宁根大学) Helmholtz Center for Environmental Research – UFZ(环境研究赫尔姆霍茨中心 – UFZ) SURF

专题命中 规划推理 :planning(abstract);分类 cs.AI

AI总结 BioAnalyst是一种针对生物多样性分析和保护规划的多模态基础模型,通过预训练和微调实现物种分布建模和生态预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04492 2025-12-05 cs.CL 57%

MSME: A Multi-Stage Multi-Expert Framework for Zero-Shot Stance Detection

MSME: 一种多阶段多专家框架用于零样本立场检测

Yuanshuo Zhang, Aohua Li, Bo Chen, Jingbo Sun, Xiaobing Zhao

机构 * footnotemark: 1(机构1)

专题命中 规划推理 :reasoning(abstract);分类 cs.CL

AI总结 MSME提出了一种多阶段多专家框架,通过知识准备、专家推理和决策聚合三个阶段,提升零样本立场检测的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22037 2025-12-05 cs.CY 50%

What AI Speaks for Your Community: Polling AI Agents for Public Opinion on Data Center Projects

人工智能为你的社区发声:通过AI代理收集数据中心项目公众意见

Zhifeng Wu, Yuelin Han, Shaolei Ren

专题命中 规划推理 :planning(abstract)

AI总结 本文提出AI代理调查框架,利用大型语言模型评估社区对数据中心项目的意见,以指导负责任的AI发展。

Comments 35 Pages. Accepted to NeurIPS 2025 Workshop on Socially Responsible and Trustworthy Foundation Models (ResponsibleFM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21164 2025-12-05 cs.HC 50%

Preference-Aligned Options from Generative AI Compensates for Age-Related Cognitive Decline in Decision Making

基于生成式AI的偏好对齐选项可缓解老年人决策中的年龄相关认知下降

Sayaka Ishibashi, Kou Tamura, Ayana Goma, Kenta Yamamoto, Kouhei Masumoto

专题命中 规划推理 :reasoning(abstract)

AI总结 本研究发现,生成式AI通过减少信息搜索的认知负荷,能缓解老年人决策中的年龄相关认知下降问题。

详情

展开后加载摘要…

URL PDF HTML 收藏