arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 11120 信号源:cs.CL, cs.AI, cs.LG

1. 规划推理 11120 篇

2511.18165 2025-11-25 cs.SE cs.AI 61%

Towards a General Framework for HTN Modeling with LLMs

迈向基于LLM的HTN建模通用框架

Israel Puerta-Merino, Carlos Núñez-Molina, Pablo Mesejo, Juan Fernández-Olivares

专题命中 规划推理 :planning(abstract,comments);分类 cs.AI

AI总结 本文提出L2HP框架,用于改进LLM在分层规划建模中的能力,通过实验发现分层规划的建模难度显著高于自动规划。

Comments 10 pages, 5 figures, to be published in the Workshop on Planning in the Era of LLMs ( LM4Plan - https://llmforplanning.github.io ) and the Workshop on Hierarchical Planning ( HPlan - https://icaps25.icaps-conference.org/program/workshops/hplan/ ), both in the International Conference on Automated Planning and Scheduling (ICAPS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07220 2025-09-10 cs.AI 61%

OmniAcc: Personalized Accessibility Assistant Using Generative AI

Siddhant Karki, Ethan Han, Nadim Mahmud, Suman Bhunia, John Femiani, Vaskar Raychoudhury

机构 * Miami University of Ohio(俄亥俄迈阿密大学)

专题命中 规划推理 :planning(abstract,comments);分类 cs.AI

Comments 11 Pages, 9 Figures, Published in the 1st Workshop on AI for Urban Planning, AAAI 2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.02167 2025-01-07 cs.AI 61%

Towards a Unified Framework for Sequential Decision Making

Carlos Núñez-Molina, Pablo Mesejo, Juan Fernández-Olivares

专题命中 规划推理 :planning(abstract,journal_ref);分类 cs.AI

Comments 10 pages, 0 figures

Journal ref Carlos Núñez Molina, Pablo Mesejo, & Juan Fernández-Olivares. (2023). Towards a Unified Framework for Sequential Decision Making. In ICAPS PRL Workshop on Bridging the Gap Between AI Planning and Reinforcement Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17109 2024-09-26 cs.CV cs.AI 61%

Unveiling Ontological Commitment in Multi-Modal Foundation Models

Mert Keser, Gesina Schwalbe, Niki Amini-Naieni, Matthias Rottmann, Alois Knoll

专题命中 规划推理 :reasoning(abstract,comments);分类 cs.AI

Comments Qualitative Reasoning Workshop 2024 (QR2024) colocated with ECAI2024, camera-ready submission; first two authors contributed equally; 10 pages, 4 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.10033 2024-06-06 cs.AI cs.GT cs.MA cs.RO 61%

Solution Concepts in Hierarchical Games under Bounded Rationality with Applications to Autonomous Driving

Atrisha Sarkar, Krzysztof Czarnecki

专题命中 规划推理 :planning(abstract,comments);分类 cs.AI

Comments Behavioral Game Theory, Motion and Path Planning, Simulating Humans

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 35(6), 5698-5708 (2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.04635 2024-03-11 cs.AI 61%

Playing Angry Birds with a Domain-Independent PDDL+ Planner

Wiktor Piotrowski, Roni Stern, Matthew Klenk, Alexandre Perez, Shiwali Mohan, Johan de Kleer, Jacob Le

专题命中 规划推理 :planning(abstract,journal_ref);分类 cs.AI

Comments 2 pages, submitted to ICAPS 2021 Demonstration Track

Journal ref Proceedings of the International Conference on Automated Planning and Scheduling (2021) Demonstration Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.07222 2023-06-08 cs.CL 61%

Saliency Map Verbalization: Comparing Feature Importance Representations from Model-free and Instruction-based Methods

Nils Feldhus, Leonhard Hennig, Maximilian Dustin Nasert, Christopher Ebert, Robert Schwarzenberg, Sebastian Möller

专题命中 规划推理 :reasoning(abstract,comments);分类 cs.CL

Comments ACL 2023 Workshop on Natural Language Reasoning and Structured Explanations (NLRSE)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10838 2023-03-21 cs.AI cs.MA 61%

Deceptive Reinforcement Learning in Model-Free Domains

Alan Lewis, Tim Miller

专题命中 规划推理 :planning(abstract,comments);分类 cs.AI

Comments 8 pages, 1 reference page, 4 appendix pages, Accepted into International Conference on Automated Planning and Scheduling (ICAPS) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.02918 2020-11-06 cs.AI 61%

Domain-independent generation and classification of behavior traces

Daniel Borrajo, Manuela Veloso

专题命中 规划推理 :planning(abstract,comments);分类 cs.AI

Comments A version of this paper appears in the Pre-prints of the Workshop in Planning for Financial Services (FinPlan) at ICAPS'20. arXiv admin note: text overlap with arXiv:2011.01826

详情

展开后加载摘要…

URL PDF HTML 收藏
0909.4441 2012-04-18 cs.AI cs.GT cs.MA 61%

Dealing with incomplete agents' preferences and an uncertain agenda in group decision making via sequential majority voting

Maria Pini, Francesca Rossi, Brent Venable, Toby Walsh

专题命中 规划推理 :reasoning(abstract,comments);分类 cs.AI

Comments Principles of Knowledge Representation and Reasoning: Proceedings of the Eleventh International Conference, KR 2008, Sydney, Australia, September 16-19, 2008

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19041 2025-09-24 cs.HC 60%

Position: Human-Robot Interaction in Embodied Intelligence Demands a Shift From Static Privacy Controls to Dynamic Learning

Shuning Zhang, Hong Jia, Simin Li, Ting Dang, Yongquan `Owen' Hu, Xin Yi, Hewu Li

专题命中 规划推理 :reasoning(abstract,comments);planning(comments)

Comments To be published in NeurIPS 2025 Workshop on Bridging Language, Agent, and World Models for Reasoning and Planning

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08804 2025-07-15 cs.HC cs.CY 60%

Cognitive Dissonance Artificial Intelligence (CD-AI): The Mind at War with Itself. Harnessing Discomfort to Sharpen Critical Thinking

Delia Deliu

专题命中 规划推理 :reasoning(abstract,comments)

Comments Presented at the 2025 ACM Workshop on Human-AI Interaction for Augmented Reasoning, Report Number: CHI25-WS-AUGMENTED-REASONING

Journal ref Proceedings of the 2025 ACM CHI Workshop on Human-AI Interaction for Augmented Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16883 2025-04-24 cs.HC 60%

Enhancing Critical Thinking with AI: A Tailored Warning System for RAG Models

Xuyang Zhu, Sejoon Chang, Andrew Kuik

专题命中 规划推理 :reasoning(abstract,comments)

Comments Presented at the 2025 ACM Workshop on Human-AI Interaction for Augmented Reasoning

Journal ref Proceedings of the 2025 ACM CHI Workshop on Human-AI Interaction for Augmented Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14996 2025-04-22 cs.HC 60%

Distributed Cognition for AI-supported Remote Operations: Challenges and Research Directions

Rune Møberg Jacobsen, Joel Wester, Helena Bøjer Djernæs, Niels van Berkel

专题命中 规划推理 :reasoning(abstract,comments)

Comments Presented at the 2025 ACM Workshop on Human-AI Interaction for Augmented Reasoning, Report Number: CHI25-WS-AUGMENTED-REASONING

Journal ref Proceedings of the 2025 ACM CHI Workshop on Human-AI Interaction for Augmented Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13684 2025-04-21 cs.HC 60%

Intelligent Interaction Strategies for Context-Aware Cognitive Augmentation

Xiangrong, Zhu, Yuan Xu, Tianjian Liu, Jingwei Sun, Yu Zhang, Xin Tong

专题命中 规划推理 :reasoning(abstract,comments)

Comments Presented at the 2025 ACM Workshop on Human-AI Interaction for Augmented Reasoning, Report Number: CHI25-WS-AUGMENTED-REASONING

Journal ref Proceedings of the 2025 ACM CHI Workshop on Human-AI Interaction for Augmented Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26991 2026-08-28 cs.AI 新提交 57%

ASIL: Replacing Screenshot-and-Click with Structured State and Semantic Actions

ASIL:用结构化状态与语义动作替代截图点击

Rui Xie, Lu Chen

机构 * Shanghai Jiao Tong University(上海交通大学) BIGAI(北京智源人工智能研究院)

专题命中 规划推理 :planning(abstract);分类 cs.AI

AI总结 本文提出ASIL智能体-软件交互层,以结构化状态与语义动作替代低效的截图点击界面,在多应用基准任务中表现优于截图点击,还可提升Qwen系列模型的性能。

Comments Accepted to Findings of EMNLP 2026. 19 pages, 7 figures: 9-page main paper followed by limitations, ethics, acknowledgments, references, and appendices A-E. Project page: https://sharryxr.github.io/ASIL/

Journal ref Findings of the Association for Computational Linguistics: EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26990 2026-08-28 cs.AI cs.MA 新提交 57%

DSA: Evidence-Aware LLM-Agent Orchestration for Multi-Market Stock Research

DSA:面向多市场股票研究的证据感知大语言模型智能体编排框架

Linsen Zhu, Yi Shi

专题命中 规划推理 :reasoning(abstract);分类 cs.AI

AI总结 该研究提出DSA框架,基于LLM智能体实现多市场股票研究的证据感知编排,通过工作流组织、配置文件差异化处理及测试验证,确保系统实现一致性。

Comments 6 pages, 2 figures, 3 tables. Code available at https://github.com/ZhuLinsen/daily_stock_analysis

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26332 2026-08-28 cs.LG 新提交 57%

Beyond Capability Benchmarks: Learning Operational Fingerprints of LLM Cloud Services from Production Incident Metadata

超越能力基准:从生产事件元数据学习LLM云服务的操作指纹

Meiwei Zhang, Eduardo Miranda, Bruce Baynes, Suvigya Jain, Wanlong Chen, Tao He, Sergey Borodavkin

机构 * Google Cloud Platform, Google LLC(谷歌云平台,谷歌有限责任公司)

专题命中 规划推理 :planning(abstract);分类 cs.LG

AI总结 该研究提出OpEmbed框架,利用生产支持案例元数据学习LLM云服务的操作指纹,在Google Cloud的大规模生产案例评估中表现优异,可用于模型上线、支持评估等场景。

Journal ref ISSRE 2026 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26226 2026-08-28 cs.AI 新提交 57%

LLM Agents for Time-Series: A Survey

面向时间序列的大语言模型(LLM)智能体:一项综述

Yilong Chen, Xiao Qin, Chenghao Liu, Liang Wu, Noelle I. Samia, Kaize Ding

机构 * Northwestern University(西北大学) Datadog AI Research(Datadog人工智能研究院) Nokia(诺基亚)

专题命中 规划推理 :reasoning(abstract);分类 cs.AI

AI总结 本综述针对LLM智能体在时间序列问题上设计差异大的问题,采用问题驱动分类法梳理现有系统,分析任务对架构等的影响,对比模型性能,为相关设计提供指南并指出未来缺口。

Comments Accepted to Findings of EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25457 2026-08-28 cs.CR cs.AI cs.MA 版本更新 57%

MACGen: Toward Functionally Correct and Secure Code Generation via Multi-Agent Collaboration

MACGen:通过多智能体协作实现功能正确且安全的代码生成

Miseon Yu, Jaehoon Choi, Younghan Lee, Yunheung Paek

机构 * Seoul National University(首尔大学) Sungshin Women’s University(诚信女子大学)

专题命中 规划推理 :planning(abstract);分类 cs.AI

AI总结 MACGen是整合规划等多环节的多智能体代码生成框架,在CWEval、BaxBench基准上,较直接提示显著提升安全代码生成的功能与安全性指标。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22631 2026-08-28 cs.LG 版本更新 57%

Learning Generalizable Behaviors for Terminal Agents

学习终端智能体的可泛化行为

Yihang Yao, Bo Pang, Xuan Phi Nguyen, Ding Zhao, Shafiq Joty, Semih Yavuz

机构 * Salesforce AI Research(Salesforce人工智能研究院) Carnegie Mellon University(卡内基梅隆大学)

专题命中 规划推理 :verifier(abstract);分类 cs.LG

AI总结 本研究针对终端智能体泛化问题,提出智能体组合泛化假设,开发训练方案River,其用不足30%的TMax环境使多规模模型在两个基准上RL增益平均提升超100%,且性能优于开源8B模型、可跨多维度泛化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22444 2026-08-28 cs.CL 版本更新 57%

Aligned Alone, Misaligned Together: Forecasting Adversarial Capture in LLM Agent Populations

单独对齐,共同错位:预测大语言模型智能体群体中的对抗捕获

Isotta Magistrali, Chen Shani

专题命中 规划推理 :reasoning(abstract);分类 cs.CL

AI总结 该研究针对LLM智能体群体,提出从群体无攻击时的良性运行状态校准响应函数,可提前预测坚定少数群体引发的对抗捕获,且捕获为临时状态,孤立对齐不等同于群体对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29353 2026-08-28 cs.AI 版本更新 57%

Nomad: Autonomous Exploration and Discovery

Nomad:自主探索与发现

Bokang Jia, Samta Kamboj, Satheesh Katipomu, Seung Hun Han, Neha Sengupta, Andrew Jackson

机构 * Inception, G42(G42 的 Inception) G42 MBZUAI(穆罕默德·本·扎耶德人工智能大学)

专题命中 规划推理 :verifier(abstract);分类 cs.AI

AI总结 Nomad通过探索优先架构,构建领域探索地图,系统遍历以平衡广度与深度,生成并验证假设,最终生成可信报告。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24889 2026-08-27 cs.AI 新提交 57%

Reliable LLM-Powered Decision Engines for Large-Scale Supply Chain Operations: Architecture, Safety, and Performance Guarantees

面向大规模供应链运营的可靠大语言模型驱动决策引擎:架构、安全性与性能保证

Nirmal Kumar Jingar

专题命中 规划推理 :reasoning(abstract);分类 cs.AI

AI总结 针对大规模供应链的不确定性等挑战,提出结合LLM与优化等的LLM-DE架构,实现更智能、安全、可扩展的供应链决策,为下一代智能供应链提供基础。

Journal ref IC_ASET 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22899 2026-08-27 cs.AI 版本更新 57%

CDEG: Learning Decision-Critical Evidence for Long-Horizon Diagnostic Agents

CDEG:为长程诊断智能体学习决策关键证据

Xiwei Dai, Zijie Meng, Zhiting Fan, Yixuan Tang, Guanyu Jiang, Ziru Niu, Zuozhu Liu

专题命中 规划推理 :reasoning(abstract);分类 cs.AI

AI总结 本研究提出基于图的框架CDEG,通过对比同一病例的成功与失败轨迹、受控反事实干预等学习决策关键证据,在多基准测试中使长程诊断智能体准确率最高提升11.5%。

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05925 2026-08-27 cs.AI 版本更新 57%

Towards World Models in Biomedical Research

迈向生物医学研究的世界模型

Guangyu Wang, Jingkun Yue, Siqi Zhang, Yu Liu, Xiaoyu Wang, Mingyuan Meng, Changwei Ji, Zongbo Han, Yulin Wang, Yang Yue, Frank Fu, Ting Chen, Song Wu, Ziwei Liu, Jiangning Song, Ming Li, Gao Huang, Xiaohong Liu, Athanasios Vasilakos, Xingcai Zhang, Ping Zhang, Yong Li

机构 * State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing, China(网络与交换技术国家重点实验室,北京邮电大学,北京,中国) Department of Engineering Science, University of Oxford, Oxford, United Kingdom(英国牛津大学工程科学系,牛津,英国) Institute of Medical Artificial Intelligence, South China Hospital, Medical School, Shenzhen University, Shenzhen, Guangdong, China(医学人工智能研究所,南方医院,医学学院,深圳大学,深圳,广东,中国) Zhongguancun Academy & Zhongguancun Institute of Artificial Intelligence, Beijing, China(中关村学院及中关村人工智能研究院,北京,中国) Beijing National Research Center for Information Science and Technology (BNRist), Tsinghua University, 100084, Beijing, China(北京信息科学与技术国家研究中心(BNRist),清华大学,100084,北京,中国) Department of Chemical and Nano Engineering, University of California, San Diego, La Jolla, CA, USA(美国加州大学圣地亚哥分校化学与纳米工程系,La Jolla,CA,美国) Nanyang Technological University, Singapore(新加坡南洋理工大学) Monash Biomedicine Discovery Institute and Department of Biochemistry and Molecular Biology, Monash University, Melbourne, Victoria, Australia(莫纳什大学生物医学发现研究所和生物化学与分子生物学系,墨尔本,维多利亚,澳大利亚) David R. Cheriton School of Computer Science, University of Waterloo, Waterloo, Ontario, Canada(加拿大滑铁卢大学戴维·R·切里顿计算机科学学校,滑铁卢,安大略,加拿大) Department of ICT and Center for AI Research, University of Agder (UiA), Jon Lilletuns vei 9, Grimstad, Norway(挪威阿格德大学(UiA)信息与通信技术系及人工智能研究中心,Jon Lilletuns vei 9,Grimstad,挪威) Department of Electronic Engineering, Tsinghua University, Beijing, China(清华大学电子工程系,北京,中国)

专题命中 规划推理 :planning(abstract);分类 cs.AI

AI总结 提出生物医学世界模型作为AI驱动发现的新范式,通过学习分子、细胞、组织和临床状态的潜在表征及干预条件动态,实现未来轨迹模拟,并探讨其在虚拟细胞、类器官、虚拟患者和手术模拟等应用中的潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14857 2026-08-27 cs.AI 版本更新 57%

World Models for Policy Refinement in StarCraft II

为《星际2》政策细化设计的世界模型

Yixin Zhang, Ziyi Wang, Yiming Rong, Haoxi Wang, Jinling Jiang, Shuang Xu, Haoran Wu, Shiyu Zhou, Bo Xu

机构 * The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学)

专题命中 规划推理 :reasoning(abstract);分类 cs.AI

AI总结 StarWM为《星际2》政策细化设计的世界模型,通过预测未来观察和结构化文本表示提升策略决策性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04879 2026-08-27 cs.CL 版本更新 57%

Mind2Report: Expert-Level Commercial Report Synthesis via Cognitive Deep Research Agent

Mind2Report: 一种用于专家级商业报告综合的认知深度研究代理

Mingyue Cheng, Daoyu Wang, Qi Liu, Shuo Yu, Xiaoyu Tao, Yuqian Wang, Chengzhong Chu, Yu Duan, Mingkang Long, Enhong Chen

机构 * State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China(认知智能国家重点实验室,中国科学技术大学) Artificial Intelligence Engineering Institute, iFLYTEK Co., Ltd(人工智能工程院,iFLYTEK公司)

专题命中 规划推理 :planning(abstract);分类 cs.CL

AI总结 Mind2Report是一种通过动态内存增强大语言模型,实现专家级商业报告综合的认知深度研究代理,优于现有基线模型。

Comments Accepted at the 35th ACM International Conference on Information and Knowledge Management (CIKM 2026), Rome, Italy

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23791 2026-08-26 eess.AS cs.AI 新提交 57%

EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis

EmoTra-TTS:用于语音合成的流畅句内情感转换

Tianchi Liu, Zeyang Song, Tianrui Wang, Zhipeng Li, Chenglin Xu, Yiwen Guo

机构 * LIGHTSPEED National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学)

专题命中 规划推理 :planning(abstract);分类 cs.AI

AI总结 EmoTra-TTS针对现有情感TTS系统与情感时间特性不匹配的问题,采用多遍流混合管道、双阶段VAD条件及方向-幅度解耦注入,实现流畅句内情感转换,且性能优于SOTA基线与商业系统。

Comments Accepted to EMNLP 2026 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24221 2026-08-26 cs.SE cs.CL cs.PL 新提交 57%

DeepRepoQA: Code Repository Question Answering with Deep Agent Exploration

DeepRepoQA:基于深度智能体探索的代码仓库问答系统

Weihan Peng, Yuling Shi, Yingwei Ma, Longfei Yun, Beijun Shen, Xiaodong Gu

机构 * Shanghai Jiao Tong University(上海交通大学) The Hong Kong University of Science and Technology(香港科技大学) University of California San Diego(加州大学圣迭戈分校)

专题命中 规划推理 :reasoning(abstract);分类 cs.CL

AI总结 针对现有代码仓库问答方法缺乏深度推理能力的问题,提出基于LLM智能体与蒙特卡洛树搜索的DeepRepoQA框架,在SWE-QA基准上实现了性能显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏