arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

2026-01-30 至 2026-01-30 共收录 112 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. 工具调用 7 篇

2601.21947 2026-01-30 cs.AI 79%

ToolWeaver: Weaving Collaborative Semantics for Scalable Tool Use in Large Language Models

ToolWeaver: 为大语言模型中的可扩展工具使用编织协作语义

Bowen Fang, Wen Ye, Yunyue Su, Jinghao Zhang, Qiang Liu, Yesheng Liu, Xin Sun, Shu Wu, Jiabing Yang, Baole Wei, Liang Wang

机构 * New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences (CASIA)(模式识别新实验室(NLPR),自动化研究所,中国科学院(CASIA)) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Zhongguancun Academy(中关村学院) Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究所)

专题命中 工具调用 :tool use(title);tool-use(abstract);分类 cs.AI

AI总结 ToolWeaver通过编码工具为层次序列,解决大语言模型中工具使用中的语义和可扩展性问题,提升工具协作学习的效率与效果。

Comments 10pages, 12 figures, Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00564 2026-01-30 cs.LG cs.AI 73%

Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining

通过联合优化的世界-动作模型扩展离线模型基于的强化学习

Jie Cheng, Ruixi Qiao, Yingwei Ma, Binhua Li, Gang Xiong, Qinghai Miao, Yongbin Li, Yisheng Lv

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Alibaba Group(阿里巴巴集团)

专题命中 工具调用 :agent(abstract);planning(abstract);分类 cs.AI、cs.LG

AI总结 JOWA通过联合优化的世界-动作模型扩展离线RL,实现高效泛化和高性能

Comments Accepted by ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22055 2026-01-30 cs.CL 70%

$G^2$-Reader: Dual Evolving Graphs for Multimodal Document QA

$G^2$-Reader: 双重演化图式用于多模态文档问答

Yaxin Du, Junru Song, Yifan Zhou, Cheng Wang, Jiahao Gu, Zimeng Chen, Menglan Chen, Wen Yao, Yang Yang, Ying Wen, Siheng Chen

机构 * Shanghai Jiao Tong University(上海交通大学) Intelligent Game and Decision Laboratory(智能游戏与决策实验室)

专题命中 工具调用 :planning(abstract);agentic(abstract);分类 cs.CL

AI总结 $G^2$-Reader通过双重图式解决多模态文档问答中的结构破坏和检索失效问题,实现66.21%的高准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21208 2026-01-30 cs.AI cs.IR 70%

When should I search more: Adaptive Complex Query Optimization with Reinforcement Learning

我应该何时搜索更多:基于强化学习的自适应复杂查询优化

Wei Wen, Sihang Deng, Tianjun Wei, Keyu Chen, Ruizhi Qiao, Xing Sun

机构 * Tencent Youtu Lab(腾讯云图实验室) The University of Hong Kong(香港大学) Nanyang Technological University(南洋理工大学)

专题命中 工具调用 :agent(abstract);agentic(abstract);分类 cs.AI

AI总结 本文提出自适应复杂查询优化框架ACQO,通过强化学习解决复杂查询优化中的子查询数量确定、文档排序与合并问题,实现更稳定高效的查询处理。

Comments 16 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.13860 2026-01-30 cs.AI math.PR q-bio.NC 57%

Active Inference Tree Search in Large POMDPs

大规模POMDPs中的主动推断树搜索

Domenico Maisto, Francesco Gregoretti, Karl Friston, Giovanni Pezzulo

专题命中 工具调用 :planning(abstract);分类 cs.AI

AI总结 本文提出主动推断树搜索方法,结合神经科学与人工智能的规划理论,实现大规模POMDP问题的高效解决。

Comments 47 pages, 9 figures, 1 Appendix of two sections with pseudocodes and one encoding example, submitted preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11311 2026-01-30 cs.AI cs.CY 57%

Prompts to Proxies: Emulating Human Preferences via a Compact LLM Ensemble

通过紧凑的LLM集合模拟人类偏好:提示到代理

Bingchen Wang, Zi-Yu Khoo, Jingtan Wang

机构 * Independent Researcher, China(中国独立研究者) School of Computing, National University of Singapore(新加坡国立大学计算机学院)

专题命中 工具调用 :agent(abstract);分类 cs.AI

AI总结 通过紧凑的LLM集合模拟人类偏好,P2P方法在无需微调和敏感数据的情况下,有效重建目标人群偏好并实现高质量预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04967 2026-01-30 quant-ph 50%

Hybrid Action Reinforcement Learning for Quantum Architecture Search

混合动作强化学习用于量子架构搜索

Jiayang Niu, Yan Wang, Jie Li, Ke Deng, Azadeh Alavi, Muhammad Usman, Yongli Ren

专题命中 工具调用 :agent(abstract)

AI总结 HyRLQAS通过混合动作强化学习框架,联合学习量子电路门放置与参数初始化,实现高效优化分子基态能量,展现优于现有方法的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 规划决策 41 篇

2601.21916 2026-01-30 cs.AI cs.CL cs.IR 89%

JADE: Bridging the Strategic-Operational Gap in Dynamic Agentic RAG

JADE:动态代理RAG中战略-操作鸿沟的弥合

Yiqun Chen, Erhan Zhang, Tianyi Hu, Shijie Wang, Zixuan Yang, Meizhi Zhong, Xiaochi Wei, Yan Gao, Yi Wu, Yao Hu, Jiaxin Mao

机构 * Renmin University of China(中国人民大学) Xiaohongshu Inc.(小红书公司) Institute of Automation,Chinese Academy of Sciences(中国科学院自动化研究所) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 规划决策 :agentic(title,abstract);agent(abstract);planning(abstract);workflow(abstract)

AI总结 JADE通过联合优化规划与执行,解决动态代理RAG中战略与操作不匹配的问题,提升多轮工作流的性能与灵活性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21936 2026-01-30 cs.AI 88%

AgenticSimLaw: A Juvenile Courtroom Multi-Agent Debate Simulation for Explainable High-Stakes Tabular Decision Making

AgenticSimLaw: 一个面向可解释高风险表格决策的青少年法庭多智能体辩论模拟

Jon Chun, Kathrine Elkins, Yong Suk Lee

专题命中 规划决策 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI

AI总结 AgenticSimLaw通过结构化多智能体辩论框架提升高风险表格决策的可解释性和稳定性,适用于需要透明和人类监督的决策任务。

Comments 18 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21113 2026-01-30 cs.AI cs.MA 88%

Planner-Auditor Twin: Agentic Discharge Planning with FHIR-Based LLM Planning, Guideline Recall, Optional Caching and Self-Improvement

计划-审核双胞胎:基于FHIR的LLM计划、指南回忆、可选缓存和自我改进的代理出院计划

Kaiyuan Wu, Aditya Nagori, Rishikesan Kamaleswaran

专题命中 规划决策 :planning(title,abstract);agentic(title,abstract);分类 cs.AI

AI总结 基于FHIR的LLM计划与审核框架,通过自我改进和缓存优化,提升临床出院计划的安全性和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10890 2026-01-30 cs.CL 87%

LLM$\times$MapReduce-V3: Enabling Interactive In-Depth Survey Generation through a MCP-Driven Hierarchically Modular Agent System

LLM×MapReduce-V3:通过MCP驱动的分层模块化代理系统实现交互式深入调研生成

Yu Chao, Siyu Lin, xiaorong wang, Zhu Zhang, Zihan Zhou, Haoyu Wang, Shuo Wang, Jie Zhou, Zhiyuan Liu, Maosong Sun

机构 * Dept. of Comp. Sci. & Tech., Institute for AI, BNRist Center, Tsinghua University(计算机科学与技术系,人工智能研究院,BNRist中心,清华大学) Peking University(北京大学) Modelbest Inc.(Modelbest公司) Nanyang Technological University(南洋理工大学)

专题命中 规划决策 :agent(title,abstract);planning(abstract);workflow(abstract);multi-agent(abstract)

AI总结 LLM×MapReduce-V3通过MCP驱动的分层模块化代理系统实现交互式深入调研生成,提升调研内容深度与长度。

Comments Accepted by EMNLP2025 System Demonstration

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21352 2026-01-30 cs.AI 86%

BEAP-Agent: Backtrackable Execution and Adaptive Planning for GUI Agents

BEAP-Agent: 可回溯执行与自适应规划用于GUI代理

Ziyu Lu, Tengjin Weng, Yiying Yang, Yuhang Zhao, Xinxin Huang, Wenhao Jiang

机构 * Guangdong Laboratory of Artificial Intelligence(广东人工智能实验室) Shenzhen University(深圳大学) Guangdong University of Technology(广东工业大学)

专题命中 规划决策 :agent(title,abstract);planning(title);分类 cs.AI

AI总结 BEAP-Agent通过可回溯执行和自适应规划机制,提升GUI代理在复杂任务探索中的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19510 2026-01-30 cs.RO cs.CL 83%

ALRM: Agentic LLM for Robotic Manipulation

ALRM:用于机器人操作的代理LLM

Vitor Gaboardi dos Santos, Ibrahim Khadraoui, Ibrahim Farhat, Hamza Yous, Samy Teffahi, Hakim Hacid

专题命中 规划决策 :agentic(title,abstract);planning(abstract);分类 cs.CL

AI总结 ALRM提出了一种基于LLM的代理框架,通过ReAct推理循环实现机器人操作的模块化执行,支持代码生成和工具规划模式,实验表明其在多步骤推理和语言多样性任务中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21504 2026-01-30 cs.RO 82%

Don't double it: Efficient Agent Prediction in Occlusions

不要双倍:遮挡中的高效代理预测

Anna Rothenhäusler, Markus Mazzola, Andreas Look, Raghu Rajan, Joschka Bödecker

专题命中 规划决策 :agent(title,abstract);planning(abstract)

AI总结 本文提出MatchInformer方法,通过整合Hungarian匹配和分离航向与运动,提升遮挡场景下的代理预测准确性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21967 2026-01-30 cs.AI cs.SE 81%

The Energy Impact of Domain Model Design in Classical Planning

经典规划中领域模型设计的能量影响

Ilche Georgievski, Serhat Tekin, Marco Aiello

机构 * University of Stuttgart(斯图加特大学)

专题命中 规划决策 :planning(title,abstract);分类 cs.AI、cs.SE

AI总结 本文研究了领域模型设计对经典规划器能耗的影响,通过配置框架分析不同领域特征对能量消耗和运行时间的影响。

Comments 2026 IEEE/ACM 5th International Conference on AI Engineering - Software Engineering for AI (CAIN '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07202 2026-01-30 cs.AI cs.LG 81%

Monte Carlo Tree Diffusion for System 2 Planning

蒙特卡洛树扩散用于系统2规划

Jaesik Yoon, Hyeonseo Cho, Doojin Baek, Yoshua Bengio, Sungjin Ahn

机构 * New York University(纽约大学)

专题命中 规划决策 :planning(title,abstract);分类 cs.AI、cs.LG

AI总结 蒙特卡洛树扩散结合扩散模型和MCTS的优势,通过树结构的去噪过程提升规划性能,实验证明其在长时间任务中优于现有方法。

Comments 23 pages, 7 figures, ICML 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11479 2026-01-30 cs.AI 79%

Health Facility Location in Ethiopia: Leveraging LLMs to Integrate Expert Knowledge into Algorithmic Planning

埃塞俄比亚卫生设施选址:利用大语言模型整合专家知识到算法规划中

Yohai Trabelsi, Guojun Xiong, Fentabil Getnet, Stéphane Verguet, Milind Tambe

机构 * John A. Paulson School of Engineering and Applied Sciences, Harvard University(工程与应用科学学院,哈佛大学) National Data Management and Analytics Center for Health, Ethiopian Public Health Institute(健康数据管理与分析中心,埃塞俄比亚公共卫生研究所) Department of Global Health and Population, Harvard T.H. Chan School of Public Health(全球卫生与人口学院,哈佛T.H. Chan公共卫生学院)

专题命中 规划决策 :planning(title,abstract);分类 cs.AI

AI总结 本文提出了一种结合大语言模型和优化技术的混合框架,用于埃塞俄比亚卫生设施选址,以整合专家知识并提升规划的公平性和数据驱动性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21598 2026-01-30 cs.AI 79%

Beyond Imitation: Reinforcement Learning for Active Latent Planning

超越模仿:用于主动潜在规划的强化学习

Zhi Zheng, Wee Sun Lee

机构 * School of Computing, National University of Singapore, Singapore(计算学院,新加坡国立大学)

专题命中 规划决策 :planning(title,abstract);分类 cs.AI

AI总结 本文提出ATP-Latent方法,通过主动规划和强化学习优化潜在空间,提升链式推理的准确性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21212 2026-01-30 cs.AI cs.CY 79%

Intelli-Planner: Towards Customized Urban Planning via Large Language Model Empowered Reinforcement Learning

Intelli-Planner: 通过大型语言模型赋能的强化学习实现定制化城市规划

Xixian Yong, Peilin Sun, Zihe Wang, Xiao Zhou

机构 * Gaoling School of Artificial Intelligence\ University of China Beijing China Gaoling School of Artificial Intelligence\ University of China

专题命中 规划决策 :planning(title,abstract);分类 cs.AI

AI总结 Intelli-Planner通过结合深度强化学习与大型语言模型,实现定制化城市规划,提升规划方案的参与度和满意度。

Comments The Web Conference 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21876 2026-01-30 cs.RO 78%

LLM-Driven Scenario-Aware Planning for Autonomous Driving

基于大语言模型的场景感知规划用于自动驾驶

He Li, Zhaowei Chen, Rui Gao, Guoliang Li, Qi Hao, Shuai Wang, Chengzhong Xu

机构 * University of Macau(澳门大学) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所) Southern University of Science and Technology(南方科技大学)

专题命中 规划决策 :planning(title,abstract)

AI总结 本文提出LAP,一种基于大语言模型的自适应规划方法,通过场景理解与联合优化实现自动驾驶中的高速与精确驾驶平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21679 2026-01-30 eess.SY cs.SY 78%

BAP-SRL: Bayesian Adaptive Priority Safe Reinforcement Learning for Vehicle Motion Planning at Mixed Traffic Intersections

BAP-SRL:基于贝叶斯自适应优先安全强化学习的车辆混合交通交叉口运动规划

Yuansheng Lian, Ke Zhang, Yaming Guo, Shen Li, Meng Li

专题命中 规划决策 :planning(title,abstract)

AI总结 BAP-SRL通过贝叶斯自适应优先安全强化学习方法,提升自动驾驶车辆在混合交通交叉口的安全性和决策效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18368 2026-01-30 cs.AI 77%

Wireless Power Transfer and Intent-Driven Network Optimization in AAVs-assisted IoT for 6G Sustainable Connectivity

面向6G可持续连接的AAV辅助物联网中的无线功率传输与意图驱动网络优化

Xiaoming He, Gaofeng Wang, Huajun Cui, Rui Yuan, Haitao Zhao

机构 * College of Internet of Things, Nanjing University of Posts and Telecommunications(物联网学院,南京邮电大学) Digital Intelligence Research Institute, PowerChina, Beijing Engineering Corporation Limited(数字智能研究院,中国电力工程顾问集团北京公司) College of Telecommunications and Information Engineering, Nanjing University of Posts and Telecommunications(电信与信息工程学院,南京邮电大学)

专题命中 规划决策 :agent(abstract);planning(abstract);multi-agent(abstract);分类 cs.AI

AI总结 本文提出意图驱动的自主网络优化框架,通过超维变换器和双动作多智能体策略优化,提升AAV辅助物联网中高维动作序列处理和低延迟执行能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21037 2026-01-30 cs.LG cs.AI cs.CL cs.CV 75%

Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning

基于框架的思考:视觉上下文和测试时扩展如何增强视频推理

Chengzu Li, Zanyi Wang, Jiaang Li, Yi Xu, Han Zhou, Huanyu Zhang, Ruichuan An, Dengyang Jiang, Zhaochong An, Ivan Vulić, Serge Belongie, Anna Korhonen

机构 * University of Cambridge(剑桥大学) Pioneer Center for AI, University of Copenhagen(人工智能先锋中心,哥本哈根大学) Hong Kong University of Science(香港科学大学) University of California San Diego(加州大学圣地亚哥分校) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Peking University(北京大学)

专题命中 规划决策 :agent(abstract);planning(abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 本文提出通过视频生成模型进行视觉推理,利用视觉上下文和测试时扩展提升视频推理能力,展示了模型在空间和时间复杂任务中的鲁棒零样本泛化能力。

Comments 8 pages, 3 figures, 3 tables (26 pages, 13 figures, 6 tables including references and appendices)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21993 2026-01-30 cs.AI cs.SE 73%

Liquid Interfaces: A Dynamic Ontology for the Interoperability of Autonomous Systems

液态接口:自主系统互操作性的动态本体

Dhiogo de Sá, Carlos Schmiedel, Carlos Pereira Lopes

专题命中 规划决策 :agent(abstract);autonomous agent(abstract);分类 cs.AI、cs.SE

AI总结 本文提出液态接口,通过动态语义协商实现自主系统间的适应性协调,提供了一种基于临时关系事件的动态本体方法。

Comments 28 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21477 2026-01-30 cs.MA cs.AI cs.LG math.OC 73%

Mean-Field Control on Sparse Graphs: From Local Limits to GNNs via Neighborhood Distributions

均场控制在稀疏图上的应用:从局部极限到GNNs via 邻居分布

Tobias Schmidt, Kai Cui

机构 * Department of Mathematics, TU Darmstadt, Darmstadt, Germany(图腾达姆施塔特大学数学系) Department of Electrical Engineering, TU Darmstadt, Darmstadt, Germany(图腾达姆施塔特大学电气工程系)

专题命中 规划决策 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种适用于稀疏图的均场控制框架,通过邻居分布理论解决了多智能体系统中的维度灾难问题,并证明了GNNs在该场景下的有效性。

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20090 2026-01-30 cs.AI 70%

Should I Have Expressed a Different Intent? Counterfactual Generation for LLM-Based Autonomous Control

我本应表达不同的意图?基于LLM的自主控制的反事实生成

Amirmohammad Farzaneh, Salvatore D'Oro, Osvaldo Simeone

机构 * Department of Engineering, King's College London, London, UK(伦敦大学金史密斯学院工程系) Intelligent Networked Systems Institute (INSI), Northeastern University, Boston, MA, USA(东北大学智能网络系统研究所) Intelligent Networked Systems Institute (INSI), Northeastern University, London, UK(伦敦大学东北大学智能网络系统研究所)

专题命中 规划决策 :agent(abstract);agentic(abstract);分类 cs.AI

AI总结 本文提出一种基于结构因果模型的反事实生成方法,通过概率性反事实生成和离线校准,为LLM驱动的自主控制提供可靠反事实推理,展示了在无线网络控制中的显著优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13421 2026-01-30 cs.AI cs.ET 70%

Virtuous Machines: Towards Artificial General Science

美德机器:迈向人工智能通用科学

Gabrielle Wehr, Reuben Rideaux, Amaya J. Fox, David R. Lightfoot, Jason Tangen, Jason B. Mattingley, Shane E. Ehrhardt

机构 * Explore Science School of Psychology(心理学学院) Queensland Brain Institute(昆士兰脑研究所) Canadian Institute of Advanced Research (CIFAR)(加拿大高级研究 institute)

专题命中 规划决策 :workflow(abstract);agentic(abstract);分类 cs.AI

AI总结 本研究提出了一种无需领域知识的AI科学家系统,能自主完成科学流程,通过实验展示了其在心理学研究中的能力,推动了人工智能在科学发现中的应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21991 2026-01-30 cs.LG cs.AI 62%

Geometry of Drifting MDPs with Path-Integral Stability Certificates

漂移MDP的几何学与路径积分稳定性证书

Zuyuan Zhang, Mahdi Imani, Tian Lan

机构 * The George Washington University(乔治·华盛顿大学) Northeastern University(东北大学)

专题命中 规划决策 :planning(abstract);分类 cs.AI、cs.LG

AI总结 本文提出通过几何方法建模漂移MDP,引入HT-RL和HT-MCTS算法,利用路径积分稳定性证书提升强化学习在非平稳环境中的跟踪与动态遗憾性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12481 2026-01-30 cs.LG cs.AI 62%

SAC-GLAM: Improving Online RL for LLM agents with Soft Actor-Critic and Hindsight Relabeling

SAC-GLAM:通过软演员-评论员和回顾重标记改进LLM代理的在线强化学习

Loris Gaven, Clement Romac, Thomas Carta, Sylvain Lamprier, Olivier Sigaud, Pierre-Yves Oudeyer

机构 * Inria (Flowers)(Inria(Flowers)) University of Bordeaux(波尔多大学) Hugging Face Univ Angers(昂热大学) LERIA SFR MATHSTIC(MATHSTIC联合研究机构) Sorbonne Université(索邦大学) ISIR

专题命中 规划决策 :agent(abstract);分类 cs.AI、cs.LG

AI总结 SAC-GLAM通过结合软演员-评论员算法和回顾重标记,改进LLM代理的在线强化学习,提升其在复杂环境中的策略学习能力。

Comments This work has been presented at the IMOL workshop at NeurIPS 2025 (https://neurips.cc/virtual/2024/101058)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21618 2026-01-30 cs.AI cs.CL 62%

Semantic Content Determines Algorithmic Performance

语义内容决定算法性能

Martiño Ríos-García, Nawaf Alampara, Kevin Maik Jablonka

机构 * Laboratory of Organic Macromolecular Chemistry (IOMC), Friedrich Schiller University Jena, Humboldtstrasse 10, 07743 Jena, Germany HIPOLE Jena (Helmholtz Institute for Polymers in Energy Applications Jena), Lessingstrasse 12-14, 07743 Jena, Germany Center for Energy Environmental Chemistry Jena (CEEC Jena), Friedrich Schiller University Jena, Philosophenweg 7a, 07743 Jena, Germany Jena Center for Soft Matter (JCSM), Friedrich Schiller University Jena, Philosophenweg 7, 07743 Jena, Germany

专题命中 规划决策 :agentic(abstract);分类 cs.AI、cs.CL

AI总结 研究揭示大语言模型的性能受输入语义影响,通过WhatCounts实验展示不同语义内容导致计数准确率显著变化,表明模型对输入意义存在隐性依赖。

详情

展开后加载摘要…

URL PDF HTML 收藏