arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 15866 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. Agent评测 15866 篇

2402.06929 2025-11-25 cs.AI 57%

Making a prototype of Seoul historical sites chatbot using Langchain

利用Langchain制作首尔历史遗址聊天机器人原型

Jae Young Suh, Minsoo Kwak, Soo Yong Kim, Hyoungseo Cho

机构 * Hanyang(翰阳大学) Konkuk(konkuk大学) Seoul National(首尔国立大学) Myongji University(明gies大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI

AI总结 本文提出利用Langchain开发首尔历史遗址聊天机器人原型,旨在通过提供准确信息提升游客对当地文化遗产的认知。

Comments 4 pages, 4 figures, draft

Journal ref Journal of Electrical Electronics Engineering, 3(1), 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15992 2025-11-21 cs.AI 57%

Detecting Sleeper Agents in Large Language Models via Semantic Drift Analysis

通过语义漂移分析检测大语言模型中的休眠代理

Shahin Zanbaghi, Ryan Rostampour, Farhan Abid, Salim Al Jarmakani

机构 * School of Computer Science, University of Windsor(计算机科学学院,温莎大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI

AI总结 通过语义漂移分析与canary基线比较,实现大语言模型中休眠代理的实时检测,准确率达92.5%。

Comments 7 pages, 3 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18591 2025-11-20 cs.LG stat.ML 57%

Bayesian Meta-Reinforcement Learning with Laplace Variational Recurrent Networks

Joery A. de Vries, Jinke He, Mathijs M. de Weerdt, Matthijs T. J. Spaan

专题命中 Agent评测 :agent(abstract);分类 cs.LG

Journal ref https://rlj.cs.umass.edu/2025/papers/Paper23.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14876 2025-11-20 cs.CR cs.CV cs.LG cs.RO 57%

Attacking Autonomous Driving Agents with Adversarial Machine Learning: A Holistic Evaluation with the CARLA Leaderboard

Henry Wong, Clement Fung, Weiran Lin, Karen Li, Stanley Chen, Lujo Bauer

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 Agent评测 :agent(abstract);分类 cs.LG

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14768 2025-11-20 cs.IR cs.AI 57%

Causally-Informed Reinforcement Learning for Adaptive Emotion-Aware Social Media Recommendation

Bhavika Jain, Robert Pitsko, Ananya Drishti, Mahfuza Farooque

机构 * Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14439 2025-11-20 cs.CL 57%

MedBench v4: A Robust and Scalable Benchmark for Evaluating Chinese Medical Language Models, Multimodal Models, and Intelligent Agents

Jinru Ding, Lu Lu, Chao Ding, Mouxiao Bian, Jiayuan Chen, Wenrao Pang, Ruiyao Chen, Xinwei Peng, Renjie Lu, Sijie Ren, Guanxu Zhu, Xiaoqin Wu, Zhiqiang Liu, Rongzhao Zhang, Luyi Jiang, Bing Han, Yunqiu Wang, Jie Xu

专题命中 Agent评测 :agentic(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14433 2025-11-19 cs.LO cs.RO cs.SE 57%

Safe-ROS: An Architecture for Autonomous Robots in Safety-Critical Domains

Diana C. Benjumea, Marie Farrell, Louise A. Dennis

机构 * Department of Computer Science The University of Manchester Manchester, UK(计算机科学系曼彻斯特大学曼彻斯特英国) University of Manchester Manchester, UK(曼彻斯特大学曼彻斯特英国)

专题命中 Agent评测 :agent(abstract);分类 cs.SE

Comments In Proceedings FMAS 2025, arXiv:2511.13245

Journal ref EPTCS 436, 2025, pp. 48-68

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02352 2025-11-19 cs.SE 57%

SWE-Sharp-Bench: A Reproducible Benchmark for C# Software Engineering Tasks

Sanket Mhatre, Yasharth Bajpai, Sumit Gulwani, Emerson Murphy-Hill, Gustavo Soares

专题命中 Agent评测 :agent(abstract);分类 cs.SE

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22995 2025-11-19 cs.LG cs.SY eess.SY 57%

A Reinforcement Learning Approach for Optimal Control in Microgrids

Davide Salaorni, Federico Bianchi, Francesco Trovò, Marcello Restelli

机构 * Politecnico di Milano(米兰理工大学) Ricerca sul Sistema Energetico(能源系统研究)

专题命中 Agent评测 :agent(abstract);分类 cs.LG

Comments 8 pages, accepted to International Joint Conference on Neural Networks 2025

Journal ref 2025 International Joint Conference on Neural Networks (IJCNN)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07965 2025-11-19 cs.AI cs.RO 57%

Autonomous Vehicle Controllers From End-to-End Differentiable Simulation

Asen Nachkov, Danda Pani Paudel, Luc Van Gool

专题命中 Agent评测 :agent(abstract);分类 cs.AI

Comments Polished and accepted at IROS 2025

Journal ref 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09562 2025-11-19 cs.CR cs.LG 57%

TooBadRL: Trigger Optimization to Boost Effectiveness of Backdoor Attacks on Deep Reinforcement Learning

Mingxuan Zhang, Oubo Ma, Kang Wei, Songze Li, Shouling Ji

机构 * Southeast University(东南大学) Zhejiang University(浙江大学)

专题命中 Agent评测 :agent(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01857 2025-11-18 math.NA cs.LG cs.NA math.OC 57%

Using Linearized Optimal Transport to Predict the Evolution of Stochastic Particle Systems

Nicholas Karris, Evangelos A. Nikitopoulos, Ioannis G. Kevrekidis, Seungjoon Lee, Alexander Cloninger

机构 * University of Michigan(密歇根大学) Johns Hopkins University(约翰霍普金斯大学) California State University, Long Beach(长滩加州州立大学)

专题命中 Agent评测 :agent(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.11922 2025-11-18 cs.CL 57%

Historical/temporal necessities/possibilities, and a logical theory of them in branching time

Fengkui Ju, Woxuan Zhou

机构 * School of Philosophy, Beijing Normal University(北京师范大学哲学学院) Institute for Logic, Language and Computation, University of Amsterdam(阿姆斯特丹大学逻辑、语言与计算研究所)

专题命中 Agent评测 :agent(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10585 2025-11-14 cs.SI cs.AI 57%

Textual understanding boost in the WikiRace

Raman Ebrahimi, Sean Fuhrman, Kendrick Nguyen, Harini Gurusankar, Massimo Franceschetti

机构 * University of California, San Diego, Electrical and Computer Engineering(加州大学圣地亚哥分校)

专题命中 Agent评测 :agent(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08866 2025-11-13 cs.CL 57%

BioVerge: A Comprehensive Benchmark and Study of Self-Evaluating Agents for Biomedical Hypothesis Generation

Fuyi Yang, Chenchen Ye, Mingyu Derek Ma, Yijia Xiao, Matthew Yang, Wei Wang

机构 * University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 Agent评测 :agent(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15201 2025-11-13 cs.RO cs.AI 57%

Survey of Vision-Language-Action Models for Embodied Manipulation

Haoran Li, Yuhui Chen, Wenbo Cui, Weiheng Liu, Kai Liu, Mingcai Zhou, Zhengtao Zhang, Dongbin Zhao

专题命中 Agent评测 :agent(abstract);分类 cs.AI

Comments in Chinese language

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08060 2025-11-12 cs.CR cs.SE 57%

From LLMs to Agents: A Comparative Evaluation of LLMs and LLM-based Agents in Security Patch Detection

Junxiao Han, Zheng Yu, Lingfeng Bao, Jiakun Liu, Yao Wan, Jianwei Yin, Shuiguang Deng, Song Han

专题命中 Agent评测 :agent(abstract);分类 cs.SE

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07831 2025-11-12 stat.ML cs.LG 57%

Distributionally Robust Online Markov Game with Linear Function Approximation

Zewu Zheng, Yuanyuan Lin

机构 * Zewu Zheng, Yuanyuan Lin(作者)

专题命中 Agent评测 :agent(abstract);分类 cs.LG

Comments To be published in the Proceedings of AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07794 2025-11-12 cs.CL 57%

Design, Results and Industry Implications of the World's First Insurance Large Language Model Evaluation Benchmark

Hua Zhou, Bing Ma, Yufei Zhang, Yi Zhao

专题命中 Agent评测 :agent(abstract);分类 cs.CL

Comments 16 pages, 11 models,1 set of evaluation framework,5 core dimensions, 54 sub-indicators, 14,430 high-quality questions

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07702 2025-11-12 cs.LG physics.comp-ph physics.flu-dyn 57%

Intelligent Optimization of Multi-Parameter Micromixers Using a Scientific Machine Learning Framework

Meraj Hassanzadeh, Ehsan Ghaderi, Mohamad Ali Bijarchi, Siamak Kazemzadeh Hannani

专题命中 Agent评测 :agent(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18905 2025-11-12 cs.AI 57%

How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective

Songsong Yu, Yuxin Chen, Hao Ju, Lianjie Jia, Fuxi Zhang, Shaofei Huang, Yuhan Wu, Rundi Cui, Binghao Ran, Zaibin Zhang, Zhedong Zheng, Zhipeng Zhang, Yifan Wang, Lin Song, Lijun Wang, Yanwei Li, Ying Shan, Huchuan Lu

机构 * ARC Lab, Tencent PCG(腾讯PCG部门ARC实验室) Shanghai Jiao Tong University(上海交通大学) University of Macau(澳门大学) Dalian University of Technology(大连理工大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 Agent评测 :planning(abstract);分类 cs.AI

Comments a comprehensive visual spatial reasoning evaluation tool, 25 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11154 2025-11-12 cs.CR cs.CL 57%

MPMA: Preference Manipulation Attack Against Model Context Protocol

Zihan Wang, Rui Zhang, Yu Liu, Wenshu Fan, Wenbo Jiang, Qingchuan Zhao, Hongwei Li, Guowen Xu

专题命中 Agent评测 :agent(abstract);分类 cs.CL

Comments This is an extended version of the copyrighted publication at AAAI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14567 2025-11-12 cs.OS cs.AI 57%

Integrating Artificial Intelligence into Operating Systems: A Survey on Techniques, Applications, and Future Directions

Yifan Zhang, Xinkui Zhao, Ziying Li, Guanjie Cheng, Jianwei Yin, Lufei Zhang, Zuoning Chen

机构 * School of Software Technology, Zhejiang University(浙江大学软件学院) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) State Key Laboratory of Mathematical Engineering and Advanced Computing(数学工程与先进计算国家重点实验室) Chinese Academy of Engineering(中国工程院) College of Cyber Security, Jinan University(暨南大学网络安全学院)

专题命中 Agent评测 :agent(abstract);分类 cs.AI

Comments 68 pages,9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07564 2025-11-12 physics.flu-dyn cs.LG 57%

Shocks Under Control: Taming Transonic Compressible Flow over an RAE2822 Airfoil with Deep Reinforcement Learning

Trishit Mondal, Ricardo Vinuesa, Ameya D. Jagtap

机构 * Aerospace Engineering Department, Worcester Polytechnic Institute(航空航天工程系,沃斯特理工学院) Department of Aerospace Engineering, University of Michigan(航空航天工程系,密歇根大学)

专题命中 Agent评测 :agent(abstract);分类 cs.LG

Comments 23 pages, 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07486 2025-11-12 cs.LG cs.SY eess.SY stat.ML 57%

Provably Efficient Sample Complexity for Robust CMDP

Sourav Ganguly, Arnob Ghosh

机构 * New Jersey Institute of Technology(新泽西理工学院)

专题命中 Agent评测 :agent(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07434 2025-11-12 q-fin.ST cs.LG q-fin.TR 57%

RL-Exec: Impact-Aware Reinforcement Learning for Opportunistic Optimal Liquidation, Outperforms TWAP and a Book-Liquidity VWAP on BTC-USD Replays

Enzo Duflot, Stanislas Robineau

机构 * Enzo Duflot, Stanislas Robineau

专题命中 Agent评测 :agent(abstract);分类 cs.LG

Comments 8 pages main text, 3 appendix pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07025 2025-11-11 cs.CL cs.IR 57%

Llama-Embed-Nemotron-8B: A Universal Text Embedding Model for Multilingual and Cross-Lingual Tasks

Yauhen Babakhin, Radek Osmulski, Ronay Ak, Gabriel Moreira, Mengyao Xu, Benedikt Schifferer, Bo Liu, Even Oldridge

专题命中 Agent评测 :planning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12059 2025-11-11 cs.CL 57%

Evaluating the Ability of Large Language Models to Reason about Cardinal Directions, Revisited

Anthony G Cohn, Robert E Blackwell

机构 * School of Computer Science, University of Leeds, UK(利兹大学计算机科学学院) Alan Turing Institute, London, UK(艾伦·图灵研究所)

专题命中 Agent评测 :agent(abstract);分类 cs.CL

Comments 8 pages, 5 figures. Accepted at QR 2025 : 38th International Workshop on Qualitative Reasoning at IJCAI. arXiv admin note: substantial text overlap with arXiv:2406.16528

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21158 2025-11-11 cs.LG 57%

Diverse Mini-Batch Selection in Reinforcement Learning for Efficient Chemical Exploration in de novo Drug Design

Hampus Gummesson Svensson, Ola Engkvist, Jon Paul Janet, Christian Tyrchan, Morteza Haghir Chehreghani

专题命中 Agent评测 :agent(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06120 2025-11-11 cs.PL cs.LG 57%

A Deep Learning Model for Predicting Transformation Legality

Avani Tiwari, Yacine Hakimi, Riyadh Baghdadi

机构 * Department of Computer Science(计算机科学系) New York University Abu Dhabi(纽约大学阿布扎克分校) Laboratoire de Méthodes de Conception de Systèmes(系统设计方法实验室) Ecole nationale Supérieure d'Informatique(信息科学高等学院)

专题命中 Agent评测 :agent(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏