arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

2025-09-23 至 2025-09-23 共收录 21 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. Agent评测 21 篇

2509.17353 2025-09-23 cs.AI eess.IV physics.med-ph 88%

Medical AI Consensus: A Multi-Agent Framework for Radiology Report Generation and Evaluation

Ahmed T. Elboardy, Ghada Khoriba, Essam A. Rashed

机构 * Graduate School of Information Science, University of Hyogo(京都大学垣田学园信息科学研究生院) Center for Informatics Science, School of Information Technology and Computer Science, Nile University(尼罗大学信息科学中心)

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI

Comments NeurIPS2025 Workshop: Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20383 2025-09-23 cs.LG cs.CL 84%

Why Are Web AI Agents More Vulnerable Than Standalone LLMs? A Security Analysis

Jeffrey Yang Fan Chiang, Seungjae Lee, Jia-Bin Huang, Furong Huang, Yizheng Chen

机构 * University of Maryland(马里兰大学)

专题命中 Agent评测 :AI agent(title,abstract);agent(abstract);分类 cs.CL、cs.LG

Comments Project website: http://vulnerable-ai-agents.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17259 2025-09-23 cs.AI 79%

Mind the Gap: Comparing Model- vs Agentic-Level Red Teaming with Action-Graph Observability on GPT-OSS-20B

Ilham Wicaksono, Zekun Wu, Rahul Patel, Theo King, Adriano Koshiyama, Philip Treleaven

机构 * University College London(伦敦大学学院) Holistic AI(整体AI)

专题命中 Agent评测 :agentic(title,abstract);分类 cs.AI

Comments Winner of the OpenAI GPT-OSS-20B Red Teaming Challenge (Kaggle, 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11049 2025-09-23 cs.CL 79%

Journalism-Guided Agentic In-Context Learning for News Stance Detection

Dahyun Lee, Jonghyeon Choi, Jiyoung Han, Kunwoo Park

专题命中 Agent评测 :agentic(title);agent(abstract);分类 cs.CL

Comments EMNLP 2025 (24 pages)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16784 2025-09-23 cs.HC 78%

Controlled Yet Natural: A Hybrid BDI-LLM Conversational Agent for Child Helpline Training

Mohammed Al Owayyed, Adarsh Denga, Willem-Paul Brinkman

专题命中 Agent评测 :agent(title,abstract)

Journal ref ACM International Conference on Intelligent Virtual Agents (IVA 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16275 2025-09-23 cs.CR cs.AI cs.SE 76%

SecureFixAgent: A Hybrid LLM Agent for Automated Python Static Vulnerability Repair

Jugal Gajjar, Kamalasankari Subramaniakuppusamy, Relsy Puthal, Kaustik Ranaware

机构 * Computer Science Department(计算机科学系) Applied Economics Department(应用经济学系)

专题命中 Agent评测 :agent(title);分类 cs.AI、cs.SE

Comments 6 pages, 3 figures, 4 tables, 1 algorithm, accepted in the Robustness and Security of Large Language Models (ROSE-LLM) special session at ICMLA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17488 2025-09-23 cs.CR cs.AI 70%

Privacy in Action: Towards Realistic Privacy Mitigation and Evaluation for LLM-Powered Agents

Shouju Wang, Fenglin Yu, Xirui Liu, Xiaoting Qin, Jue Zhang, Qingwei Lin, Dongmei Zhang, Saravan Rajmohan

机构 * Wuhan University(武汉大学) Microsoft(微软公司)

专题命中 Agent评测 :agent(abstract);agentic(abstract);分类 cs.AI

Comments To appear at EMNLP 2025 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16610 2025-09-23 cs.CL 70%

LLMsPark: A Benchmark for Evaluating Large Language Models in Strategic Gaming Contexts

Junhao Chen, Jingbo Sun, Xiang Li, Haidong Xin, Yuhao Xue, Yibin Xu, Hao Zhao

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) School of Software and Microelectronics, Peking University(北京大学软件与微电子学院) School of Computer Science and Engineering, Northeastern University(东北大学计算机科学与工程学院) Tongji University(同济大学) AIR, Tsinghua University(清华大学人工智能研究院) BAAI(百度人工智能研究院)

专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.CL

Comments Accepted by EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16866 2025-09-23 cs.AI cs.CL cs.LG 67%

seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMs

Mohammad Ramezanali, Mo Vazifeh, Paolo Santi

机构 * Salesforce AI Palo Alto(Salesforce AI 巴尔的摩) Capital One MIT(麻省理工学院)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16232 2025-09-23 q-bio.NC cs.HC 67%

Emotions are Recognized Patterns of Cognitive Activities

Yue Jin

专题命中 Agent评测 :agent(abstract);autonomous agent(abstract)

Comments 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00742 2025-09-23 cs.CL cs.LG 62%

Applying Psychometrics to Large Language Model Simulated Populations: Recreating the HEXACO Personality Inventory Experiment with Generative Agents

Sarah Mercer, Daniel P. Martin, Phil Swatton

机构 * The Alan Turing Institute(艾伦·图灵研究所)

专题命中 Agent评测 :agent(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16325 2025-09-23 cs.CL cs.AI cs.HC 62%

Overhearing LLM Agents: A Survey, Taxonomy, and Roadmap

Andrew Zhu, Chris Callison-Burch

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL

Comments 8 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17728 2025-09-23 cs.LG cs.MA 57%

A non-smooth regularization framework for learning over multitask graphs

Yara Zgheib, Luca Calatroni, Marc Antonini, Roula Nassif

机构 * Université Côte d’Azur, I3S Laboratory, CNRS, France(法国大学皮特里特大学I3S实验室,CNRS) Machine Learning Genoa Center (MaLGa), Department of Computer Science, University of Genoa(热那亚大学计算机科学系机器学习中心(MaLGa))

专题命中 Agent评测 :agent(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17092 2025-09-23 cs.LG 57%

On the Limits of Tabular Hardness Metrics for Deep RL: A Study with the Pharos Benchmark

Michelangelo Conserva, Remo Sasso, Paulo Rauber

机构 * School of Electronic Engineering and Computer Science(电子工程与计算机科学学院)

专题命中 Agent评测 :agent(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07501 2025-09-23 cs.GT cs.CC cs.FL cs.LO cs.MA 50%

The Complexity of Pure Strategy Relevant Equilibria in Concurrent Games

Purandar Bhaduri

专题命中 Agent评测 :agent(abstract)

Comments In Proceedings GandALF 2025, arXiv:2509.13258

Journal ref EPTCS 428, 2025, pp. 62-75

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17432 2025-09-23 physics.med-ph 50%

In vivo and predictive interplay evaluation methodology for lung and esophageal cancer patients treated in free breathing with IMPT

Giorgio Cartechini, Esther Kneepkens, Gloria Vilches-Freixas, Indra Lubken, Marije Velders, Sebastiaan Nijsten, Mirko Unipan, Ilaria Rinaldi

专题命中 Agent评测 :planning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17044 2025-09-23 cs.CV 50%

AgriDoctor: A Multimodal Intelligent Assistant for Agriculture

Mingqing Zhang, Zhuoning Xu, Peijie Wang, Rongji Li, Liang Wang, Qiang Liu, Jian Xu, Xuyao Zhang, Shu Wu, Liang Wang

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 Agent评测 :agent(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16895 2025-09-23 cs.IR 50%

Temporal-Aware User Behaviour Simulation with Large Language Models for Recommender Systems

Xinye Wanyan, Danula Hettiachchi, Chenglong Ma, Ziqi Xu, Jeffrey Chan

专题命中 Agent评测 :agent(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01500 2025-09-23 nlin.AO physics.soc-ph 50%

The role of zealots in the spread of linguistic traits

Vivian Dornelas, Celia Anteneodo, Renan Nunes, Els Heinsalu, Marco Patriarca

专题命中 Agent评测 :agent(abstract)

Comments 13 pages with 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02651 2025-09-23 cond-mat.soft 50%

Slow modulation of the contraction patterns in Physarum polycephalum

Raphael Saiseau, Valentin Busson, Marc Durand

专题命中 Agent评测 :agent(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16450 2025-09-23 cs.GT econ.TH 50%

On the Existence and Complexity of Core-Stable Data Exchanges

Jiaxin Song, Pooja Kulkarni, Parnian Shahkar, Bhaskar Ray Chaudhury

专题命中 Agent评测 :agent(abstract)

Comments 28 pages, 5 figures, accepted by NeurIPS'25

详情

展开后加载摘要…

URL PDF HTML 收藏