arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

2025-10-29 至 2025-10-29 共收录 17 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. Agent评测 17 篇

2510.24397 2025-10-29 cs.AI 85%

APTBench: Benchmarking Agentic Potential of Base LLMs During Pre-Training

Jiarui Qin, Yunjia Xi, Junjie Huang, Renting Rui, Di Yin, Weiwen Liu, Yong Yu, Weinan Zhang, Xing Sun

机构 * Tencent Youtu Lab(腾讯优图实验室) Shanghai Jiao Tong University(上海交通大学)

专题命中 Agent评测 :agentic(title,abstract);agent(abstract);planning(abstract);分类 cs.AI

Comments 46 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24438 2025-10-29 cs.CL cs.AI cs.CY cs.MA 81%

Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content

Abdullah Mushtaq, Rafay Naeem, Ezieddin Elmahjub, Ibrahim Ghaznavi, Shawqi Al-Maliki, Mohamed Abdallah, Ala Al-Fuqaha, Junaid Qadir

机构 * Information Technology University(信息科技大学) Qatar University(卡塔尔大学) Hamad Bin Khalifa University(哈马德·本·卡西姆大学)

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI、cs.CL

Comments Accepted at 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: 5th Muslims in Machine Learning (MusIML) Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24259 2025-10-29 cs.CL cs.RO 79%

Can LLMs Translate Human Instructions into a Reinforcement Learning Agent's Internal Emergent Symbolic Representation?

Ziqi Ma, Sao Mai Nguyen, Philippe Xu

机构 * U2IS, ENSTA, IP-Paris(U2IS、ENSTA、IP-巴黎)

专题命中 Agent评测 :agent(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23875 2025-10-29 cs.HC 78%

Large Language Model Agent Personality and Response Appropriateness: Evaluation by Human Linguistic Experts, LLM-as-Judge, and Natural Language Processing Model

Eswari Jayakumar, Niladri Sekhar Dash, Debasmita Mukherjee

专题命中 Agent评测 :agent(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21071 2025-10-29 econ.GN cs.MA q-fin.EC 78%

Central Bank Digital Currency, Flight-to-Quality, and Bank-Runs in an Agent-Based Model

Emilio Barucci, Andrea Gurgone, Giulia Iori, Michele Azzone

专题命中 Agent评测 :agent(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24317 2025-10-29 cs.CR 71%

Cybersecurity AI Benchmark (CAIBench): A Meta-Benchmark for Evaluating Cybersecurity AI Agents

María Sanz-Gómez, Víctor Mayoral-Vilches, Francesco Balassone, Luis Javier Navarrete-Lozano, Cristóbal R. J. Veas Chavez, Maite del Mundo de Torres

专题命中 Agent评测 :AI agent(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17734 2025-10-29 cs.LG 70%

URB -- Urban Routing Benchmark for RL-equipped Connected Autonomous Vehicles

Ahmet Onur Akman, Anastasia Psarou, Michał Hoffmann, Łukasz Gorczyca, Łukasz Kowalski, Paweł Gora, Grzegorz Jamróz, Rafał Kucharski

机构 * Doctoral School of Exact and Natural Sciences, Jagiellonian University(杰尔吉利亚大学精确与自然科学博士学院) Faculty of Mathematics and Computer Science, Jagiellonian University(杰尔吉利亚大学数学与计算机科学学院) Urban Policy Observatory, Institute of Urban and Regional Development(城市政策观察所,城市与区域发展研究所)

专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.LG

Comments Accepted at the 39th Conference on Neural Information Processing Systems (NeurIPS 2025), Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24356 2025-10-29 cs.LG cs.AI stat.ML 62%

Perception Learning: A Formal Separation of Sensory Representation Learning from Decision Learning

Suman Sanyal

机构 * Big Data Analytics, Goa Institute of Management, Goa, India(大数据分析,果阿管理学院,果阿,印度)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16175 2025-10-29 cs.LG cs.AI 62%

The Formalism-Implementation Gap in Reinforcement Learning Research

Pablo Samuel Castro

机构 * Google DeepMind(谷歌DeepMind) Université de Montréal(蒙特利尔大学) Mila - Québec AI Institute(魁北克AI研究院)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24695 2025-10-29 cs.CL 61%

AgentFrontier: Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesis

Xuanzhong Chen, Zile Qiao, Guoxin Chen, Liangcai Su, Zhen Zhang, Xinyu Wang, Pengjun Xie, Fei Huang, Jingren Zhou, Yong Jiang

机构 * Tongyi Lab(通义实验室) Alibaba Group(阿里巴巴集团)

专题命中 Agent评测 :agent(abstract,comments);分类 cs.CL

Comments https://tongyi-agent.github.io/blog/introducing-tongyi-deep-research/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24432 2025-10-29 cs.LG 57%

Fill in the Blanks: Accelerating Q-Learning with a Handful of Demonstrations in Sparse Reward Settings

Seyed Mahdi Basiri Azad, Joschka Boedecker

机构 * Faculty of Engineering University of Freiburg(工程学院弗赖堡大学)

专题命中 Agent评测 :agent(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24287 2025-10-29 eess.SP cs.LG 57%

Towards actionable hypotension prediction -- predicting catecholamine therapy initiation in the intensive care unit

Richard Koebe, Noah Saibel, Juan Miguel Lopez Alcaraz, Simon Schäfer, Nils Strodthoff

机构 * University Clinic of Anesthesiology, Intensive Care Medicine, Emergency Medicine, and Pain Therapy(麻醉学、重症医学、急诊医学和疼痛治疗大学诊所) AI4Health Division(AI4Health部门) Carl von Ossietzky Universität(卡尔·冯·奥西特茨基大学)

专题命中 Agent评测 :agent(abstract);分类 cs.LG

Comments 27 pages, 8 figures, source code under https://github.com/AI4HealthUOL/actionable-hypotension

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24958 2025-10-29 cs.CL 57%

The Dialogue That Heals: A Comprehensive Evaluation of Doctor Agents' Inquiry Capability

Linlu Gong, Ante Wang, Yunghwei Lai, Weizhi Ma, Yang Liu

机构 * Department of Computer Science and Technology, Tsinghua University, China(清华大学计算机科学与技术系) Institute for AI Industry Research (AIR), Tsinghua University, China(清华大学人工智能产业研究院)

专题命中 Agent评测 :agent(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24467 2025-10-29 q-fin.TR math.OC q-fin.MF q-fin.ST 50%

The Omniscient, yet Lazy, Investor

Stanisław M. S. Halkiewicz

专题命中 Agent评测 :agent(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24447 2025-10-29 physics.soc-ph cs.SI 50%

Pair Approximation Meets Reality: Diffusion of Innovation in Organizational Networks within the biased-independence q-Voter Model

Angelika Abramiuk-Szurlej, Katarzyna Sznajd-Weron

专题命中 Agent评测 :agent(abstract)

Comments 13 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03520 2025-10-29 cs.CV 50%

Is Sora a World Simulator? A Comprehensive Survey on General World Models and Beyond

Zheng Zhu, Xiaofeng Wang, Wangbo Zhao, Chen Min, Bohan Li, Nianchen Deng, Min Dou, Yuqi Wang, Botian Shi, Kai Wang, Chi Zhang, Yang You, Zhaoxiang Zhang, Dawei Zhao, Liang Xiao, Jian Zhao, Jiwen Lu, Guan Huang

专题命中 Agent评测 :autonomous agent(abstract)

Comments This survey will be regularly updated at: https://github.com/GigaAI-research/General-World-Models-Survey

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24107 2025-10-29 physics.soc-ph cs.GT cs.SI nlin.AO 50%

Exploring Emergent Topological Properties in Socio-Economic Networks through Learning Heterogeneity

Chanuka Karavita, Zehua Lyu, Dharshana Kasthurirathna, Mahendra Piraveenan

专题命中 Agent评测 :agent(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏