arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 15866 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. Agent评测 15866 篇

2511.05496 2025-11-11 cs.IR cs.AI 57%

DOCUEVAL: An LLM-based AI Engineering Tool for Building Customisable Document Evaluation Workflows

Hao Zhang, Qinghua Lu, Liming Zhu

机构 * CSIRO’s Data61(CSIRO数据61)

专题命中 Agent评测 :workflow(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04423 2025-11-11 cs.CL 57%

Evaluating, Synthesizing, and Enhancing for Customer Support Conversation

Jie Zhu, Huaixia Dou, Junhui Li, Lifan Guo, Feng Chen, Chi Zhang, Fang Kong

专题命中 Agent评测 :agent(abstract);分类 cs.CL

Comments Accepted by AAAI-2026

Journal ref The Association for the Advancement of Artificial Intelligence (AAAI),2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05231 2025-11-10 physics.soc-ph cs.LG cs.MA 57%

A differentiable model of supply-chain shocks

Saad Hamid, José Moran, Luca Mungo, Arnau Quera-Bofarull, Sebastian Towers

机构 * Aioi R&D Lab(Aioi 研发实验室) Macrocosm Institute for New Economic Thinking at the Oxford Martin School, University of Oxford(牛津大学奥克斯马丁学院新经济思想研究所) FLAIR, Foerster Lab for AI Research, University of Oxford(牛津大学FLAIR人工智能研究实验室) Department of Engineering, University of Oxford(牛津大学工程学院)

专题命中 Agent评测 :agent(abstract);分类 cs.LG

Comments Accepted to 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Differentiable Systems and Scientific Machine Learning (EurIPS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07325 2025-11-07 cs.LG 57%

CancerGUIDE: Cancer Guideline Understanding via Internal Disagreement Estimation

Alyssa Unell, Noel C. F. Codella, Sam Preston, Peniel Argaw, Wen-wai Yim, Zelalem Gero, Cliff Wong, Rajesh Jena, Eric Horvitz, Amanda K. Hall, Ruican Rachel Zhong, Jiachen Li, Shrey Jain, Mu Wei, Matthew Lungren, Hoifung Poon

机构 * Microsoft Research Stanford University(微软研究院斯坦福大学) Microsoft Health and Life Sciences Project Mentors(微软健康与生命科学项目导师) Microsoft Health and Life Sciences(微软健康与生命科学) Microsoft Research(微软研究院) Microsoft Research University of Washington(微软研究院华盛顿大学) Microsoft Health and Life Sciences Equal Leadership(微软健康与生命科学同等领导)

专题命中 Agent评测 :agent(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03730 2025-11-07 cs.HC cs.AI 57%

Not All Explanations are Created Equal: Investigating the Pitfalls of Current XAI Evaluation

Joe Shymanski, Jacob Brue, Sandip Sen

专题命中 Agent评测 :agent(abstract);分类 cs.AI

Comments The authors' accepted manuscript of Chapter 9 in Bi-directionality in Human-AI Collaborative Systems (Springer, 2025). The final published version is available at https://doi.org/10.1016/B978-0-44-340553-2.00015-0. 27 pages, 12 figures, 3 tables

Journal ref William Lawless, Ranjeev Mittu, Donald Sofge, Marco Brambilla, Bi-directionality in Human-AI Collaborative Systems, 2025, Pages 227-251

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23558 2025-11-07 cs.CR cs.AI 57%

Transferable & Stealthy Ensemble Attacks: A Black-Box Jailbreaking Framework for Large Language Models

Yiqi Yang, Hongye Fu

专题命中 Agent评测 :agent(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03186 2025-11-06 cs.AI 57%

Adobe Summit Concierge Evaluation with Human in the Loop

Yiru Chen, Sally Fang, Sai Sree Harsha, Dan Luo, Vaishnavi Muppala, Fei Wu, Shun Jiang, Kun Qian, Yunyao Li

机构 * Adobe Inc.(Adobe公司)

专题命中 Agent评测 :workflow(abstract);分类 cs.AI

Comments Accepted by 6th Workshop on Data Science with Human in the Loop @ VLDB 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03047 2025-11-06 cs.LG 57%

Unsupervised Evaluation of Multi-Turn Objective-Driven Interactions

Emi Soroka, Tanmay Chopra, Krish Desai, Sanjay Lall

机构 * Department of Electrical Engineering Stanford University(电气工程系 斯坦福大学) Emissary Technologies

专题命中 Agent评测 :AI agent(abstract);分类 cs.LG

Comments Under review at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07003 2025-11-05 cs.RO cs.AI cs.HC 57%

Enhancing Robot Assistive Behaviour with Reinforcement Learning and Theory of Mind

Antonio Andriella, Giovanni Falcone, Silvia Rossi

机构 * Artificial Intelligence Research Institute (IIIA-CSIC)(人工智能研究 institute(IIIA-CSIC)) Department of Electrical Engineering and Information Technologies(电气工程与信息科技系) University of Naples Federico II(那不勒斯费德里科二世大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02094 2025-11-05 cs.AI 57%

Automated Reward Design for Gran Turismo

Michel Ma, Takuma Seno, Kaushik Subramanian, Peter R. Wurman, Peter Stone, Craig Sherstan

机构 * Mila, University of Montreal(蒙特利尔大学Mila) Turing Inc.(图灵公司) Sony AI(索尼人工智能) Sony AI, UT Austin(索尼人工智能、得克萨斯大学奥斯汀分校)

专题命中 Agent评测 :agent(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01757 2025-11-04 cs.SE 57%

Towards LLM-Powered Task-Aware Retrieval of Scientific Workflows for Galaxy

Shamse Tasnim Cynthia, Banani Roy

专题命中 Agent评测 :workflow(abstract);分类 cs.SE

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01541 2025-11-04 cs.CV cs.AI 57%

Driving scenario generation and evaluation using a structured layer representation and foundational models

Arthur Hubert, Gamal Elghazaly, Raphaël Frank

机构 * Interdisciplinary Centre for Security, Reliability and Trust (SnT), University of Luxembourg(安全、可靠性与信任跨学科中心(SnT),卢森堡大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00048 2025-11-04 cs.AI cs.CY 57%

GEPOC Parameters -- Open Source Parametrisation and Validation for Austria, Version 2.0

Martin Bicher, Maximilian Viehauser, Daniele Giannandrea, Hannah Kastinger, Dominik Brunmeir, Claire Rippinger, Christoph Urach, Niki Popper

机构 * TU Wien, Institute of Information Systems Engineering(维也纳技术大学信息系统工程学院) TU Wien, Institute of Statistics and Mathematical Methods in Economics(维也纳技术大学统计与经济学数学方法研究所) dwh GmbH, dwh simulation services(dwh GmbH 模拟服务)

专题命中 Agent评测 :agent(abstract);分类 cs.AI

Comments 134 pages, 75 figures, 19 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00899 2025-11-04 cs.LO cs.AI math.LO 57%

Dynamic Logic of Trust-Based Beliefs

Junli Jiang, Pavel Naumov, Wenxuan Zhang

机构 * Institute of Logic and Intelligence, Southwest University, China(逻辑与智能研究所,西南大学,中国) University of Southampton, United Kingdom(南安普顿大学,英国) Independent Scholar, United States(独立学者,美国)

专题命中 Agent评测 :agent(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00768 2025-11-04 cs.SI cs.LG 57%

A Framework Based on Graph Cellular Automata for Similarity Evaluation in Urban Spatial Networks

Peiru Wu, Maojun Zhai, Lingzhu Zhang

专题命中 Agent评测 :planning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00112 2025-11-04 cs.RO cs.AI 57%

Real-DRL: Teach and Learn in Reality

Yanbing Mao, Yihao Cai, Lui Sha

机构 * Engineering Technology Division(工程科技部门) Wayne State University(韦恩州立大学) Department of Electrical and Computer Engineering(电气与计算机工程系) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 Agent评测 :agent(abstract);分类 cs.AI

Comments 37 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18384 2025-11-03 cs.CR cs.AI 57%

Dynamic Risk Assessments for Offensive Cybersecurity Agents

Boyi Wei, Benedikt Stroebl, Jiacen Xu, Joie Zhang, Zhou Li, Peter Henderson

机构 * Princeton University(普林斯顿大学) Microsoft(微软公司) University of California Irvine(加州大学尔湾分校)

专题命中 Agent评测 :agent(abstract);分类 cs.AI

Comments 26 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26494 2025-10-31 cs.SI cs.AI 57%

Simulating and Experimenting with Social Media Mobilization Using LLM Agents

Sadegh Shirani, Mohsen Bayati

机构 * Graduate School of Business, Stanford University(斯坦福大学商学院)

专题命中 Agent评测 :agent(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26238 2025-10-31 cs.AI 57%

Questionnaire meets LLM: A Benchmark and Empirical Study of Structural Skills for Understanding Questions and Responses

Duc-Hai Nguyen, Vijayakumar Nanjappan, Barry O'Sullivan, Hoang D. Nguyen

机构 * University College Cork(科克大学)

专题命中 Agent评测 :workflow(abstract);分类 cs.AI

Comments 14 pages, 3 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26057 2025-10-31 cs.AI cs.HC 57%

Can AI be Accountable?

Andrew L. Kun

专题命中 Agent评测 :agent(abstract);分类 cs.AI

Comments To be published as a chapter in Daniele Quercia and Marios Constantinides (Eds.). Operationalizing Responsible AI. Cambridge University Press. Forthcoming

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24014 2025-10-31 cs.CL 57%

TEXT2DB: Integration-Aware Information Extraction with Large Language Model Agents

Yizhu Jiao, Sha Li, Sizhe Zhou, Heng Ji, Jiawei Han

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 Agent评测 :agent(abstract);分类 cs.CL

Comments Source code: https://github.com/yzjiao/Text2DB

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11695 2025-10-31 cs.CL 57%

When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents

Lingfei Qian, Xueqing Peng, Yan Wang, Vincent Jim Zhang, Huan He, Hanley Smith, Yi Han, Yueru He, Haohang Li, Yupeng Cao, Yangyang Yu, Alejandro Lopez-Lira, Peng Lu, Jian-Yun Nie, Guojun Xiong, Jimin Huang, Sophia Ananiadou

机构 * Georgia Institute of Technology(佐治亚理工学院) Columbia University(哥伦比亚大学) Stevens Institute of Technology(史蒂文斯理工学院) University of Florida(佛罗里达大学) Harvard University(哈佛大学) National Centre for Text Mining, University of Manchester(曼彻斯特大学文本挖掘中心)

专题命中 Agent评测 :agent(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24432 2025-10-29 cs.LG 57%

Fill in the Blanks: Accelerating Q-Learning with a Handful of Demonstrations in Sparse Reward Settings

Seyed Mahdi Basiri Azad, Joschka Boedecker

机构 * Faculty of Engineering University of Freiburg(工程学院弗赖堡大学)

专题命中 Agent评测 :agent(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24287 2025-10-29 eess.SP cs.LG 57%

Towards actionable hypotension prediction -- predicting catecholamine therapy initiation in the intensive care unit

Richard Koebe, Noah Saibel, Juan Miguel Lopez Alcaraz, Simon Schäfer, Nils Strodthoff

机构 * University Clinic of Anesthesiology, Intensive Care Medicine, Emergency Medicine, and Pain Therapy(麻醉学、重症医学、急诊医学和疼痛治疗大学诊所) AI4Health Division(AI4Health部门) Carl von Ossietzky Universität(卡尔·冯·奥西特茨基大学)

专题命中 Agent评测 :agent(abstract);分类 cs.LG

Comments 27 pages, 8 figures, source code under https://github.com/AI4HealthUOL/actionable-hypotension

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24958 2025-10-29 cs.CL 57%

The Dialogue That Heals: A Comprehensive Evaluation of Doctor Agents' Inquiry Capability

Linlu Gong, Ante Wang, Yunghwei Lai, Weizhi Ma, Yang Liu

机构 * Department of Computer Science and Technology, Tsinghua University, China(清华大学计算机科学与技术系) Institute for AI Industry Research (AIR), Tsinghua University, China(清华大学人工智能产业研究院)

专题命中 Agent评测 :agent(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23011 2025-10-28 cs.CL cs.CY 57%

LangLingual: A Personalised, Exercise-oriented English Language Learning Tool Leveraging Large Language Models

Sammriddh Gupta, Sonit Singh, Aditya Joshi, Mira Kim

专题命中 Agent评测 :agent(abstract);分类 cs.CL

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22685 2025-10-28 q-fin.CP cs.AI cs.MA q-fin.TR 57%

TABL-ABM: A Hybrid Framework for Synthetic LOB Generation

Ollie Olby, Rory Baggott, Namid Stillman

机构 * Simudyne

专题命中 Agent评测 :agent(abstract);分类 cs.AI

Comments 8 pages, 5 figures, accepted to the Workshop on AI in Finance at ECAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22443 2025-10-28 cs.CV cs.LG 57%

Benchmarking Egocentric Multimodal Goal Inference for Assistive Wearable Agents

Vijay Veerabadran, Fanyi Xiao, Nitin Kamra, Pedro Matias, Joy Chen, Caley Drooff, Brett D Roads, Riley Williams, Ethan Henderson, Xuanyi Zhao, Kevin Carlberg, Joseph Tighe, Karl Ridgeway

机构 * Meta Reality Labs(Meta 现实实验室) Meta FAIR

专题命中 Agent评测 :agent(abstract);分类 cs.LG

Comments Accepted as a spotlight paper at the 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21347 2025-10-28 cs.LG 57%

OVERT: A Benchmark for Over-Refusal Evaluation on Text-to-Image Models

Ziheng Cheng, Yixiao Huang, Hui Xu, Somayeh Sojoudi, Xuandong Zhao, Dawn Song, Song Mei

机构 * UC Berkeley(加州大学伯克利分校)

专题命中 Agent评测 :workflow(abstract);分类 cs.LG

Comments NeurIPS 2025 (D&B Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22045 2025-10-28 cs.CV cs.AI 57%

VLM-SlideEval: Evaluating VLMs on Structured Comprehension and Perturbation Sensitivity in PPT

Hyeonsu Kang, Emily Bao, Anjan Goswami

专题命中 Agent评测 :agentic(abstract);分类 cs.AI

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Evaluating the Evolving LLM Lifecycle - Benchmarks, Emergent Abilities, and Scaling

详情

展开后加载摘要…

URL PDF HTML 收藏