arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 15756 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. Agent评测 15756 篇

2506.02720 2025-10-27 cs.AI cs.CL 62%

LocalGPT: Benchmarking and Advancing Large Language Models for Local Life Services in Meituan

Xiaochong Lan, Jie Feng, Jiahuan Lei, Xinlei Shi, Yong Li

机构 * Department of Electronic\ , BNRist,\ University Beijing, China Department of Electronic\ , BNRist,\ University

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL

Comments KDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06771 2025-10-27 cs.AI cs.CV cs.LG 62%

Proactive Agents for Multi-Turn Text-to-Image Generation Under Uncertainty

Meera Hahn, Wenjun Zeng, Nithish Kannen, Rich Galt, Kartikeya Badola, Been Kim, Zi Wang

机构 * Google DeepMind(谷歌DeepMind)

专题命中 Agent评测 :workflow(abstract);分类 cs.AI、cs.LG

Journal ref International Conference on Machine Learning, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20270 2025-10-24 cs.LG cs.CL 62%

ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases

Ziqian Zhong, Aditi Raghunathan, Nicholas Carlini

专题命中 Agent评测 :agent(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19892 2025-10-24 cs.CL cs.AI 62%

Can They Dixit? Yes they Can! Dixit as a Playground for Multimodal Language Model Capabilities

Nishant Balepur, Dang Nguyen, Dayeon Ki

机构 * University of Maryland(马里兰大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL

Comments Accepted as a Spotlight paper at the EMNLP 2025 Wordplay Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19631 2025-10-23 cs.AI cs.CL cs.MA 62%

HSCodeComp: A Realistic and Expert-level Benchmark for Deep Search Agents in Hierarchical Rule Application

Yiqian Yang, Tian Lan, Qianghuai Jia, Li Zhu, Hui Jiang, Hang Zhu, Longyue Wang, Weihua Luo, Kaifu Zhang

机构 * Alibaba International Digital Commerce(阿里巴巴国际数字商业)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03931 2025-10-23 cs.CL cs.AI 62%

NAACL2025 Tutorial: Adaptation of Large Language Models

Zixuan Ke, Yifei Ming, Shafiq Joty

机构 * Salesforce AI Research(Salesforce人工智能研究)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL

Comments NAACL2025 Tutorial

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18488 2025-10-22 cs.AI cs.SE 62%

AndroidControl-Curated: Revealing the True Potential of GUI Agents through Benchmark Purification

Ho Fai Leung, Xiaoyan Xi, Fei Zuo

机构 * BMW ArcherMind Information Technology Co. Ltd. (BA TechWorks)(宝马阿彻Mind信息技术有限公司(BA TechWorks))

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.SE

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18143 2025-10-22 cs.AI cs.LG 62%

Learning from Generalization Patterns: An Evaluation-Driven Approach to Enhanced Data Augmentation for Fine-Tuning Small Language Models

Huan Song, Deeksha Razdan, Yiyue Qian, Arijit Ghosh Chowdhury, Parth Patwa, Aman Chadha, Shinan Zhang, Sharlina Keshava, Hannah Marlowe

机构 * AWS Generative AI Innovation Center(AWS生成式AI创新中心)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

Comments Neural Information Processing Systems (NeurIPS 2025) Workshop: Evaluating the Evolving LLM Lifecycle

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.19398 2025-10-22 cs.AI cs.GT cs.LG 62%

Do LLMs Strategically Reveal, Conceal, and Infer Information? A Theoretical and Empirical Analysis in The Chameleon Game

Mustafa O. Karabag, Jan Sobotka, Ufuk Topcu

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19850 2025-10-21 cs.LG cs.AI cs.RO 62%

DISCOVER: Automated Curricula for Sparse-Reward Reinforcement Learning

Leander Diaz-Bone, Marco Bagatella, Jonas Hübotter, Andreas Krause

机构 * ETH Zürich, Switzerland(苏黎世联邦理工学院,瑞士) Max Planck Institute for Intelligent Systems, Germany(智能系统马克斯·普朗克研究所,德国)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13636 2025-10-21 cs.LG cs.AI cs.GT 62%

Incentivizing Truthful Language Models via Peer Elicitation Games

Baiting Chen, Tong Zhu, Jiale Han, Lexin Li, Gang Li, Xiaowu Dai

机构 * UCLA(加州大学洛杉矶分校) UC Berkeley(伯克利大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26490 2025-10-20 cs.CL cs.AI 62%

VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications

Wei He, Yueqing Sun, Hongyan Hao, Xueyuan Hao, Zhikang Xia, Qi Gu, Chengcheng Han, Dengchang Zhao, Hui Su, Kefeng Zhang, Man Gao, Xi Su, Xiaodong Cai, Xunliang Cai, Yu Yang, Yunke Zhao

机构 * Meituan LongCat Team(美团LongCat团队)

专题命中 Agent评测 :AI agent(abstract);分类 cs.AI、cs.CL

Comments The code, dataset, and leaderboard are available at https://vitabench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20934 2025-10-17 cs.SE cs.AI 62%

Leveraging LLMs, IDEs, and Semantic Embeddings for Automated Move Method Refactoring

Abhiram Bellur, Fraol Batole, Mohammed Raihan Ullah, Malinda Dilhara, Yaroslav Zharov, Timofey Bryksin, Kai Ishikawa, Haifeng Chen, Masaharu Morimoto, Shota Motoura, Takeo Hosomi, Tien N. Nguyen, Hridesh Rajan, Nikolaos Tsantalis, Danny Dig

机构 * University of Colorado(科罗拉多大学) Tulane University(路易斯安那州立大学) Amazon Web Services(亚马逊网络服务) JetBrains Research(JetBrains研究) NEC Corporation(日本电报电话公司) NEC Laboratories America(日本电报电话美洲实验室) University of Texas at Dallas(德克萨斯大学达拉斯分校) Concordia University(康科迪亚大学) University of Colorado, JetBrains Research(科罗拉多大学,JetBrains研究)

专题命中 Agent评测 :workflow(abstract);分类 cs.AI、cs.SE

Comments Published at the International Conference on Software Maintenance and Evolution (ICSME'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04426 2025-10-17 cs.CL cs.AI cs.CY 62%

The simulation of judgment in LLMs

Edoardo Loru, Jacopo Nudo, Niccolò Di Marco, Alessandro Santirocchi, Roberto Atzeni, Matteo Cinelli, Vincenzo Cestari, Clelia Rossi-Arnaud, Walter Quattrociocchi

机构 * Department of Computer, Control and Management Engineering, Sapienza University of Rome(计算机、控制与管理工程系,罗马萨皮恩扎大学) Department of Computer Science, Sapienza University of Rome(计算机科学系,罗马萨皮恩扎大学) Department of Legal, Social, and Educational Sciences, Tuscia University(法律、社会与教育科学系,图斯西亚大学) Department of Psychology, Sapienza University of Rome(心理学系,罗马萨皮恩扎大学)

专题命中 Agent评测 :agentic(abstract);分类 cs.AI、cs.CL

Comments Please refer to published version: https://doi.org/10.1073/pnas.2518443122

Journal ref Proc. Natl. Acad. Sci. U.S.A. 122 (42) e2518443122, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13052 2025-10-16 cs.LG cs.AI cs.SY eess.SP eess.SY math.OC 62%

Time-Varying Optimization for Streaming Data Via Temporal Weighting

Muhammad Faraz Ul Abrar, Nicolò Michelusi, Erik G. Larsson

机构 * School of Electrical, Computer and Energy Engineering, Arizona State University(亚利桑那州立大学电气、计算机与能源工程学院) Department of Electrical Engineering (ISY), Linköping University(利尔贝里大学电气工程系)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

Comments Accepted at IEEE Asilomar, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01540 2025-10-16 cs.LG cs.AI 62%

BoxingGym: Benchmarking Progress in Automated Experimental Design and Model Discovery

Kanishk Gandhi, Michael Y. Li, Lyle Goodyear, Agam Bhatia, Louise Li, Aditi Bhaskar, Mohammed Zaman, Noah D. Goodman

机构 * Stanford University(斯坦福大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

Comments KG and MYL contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10586 2025-10-14 cs.LG cs.AI cs.IT math.IT q-bio.NC 62%

Compositional Symmetry as Compression: Lie Pseudogroup Structure in Algorithmic Agents

Giulio Ruffini

机构 * Neuroelectrics Starlab BCOM (Barcelona Spain)(神经电学星实验室BCOM(巴塞罗那西班牙))

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

Comments Submitted to NeurReps 2025 (https://www.neurreps.org)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09738 2025-10-14 cs.CL cs.AI 62%

Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement

Steve Han, Gilberto Titericz Junior, Tom Balough, Wenfei Zhou

机构 * NVIDIA Corporation(NVIDIA公司)

专题命中 Agent评测 :agentic(abstract);分类 cs.AI、cs.CL

Comments 10 pages, 1 figure, 4 tables, under review as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20036 2025-10-14 cs.SE cs.AI 62%

Agents in the Sandbox: End-to-End Crash Bug Reproduction for Minecraft

Eray Yapağcı, Yavuz Alp Sencer Öztürk, Eray Tüzün

机构 * Electronics Engineering Department Bilkent University Ankara, Turkey(电子工程系比尔肯特大学安卡拉土耳其) Computer Engineering Department Bilkent University Ankara, Turkey(计算机工程系比尔肯特大学安卡拉土耳其)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.SE

Comments Accepted into ASE'25 (Automated Software Engineering). https://conf.researchr.org/details/ase-2025/ase-2025-papers/244/Agents-in-the-Sandbox-End-to-End-Crash-Bug-Reproduction-for-Minecraft

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02360 2025-10-09 cs.CL cs.AI 62%

Spiral of Silence in Large Language Model Agents

Mingze Zhong, Meng Fang, Zijing Shi, Yuxuan Huang, Shunfeng Zheng, Yali Du, Ling Chen, Jun Wang

机构 * AAII, University of Technology Sydney(AAII,悉尼大学) University of Liverpool(利物浦大学) King’s College London(伦敦国王学院) University College London(伦敦大学学院)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL

Comments Accepted to EMNLP 2025 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24031 2025-10-09 cs.LG cs.AI cs.CV cs.MA 62%

GPS-MTM: Capturing Pattern of Normalcy in GPS-Trajectories with self-supervised learning

Umang Garg, Bowen Zhang, Anantajit Subrahmanya, Chandrakanth Gudavalli, BS Manjunath

机构 * ECE Department(电子工程系) University of California(加州大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

Comments 4 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05147 2025-10-08 cs.SE cs.LG stat.ML 62%

Adaptive Reinforcement Learning for Dynamic Configuration Allocation in Pre-Production Testing

Yu Zhu

机构 * University of California, Santa Cruz(加州大学圣克鲁兹分校)

专题命中 Agent评测 :agent(abstract);分类 cs.LG、cs.SE

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04786 2025-10-07 cs.LG cs.AI 62%

Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning

Jonas Hübotter, Leander Diaz-Bone, Ido Hakimi, Andreas Krause, Moritz Hardt

机构 * ETH Zürich(苏黎世联邦理工学院) Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究所)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04624 2025-10-07 cs.GT cs.AI cs.LG cs.MA econ.TH 62%

Fairness in Repeated Matching: A Maximin Perspective

Eugene Lim, Tzeh Yuan Neoh, Nicholas Teh

机构 * National University of Singapore(新加坡国立大学) Harvard University(哈佛大学) University of Oxford(牛津大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04032 2025-10-07 cs.CL cs.AI 62%

Small Language Models for Emergency Departments Decision Support: A Benchmark Study

Zirui Wang, Jiajun Wu, Braden Teitge, Jessalyn Holodinsky, Steve Drew

机构 * Department of Electrical and Software Engineering, University of Calgary, Calgary, AB, Canada(电气与软件工程系,卡尔加里大学) Department of Emergency Medicine, University of Calgary, Calgary, AB, Canada(急诊医学系,卡尔加里大学) Rockview General Hospital, Calgary, AB, Canada(罗克维尔医院)

专题命中 Agent评测 :workflow(abstract);分类 cs.AI、cs.CL

Comments Accepted to 2025 IEEE International Conference on Autonomous and Trusted Computing (ATC 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03340 2025-10-07 cs.LG cs.AI cs.CY q-bio.PE 62%

Learning Pareto-Optimal Pandemic Intervention Policies with MORL

Marian Chen, Miri Zilka

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02343 2025-10-06 cs.CL cs.AI 62%

$\texttt{BluePrint}$: A Social Media User Dataset for LLM Persona Evaluation and Training

Aurélien Bück-Kaeffer, Je Qin Chooi, Dan Zhao, Maximilian Puelma Touzel, Kellin Pelrine, Jean-François Godbout, Reihaneh Rabbany, Zachary Yang

机构 * McGill University(麦吉尔大学) Mila - Quebec Artificial Intelligence Institute(魁北克人工智能研究所) Harvard College(哈佛学院) NYU(纽约大学) Université de Montréal(蒙特利尔大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL

Comments 8 pages, 4 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02326 2025-10-06 cs.CL cs.AI 62%

Hallucination-Resistant, Domain-Specific Research Assistant with Self-Evaluation and Vector-Grounded Retrieval

Vivek Bhavsar, Joseph Ereifej, Aravanan Gurusami

机构 * CTO Office, Coherent Corporation(Coherent Corporation 技术总监办公室)

专题命中 Agent评测 :workflow(abstract);分类 cs.AI、cs.CL

Comments 21 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03206 2025-10-03 cs.CL cs.AI 62%

Enhancing Personalized Multi-Turn Dialogue with Curiosity Reward

Yanming Wan, Jiaxing Wu, Marwa Abdulhai, Lior Shani, Natasha Jaques

机构 * Google DeepMind(谷歌DeepMind) University of Washington(华盛顿大学) Google Research(谷歌研究) University of California, Berkeley(伯克利加州大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03054 2025-10-03 cs.AI cs.LG 62%

Goal Recognition Design for General Behavioral Agents using Machine Learning

Robert Kasumba, Guanghui Yu, Chien-Ju Ho, Sarah Keren, William Yeoh

机构 * Washington University in Saint Louis(华盛顿大学圣路易斯分校) Technion – Israel Institute of Technology(技术学院-以色列理工学院)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

Journal ref Transactions on Machine Learning Research, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏