arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 15756 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. Agent评测 15756 篇

2412.03920 2025-07-22 cs.CL cs.AI 62%

A Survey on Large Language Model-Based Social Agents in Game-Theoretic Scenarios

Xiachong Feng, Longxu Dou, Ella Li, Qinghao Wang, Haochuan Wang, Yu Guo, Chang Ma, Lingpeng Kong

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09032 2025-07-22 math.GT cs.AI cs.LG 62%

The unknotting number, hard unknot diagrams, and reinforcement learning

Taylor Applebaum, Sam Blackwell, Alex Davies, Thomas Edlich, András Juhász, Marc Lackenby, Nenad Tomašev, Daniel Zheng

机构 * Google DeepMind, London, UK(谷歌深思(DeepMind),伦敦,英国) Mathematical Institute, University of Oxford, Andrew Wiles Building, Radcliffe Observatory Quarter, Woodstock Road, Oxford, OX2 6GG, UK(牛津大学数学研究所)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

Comments 30 pages, 17 figures; to appear in Experimental Mathematics

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08208 2025-07-21 cs.CL cs.AI 62%

ASTRID -- An Automated and Scalable TRIaD for the Evaluation of RAG-based Clinical Question Answering Systems

Mohita Chowdhury, Yajie Vera He, Jared Joselowitz, Aisling Higham, Ernest Lim

机构 * Ufonia Limited(乌菲尼亚有限公司) University of York(约克大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL

Comments 29 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.16567 2025-07-21 cs.RO cs.AI cs.CV cs.LG 62%

DeFIX: Detecting and Fixing Failure Scenarios with Reinforcement Learning in Imitation Learning Based Autonomous Driving

Resul Dagdanov, Feyza Eksen, Halil Durmus, Ferhat Yurdakul, Nazim Kemal Ure

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

Comments 6 pages, 4 figures, 2 tables, published in IEEE International Conference on Intelligent Transportation Systems (ITSC), October 12, 2022, Macau, China

Journal ref IEEE 25th International Conference on Intelligent Transportation Systems (ITSC), 2022, pp. 4215-4220

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13275 2025-07-18 cs.CL cs.AI cs.IR 62%

Overview of the TalentCLEF 2025: Skill and Job Title Intelligence for Human Capital Management

Luis Gasco, Hermenegildo Fabregat, Laura García-Sardiña, Paula Estrella, Daniel Deniz, Alvaro Rodrigo, Rabih Zbib

机构 * Avature Machine Learning(Avature机器学习)

专题命中 Agent评测 :planning(abstract);分类 cs.AI、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19982 2025-07-17 cs.CL cs.AI 62%

TD-EVAL: Revisiting Task-Oriented Dialogue Evaluation by Combining Turn-Level Precision with Dialogue-Level Comparisons

Emre Can Acikgoz, Carl Guo, Suvodip Dey, Akul Datta, Takyoung Kim, Gokhan Tur, Dilek Hakkani-Tür

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16425 2025-07-17 eess.IV cs.AI cs.CV cs.LG 62%

Patherea: Cell Detection and Classification for the 2020s

Dejan Štepec, Maja Jerše, Snežana Đokić, Jera Jeruc, Nina Zidar, Danijel Skočaj

机构 * Faculty of Computer and Information Science, University of Ljubljana(计算机与信息科学学院,卢布尔雅那大学) Institute of Pathology, Faculty of Medicine, University of Ljubljana(病理学研究所,医学学院,卢布尔雅那大学) Institute of Oncology Ljubljana(卢布尔雅那肿瘤研究所)

专题命中 Agent评测 :workflow(abstract);分类 cs.AI、cs.LG

Comments Submitted to Medical Image Analysis

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08161 2025-07-16 cs.LG cs.AI 62%

Rethinking the Foundations for Continual Reinforcement Learning

Esraa Elelimy, David Szepesvari, Martha White, Michael Bowling

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

Journal ref RLDM, 2025. RLC,2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18099 2025-07-09 cs.AI cs.CL 62%

Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge

Swarnadeep Saha, Xian Li, Marjan Ghazvininejad, Jason Weston, Tianlu Wang

机构 * FAIR at Meta(Meta 的 FAIR)

专题命中 Agent评测 :planning(abstract);分类 cs.AI、cs.CL

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02920 2025-07-08 cs.HC cs.AI cs.LG 62%

Visual-Conversational Interface for Evidence-Based Explanation of Diabetes Risk Prediction

Reza Samimi, Aditya Bhattacharya, Lucija Gosak, Gregor Stiglic, Katrien Verbert

机构 * University of Maribor(马拉堡大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

Comments 18 pages, 5 figures, 7th ACM Conference on Conversational User Interfaces

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02910 2025-07-08 cs.LG cs.AI stat.ML 62%

Causal-Paced Deep Reinforcement Learning

Geonwoo Cho, Jaegyun Im, Doyoon Kim, Sundong Kim

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

Comments Workshop on Causal Reinforcement Learning, Reinforcement Learning Conference (RLC) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.06827 2025-07-04 cs.LG cs.AI 62%

Kernel Density Bayesian Inverse Reinforcement Learning

Aishwarya Mandyam, Didong Li, Jiayu Yao, Diana Cai, Andrew Jones, Barbara E. Engelhardt

机构 * Stanford University(斯坦福大学) University of North Carolina(北卡罗来纳大学) Columbia University(哥伦比亚大学) Princeton University(普林斯顿大学) Flatiron Institute(Flatiron研究所) Gladstone Institutes(Gladstone研究所)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09992 2025-07-01 cs.CV cs.AI cs.CL 62%

MMInA: Benchmarking Multihop Multimodal Internet Agents

Shulin Tian, Ziniu Zhang, Liangyu Chen, Ziwei Liu

机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL

Comments ACL 2025 findings. The live leaderboard is at https://mmina.cliangyu.com/

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.14547 2025-07-01 cs.LG cs.AI 62%

DSAC: Distributional Soft Actor-Critic for Risk-Sensitive Reinforcement Learning

Xiaoteng Ma, Junyao Chen, Li Xia, Jun Yang, Qianchuan Zhao, Zhengyuan Zhou

机构 * Department of Automation, Tsinghua University(清华大学自动化系) School of Engineering and Applied Science, Columbia University(哥伦比亚大学工程与应用科学学院) School of Business, Sun Yat-sen University(中山大学商学院) Stern School of Business, New York University(纽约大学斯特恩商学院)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

Comments Accecpted by Journal of Artificial Intelligence Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21976 2025-06-30 cs.LG cs.AI cs.CV cs.MA cs.RO 62%

SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model

Shuhan Tan, John Lambert, Hong Jeon, Sakshum Kulshrestha, Yijing Bai, Jing Luo, Dragomir Anguelov, Mingxing Tan, Chiyu Max Jiang

机构 * Waymo LLC(Waymo公司) UT Austin(德克萨斯大学奥斯汀分校)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21967 2025-06-30 cs.CL cs.LG 62%

More Vulnerable than You Think: On the Stability of Tool-Integrated LLM Agents

Weimin Xiong, Ke Wang, Yifan Song, Hanchao Liu, Sai Zhou, Wei Peng, Sujian Li

专题命中 Agent评测 :agent(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21552 2025-06-27 cs.CV cs.AI cs.LG cs.MM cs.RO 62%

Whole-Body Conditioned Egocentric Video Prediction

Yutong Bai, Danny Tran, Amir Bar, Yann LeCun, Trevor Darrell, Jitendra Malik

机构 * UC Berkeley (BAIR)(伯克利大学(BAIR)) FAIR, Meta(Meta 公司 FAIR 实验室) New York University(纽约大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

Comments Project Page: https://dannytran123.github.io/PEVA

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.04345 2025-06-27 cs.LG cs.AI 62%

Continual Learning as Computationally Constrained Reinforcement Learning

Saurabh Kumar, Henrik Marklund, Ashish Rao, Yifan Zhu, Hong Jun Jeon, Yueyang Liu, Benjamin Van Roy

机构 * Department of Computer Science, Stanford University(计算机科学系,斯坦福大学) Department of Electrical Engineering, Stanford University(电气工程系,斯坦福大学) Department of Management Science and Engineering, Stanford University(管理科学与工程系,斯坦福大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17929 2025-06-24 cs.LG cs.AI 62%

ASTER: Adaptive Spatio-Temporal Early Decision Model for Dynamic Resource Allocation

Shulun Chen, Wei Shao, Flora D. Salim, Hao Xue

机构 * University of New South Wales(新南威尔士大学) Data61, CSIRO(Data61,CSIRO)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

Comments ASTER: Adaptive Spatio-Temporal Early Decision Model for Dynamic Resource Allocation

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00691 2025-06-24 cs.LG cs.AI 62%

Optimizing Sensory Neurons: Nonlinear Attention Mechanisms for Accelerated Convergence in Permutation-Invariant Neural Networks for Reinforcement Learning

Junaid Muzaffar, Khubaib Ahmed, Ingo Frommholz, Zeeshan Pervez, Ahsan ul Haq

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

Comments there was an error with the figures and the algorithm, working on it to correct it, will publish with updated and correct algorithm and results

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.02572 2025-06-24 cs.CR cs.AI cs.ET cs.LG 62%

GenDFIR: Advancing Cyber Incident Timeline Analysis Through Retrieval Augmented Generation and Large Language Models

Fatma Yasmine Loumachi, Mohamed Chahine Ghanem, Mohamed Amine Ferrag

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

Comments 24 pages V5.3

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16753 2025-06-23 cs.LG cs.AI cs.RO 62%

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation

Kosuke Nakanishi, Akihiro Kubo, Yuji Yasui, Shin Ishii

机构 * Department of Information Science, Kyoto University, Kyoto, Japan(京都大学信息科学系) ATR Neural Information Processing Laboratories, Kyoto, Japan(ATR神经信息处理实验室) International Research Center for Neurointelligence, The University of Tokyo, Tokyo, Japan(神经智能国际研究中心)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

Comments ICML2025 poster, 39 pages, 6 figures, 13 tables. arXiv admin note: text overlap with arXiv:2409.00418

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16473 2025-06-23 cs.HC cs.AI cs.CL 62%

Do We Talk to Robots Like Therapists, and Do They Respond Accordingly? Language Alignment in AI Emotional Support

Sophie Chiang, Guy Laban, Hatice Gunes

机构 * Department of Computer Science Technology, University of Cambridge Cambridge UK Technology, University of Cambridge

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01437 2025-06-23 cs.LG cs.AI 62%

Eau De $Q$-Network: Adaptive Distillation of Neural Networks in Deep Reinforcement Learning

Théo Vincent, Tim Faust, Yogesh Tripathi, Jan Peters, Carlo D'Eramo

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

Comments Published at RLC 2025: https://openreview.net/forum?id=Bb84iBj4wU#discussion

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13803 2025-06-18 cs.AI cs.LG 62%

Causality in the human niche: lessons for machine learning

Richard D. Lange, Konrad P. Kording

机构 * Dept. of Computer Science, Rochester Institute of Technology(计算机科学系,罗切斯特理工学院) Cognitive Science PhD Program, Rochester Institute of Technology(认知科学博士项目,罗切斯特理工学院) Center for Vision Science, University of Rochester(视觉科学中心,罗切斯特大学) Dept. of Neurobiology, University of Pennsylvania(神经生物学系,宾夕法尼亚大学) CIFAR LMB

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

Comments 23 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19824 2025-06-17 cs.AR cs.AI cs.SE 62%

AnalogXpert: Automating Analog Topology Synthesis by Incorporating Circuit Design Expertise into Large Language Models

Haoyi Zhang, Shizhao Sun, Yibo Lin, Runsheng Wang, Jiang Bian

机构 * School of Integrated Circuits, Peking University(集成电路学院,北京大学) Institute of EDA, Peking University(EDA研究院,北京大学) Beijing Advanced Innovation Center for Integrated Circuits(北京集成电路先进创新中心) Microsoft Research Asia(微软亚洲研究院)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.SE

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15014 2025-06-17 cs.AI cs.CY cs.LG 62%

Acting for the Right Reasons: Creating Reason-Sensitive Artificial Moral Agents

Kevin Baum, Lisa Dargasz, Felix Jahn, Timo P. Gros, Verena Wolf

机构 * Neuro-Mechanistic Modeling(神经机制建模) Research Center for Artificial Intelligence (DFKI)(人工智能研究所以及德国人工智能研究中心(DFKI))

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

Comments 8 pages, 2 figures, Workshop paper accepted to FEAR24 (IFM Workshop)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10816 2025-06-16 cs.LG cs.AI 62%

BalanceBenchmark: A Survey for Multimodal Imbalance Learning

Shaoxuan Xu, Menglu Cui, Chengxiang Huang, Hongfa Wang, Di Hu

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) Shanghai University of Finance and Economics(上海财经大学) Beijing University of Posts and Telecommunications(北京邮电大学) Tencent Data Platform(腾讯数据平台) Tsinghua Shenzhen International Graduate School(清华大学深圳国际 Graduate School)

专题命中 Agent评测 :workflow(abstract);分类 cs.AI、cs.LG

Comments 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02186 2025-06-16 cs.LG cs.AI 62%

Evolution Guided Generative Flow Networks

Zarif Ikram, Ling Pan, Dianbo Liu

机构 * National University of Singapore(新加坡国立大学) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

Comments Transaction of machine learning research

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.11118 2025-06-13 cs.GT cs.AI cs.LG 62%

Incentivizing Quality Text Generation via Statistical Contracts

Eden Saig, Ohad Einav, Inbal Talgam-Cohen

机构 * Technion – Israel Institute of Technology(技术ion–以色列理工学院) Tel Aviv University(特拉维夫大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏