arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

2026-02-16 至 2026-02-16 共收录 12 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. Agent评测 12 篇

2508.12685 2026-02-16 cs.CL cs.AI cs.LG 85%

ToolACE-MT: Non-Autoregressive Generation for Agentic Multi-Turn Interaction

ToolACE-MT:非自回归生成用于代理多轮交互

Xingshan Zeng, Weiwen Liu, Lingzhi Wang, Liangyou Li, Fei Mi, Yasheng Wang, Lifeng Shang, Xin Jiang, Qun Liu

机构 * Huawei Technologies Co., Ltd(华为技术有限公司) Shanghai Jiao Tong University(上海交通大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))

专题命中 Agent评测 :agentic(title,abstract);agent(abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 ToolACE-MT通过非自回归生成方法高效构建高质量多轮代理对话,解决传统自回归方法效率低下的问题。

Comments Accepted by ICLR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07978 2026-02-16 cs.AI cs.CL cs.LG 82%

VoiceAgentBench: Are Voice Assistants ready for agentic tasks?

VoiceAgentBench: 聊天助手是否准备好处理代理任务?

Dhruv Jain, Harshit Shukla, Gautam Rajeev, Ashish Kulkarni, Chandra Khatri, Shubham Agarwal

机构 * OLA Electric(OLA电讯) Krutrim AI(Krutrim人工智能)

专题命中 Agent评测 :agentic(title,abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 VoiceAgentBench评估语音模型在代理任务中的表现,发现ASR-LLM在英语任务中表现优于端到端SpeechLMs,但两者在多语言和安全评估中均存在局限。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18080 2026-02-16 cs.HC cs.AI cs.SE 81%

From Prompt to Product: A Human-Centered Benchmark of Agentic App Generation Systems

从提示到产品:一种以人为中心的代理应用生成系统基准测试

Marcos Ortiz, Justin Hill, Collin Overbay, Ingrida Semenec, Frederic Sauve-Hoover, Jim Schwoebel, Joel Shor

机构 * Quome Inc.(Quome公司) Move37 Labs(Move37实验室)

专题命中 Agent评测 :agentic(title,abstract);分类 cs.AI、cs.SE

AI总结 本文提出以人为中心的基准测试,评估提示到应用系统,发现Firebase Studio在易用性、信任和视觉吸引力等方面表现最佳。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12315 2026-02-16 cs.IR cs.AI 79%

AgenticShop: Benchmarking Agentic Product Curation for Personalized Web Shopping

AgenticShop: 评估面向个性化网络购物的智能代理产品整理

Sunghwan Kim, Ryang Heo, Yongsik Seo, Jinyoung Yeo, Dongha Lee

机构 * Department of Artificial Intelligence Yonsei University Seoul Republic of Korea(人工智能系 首尔国立庆尚大学 首尔 大韩民国) ParamitaAI Seoul Republic of Korea(ParamitaAI 首尔 大韩民国) Yonsei University(首尔国立庆尚大学) ParamitaAI

专题命中 Agent评测 :agentic(title,abstract);分类 cs.AI

AI总结 AgenticShop是首个评估智能代理系统在开放网络环境中进行个性化产品整理的基准测试,通过真实购物场景和多样用户资料,验证代理系统在复杂购物情境中的适应能力。

Comments Accepted at WWW 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12544 2026-02-16 cs.AI 74%

Scaling Web Agent Training through Automatic Data Generation and Fine-grained Evaluation

通过自动数据生成和细粒度评估扩展网络代理训练

Lajanugen Logeswaran, Jaekyeom Kim, Sungryull Sohn, Creighton Glasscock, Honglak Lee

机构 * LG AI Research(LG人工智能研究)

专题命中 Agent评测 :agent(title);分类 cs.AI

AI总结 本文提出了一种基于约束的评估框架,通过自动数据生成和细粒度评估扩展网络代理训练,提升了训练数据质量和模型性能。

Comments COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12517 2026-02-16 cs.LG cs.AI cs.MA math.OC 73%

Bench-MFG: A Benchmark Suite for Learning in Stationary Mean Field Games

Bench-MFG:用于在平稳均场博弈中学习的基准测试套件

Lorenzo Magnino, Jiacheng Shen, Matthieu Geist, Olivier Pietquin, Mathieu Laurière

机构 * University of Cambridge(剑桥大学) NYU Shanghai(纽约大学上海分校) NYU Center for Data Science(纽约大学数据科学中心) Earth Species Project(地球物种计划) NYU-ECNU Institute of Mathematical Sciences at NYU Shanghai(纽约大学上海分校数学科学研究所)

专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG

AI总结 Bench-MFG提出了一套用于评估MFG学习方法的基准测试套件,通过分类问题类型和生成随机实例,提供标准化的实验框架和评估指南。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07346 2026-02-16 hep-th 71%

Higgs Field as Architect of a Geodesically Complete Universe and Agent for New Physics in Interiors of Black Holes

希格斯场作为几何完备宇宙的架构师及黑洞内部新物理的代理

Itzhak Bars

专题命中 Agent评测 :agent(title)

AI总结 希格斯场在极端引力区域创造反引力区域,解决黑洞信息悖论并恢复电弱对称性。

Comments 7 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08543 2026-02-16 cs.CL cs.AI cs.IR 62%

GISA: A Benchmark for General Information-Seeking Assistant

GISA:通用信息检索助手的基准测试

Yutao Zhu, Xingshuo Zhang, Maosen Zhang, Jiajie Jin, Liancheng Zhang, Xiaoshuai Song, Kangzhi Zhao, Wencong Zeng, Ruiming Tang, Han Li, Ji-Rong Wen, Zhicheng Dou

机构 * Renmin University of China(中国人民大学) Kuaishou Technology(快手科技)

专题命中 Agent评测 :planning(abstract);分类 cs.AI、cs.CL

AI总结 GISA是一个针对通用信息检索助手的基准测试,包含373个人工设计的查询,旨在评估信息检索任务中深度推理与信息聚合的能力。

Comments Project repo: https://github.com/RUC-NLPIR/GISA

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03834 2026-02-16 cs.LG cs.AI cs.GT 62%

From Leiden to Pleasure Island: The Constant Potts Model for Community Detection as a Hedonic Game

从莱登到欢乐岛:常数Potts模型在社区检测中的应用作为享乐博弈

Lucas Lopes Felipe, Konstantin Avrachenkov, Daniel Sadoc Menasche

机构 * Federal University of Rio de Janeiro (UFRJ)(里约热内卢联邦大学) National Institute for Research in Digital Science and Technology (Inria)(数字科学与技术国家研究院)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

AI总结 本文提出基于享乐博弈的常数Potts模型用于社区检测,通过伪多项式时间收敛到平衡划分,提升鲁棒性和准确性。

Comments Manuscript submitted to Physica A: Statistical Mechanics and its Applications

Journal ref Felipe, L. L., Avrachenkov, K., & Menasché, D. S. (2025). From Leiden to Pleasure Island: The Constant Potts Model for community detection as a hedonic game. Physica A: Statistical Mechanics and its Applications, 130989

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12763 2026-02-16 cs.HC cs.AI 57%

"Not Human, Funnier": How Machine Identity Shapes Humor Perception in Online AI Stand-up Comedy

不是人类,更有趣:机器身份如何塑造在线AI单口喜剧的幽默感知

Xuehan Huang, Canwen Wang, Yifei Hao, Daijin Yang, Ray LC

机构 * The University of Hong Kong Hong Kong, SAR China Carnegie Mellon University\ -Computer Interaction Institute Pittsburgh United States East China Normal University Shanghai China Northeastern University\ of Art, Media City University of Hong Kong\ for Narrative Spaces Hong Kong, SAR China The University of Hong Kong Carnegie Mellon University\ -Computer Interaction Institute East China Normal University City University of Hong Kong\ for Narrative Spaces

专题命中 Agent评测 :agent(abstract);分类 cs.AI

AI总结 本研究探讨了AI身份如何影响幽默感知,通过设计基于机器身份的代理,发现其在单口喜剧表演中比基线GPT代理更有趣,提出人机集成系统应明确利用AI的独特身份。

Comments 27 pages, 5 figures. Conditionally Accepted to CHI '26

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13038 2026-02-16 physics.soc-ph 50%

Modelling human activities in a system of cities

在城市系统中建模人类活动

Guo-Shiuan Lin, Denise Hertwig, Megan McGrory, Tiancheng Ma, Stefán Thor Smith, Maider Llaguno-Munitxa, Sue Grimmond, Gabriele Manoli

专题命中 Agent评测 :agent(abstract)

AI总结 本研究利用DAVE模型模拟瑞士瓦德和日内瓦地区的人口行为和移动模式,验证了通过居民行为统计数据驱动城市系统建模的可能性,并展示了可持续性和健康指标的分析。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12847 2026-02-16 cs.AR 50%

DPUConfig: Optimizing ML Inference in FPGAs Using Reinforcement Learning

DPUConfig: 使用强化学习优化FPGA上的机器学习推理

Alexandros Patras, Spyros Lalis, Christos D. Antonopoulos, Nikolaos Bellas

专题命中 Agent评测 :agent(abstract)

AI总结 DPUConfig通过强化学习动态优化FPGA上的机器学习推理配置,提升能效和性能。

Comments 8 pages, 6 figures, to appear in the proceedings of DATE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏