arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

2025-11-06 至 2025-11-06 共收录 11 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. Agent评测 11 篇

2511.03434 2025-11-06 cs.HC cs.AI cs.MA cs.NI cs.SI 90%

Inter-Agent Trust Models: A Comparative Study of Brief, Claim, Proof, Stake, Reputation and Constraint in Agentic Web Protocol Design-A2A, AP2, ERC-8004, and Beyond

Botao 'Amber' Hu, Helena Rong

专题命中 Agent评测 :agentic(title,abstract);agent(title,abstract);AI agent(abstract);分类 cs.AI

Comments Submitted to AAAI 2026 Workshop on Trust and Control in Agentic AI (TrustAgent)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19738 2025-11-06 cs.AI 84%

Misalignment Bounty: Crowdsourcing AI Agent Misbehavior

Rustem Turtayev, Natalia Fedorova, Oleg Serikov, Sergey Koldyba, Lev Avagyan, Dmitrii Volkov

专题命中 Agent评测 :agent(title);AI agent(title);分类 cs.AI

Comments Add Limitations section

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02885 2025-11-06 cs.SE cs.AI 84%

AgentSLA : Towards a Service Level Agreement for AI Agents

Gwendal Jouneaux, Jordi Cabot

机构 * Luxembourg Institute of Science and Technology(卢森堡科学与技术研究院)

专题命中 Agent评测 :AI agent(title,abstract);agent(abstract);分类 cs.AI、cs.SE

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03051 2025-11-06 cs.AI cs.IR 83%

No-Human in the Loop: Agentic Evaluation at Scale for Recommendation

Tao Zhang, Kehui Yao, Luyi Ma, Jiao Chen, Reza Yousefi Maragheh, Kai Zhao, Jianpeng Xu, Evren Korpeoglu, Sushant Kumar, Kannan Achan

机构 * Walmart Global Tech(沃尔玛全球技术)

专题命中 Agent评测 :agentic(title);agent(abstract);multi-agent(abstract);分类 cs.AI

Comments 4 page, NeurIPS 2025 Workshop: Evaluating the Evolving LLM Lifecycle

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02919 2025-11-06 cs.CL 79%

Cache Mechanism for Agent RAG Systems

Shuhang Lin, Zhencan Peng, Lingyao Li, Xiao Lin, Xi Zhu, Yongfeng Zhang

机构 * Rutgers University(新泽西州立大学) University of South Florida(佛罗里达州立大学) University of Illinois Urbana–Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 Agent评测 :agent(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20462 2025-11-06 cs.AI 79%

TAMO: Fine-Grained Root Cause Analysis via Tool-Assisted LLM Agent with Multi-Modality Observation Data in Cloud-Native Systems

Xiao Zhang, Qi Wang, Mingyi Li, Yuan Yuan, Mengbai Xiao, Fuzhen Zhuang, Dongxiao Yu

机构 * School of Computer Science and Technology, Shandong University(山东大学计算机科学与技术学院) Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院)

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17619 2025-11-06 cs.IR 78%

Human vs. Agent in Task-Oriented Conversations

Zhefan Wang, Ning Geng, Zhiqiang Guo, Weizhi Ma, Min Zhang

专题命中 Agent评测 :agent(title,abstract)

Comments SIGIR-AP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03186 2025-11-06 cs.AI 57%

Adobe Summit Concierge Evaluation with Human in the Loop

Yiru Chen, Sally Fang, Sai Sree Harsha, Dan Luo, Vaishnavi Muppala, Fei Wu, Shun Jiang, Kun Qian, Yunyao Li

机构 * Adobe Inc.(Adobe公司)

专题命中 Agent评测 :workflow(abstract);分类 cs.AI

Comments Accepted by 6th Workshop on Data Science with Human in the Loop @ VLDB 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03047 2025-11-06 cs.LG 57%

Unsupervised Evaluation of Multi-Turn Objective-Driven Interactions

Emi Soroka, Tanmay Chopra, Krish Desai, Sanjay Lall

机构 * Department of Electrical Engineering Stanford University(电气工程系 斯坦福大学) Emissary Technologies

专题命中 Agent评测 :AI agent(abstract);分类 cs.LG

Comments Under review at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26091 2025-11-06 econ.TH 50%

TEE-BFT: Pricing the Security of Data Center Execution Assurance

Alex Shamis, Matt Stephenson, Linfeng Zhou

专题命中 Agent评测 :agent(abstract)

Comments Missing section + image issue

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02942 2025-11-06 math.LO 50%

Ignorance as an excuse, formally

Ekaterina Kubyshkina, Marcio Kléos Pereira, Mattia Petrolo

专题命中 Agent评测 :agent(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏