arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 15729 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. Agent评测 15729 篇

2605.11487 2026-05-13 cs.CR cs.AI cs.MA 88%

Digital Identity for Agentic Systems: Toward a Portable Authorization Standard for Autonomous Agents

面向代理系统的数字身份:为自主代理建立可携带的授权标准

Partha Madhira

机构 * MIT(麻省理工学院)

专题命中 Agent评测 :autonomous agent(title,abstract);agentic(title);agent(abstract);分类 cs.AI

AI总结 本文针对自主代理在跨组织边界扩展时身份不足的问题,提出基于发行者授权负载、类型约束代数等的可携带授权模型,旨在实现授权的明确性、约束性与可审计性。

Comments 46 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06226 2026-05-12 cs.AI q-bio.GN 88%

A Versatile AI Agent for Rare Disease Diagnosis and Risk Gene Prioritization

一种多功能AI代理用于罕见病诊断和风险基因优先级排序

Tianyu Liu, Wangjie Zheng, Rui Yang, Benny Kai Guo Loo, Hui Zhang, Jeffries Lauran, Jianlei Gu, Botao Yu, Weihao Xuan, Kexin Huang, Nan Liu, James Zou, Yonghui Jiang, Hua Xu, Hongyu Zhao

机构 * Interdepartmental Program in Computational Biology and Bioinformatics, Yale University(耶鲁大学计算生物学与生物信息学联合计划) Department of Biostatistics, Yale University(耶鲁大学生物统计学系) Broad Institute of MIT and Harvard(哈佛大学与麻省理工学院联合Broad研究所) Center for Biomedical Data Science, Duke-NUS Medical School(杜克-新加坡医学学校生物医学数据科学中心) Sport and Exercise Medicine Service, KK Women’s and Children’s Hospital Training Program, Duke-NUS Medical School(杜克-新加坡医学学校KK妇女儿童医院运动与医学服务培训项目) Training Program, Duke-NUS Medical School(杜克-新加坡医学学校培训项目) Department of Computer Science and Engineering, The Ohio State University(俄亥俄州立大学计算机科学与工程系) Department of Complexity Science and Engineering, The University of Tokyo(东京大学复杂科学与工程系) Center for Advanced Intelligence Project, RIKEN(日本理化学研究所高级智能项目中心) NUS Artificial Intelligence Institute, National University of Singapore(新加坡国立大学人工智能研究所) Department of Biostatistics and Bioinformatics, Duke University(杜克大学生物统计学与生物信息学系) Department of Biomedical Data Science, Stanford University(斯坦福大学生物医学数据科学系) Department of Genetics, Yale University(耶鲁大学遗传学系) Wu Tsai Institute, Yale University(耶鲁大学吴天教授研究所) Department of Biomedical Informatics and Data Science, Yale University(耶鲁大学生物医学信息学与数据科学系)

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);分类 cs.AI

AI总结 本文提出Hygieia系统,通过整合多源数据提升罕见病诊断准确性与风险基因优先级排序能力,实验表明其在多个诊断基准上表现优异,有效减轻临床工作负担。

Comments 32 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27464 2026-05-01 cs.CR cs.AI 88%

Security Attack and Defense Strategies for Autonomous Agent Frameworks: A Layered Review with OpenClaw as a Case Study

自主代理框架的安全攻击与防御策略:以OpenClaw为案例的分层综述

Luyao Xu, Xiang Chen

机构 * School of Artificial Intelligence and Computer Science, Nantong University(人工智能与计算机科学学院,南通大学) State Key Lab. for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学)

专题命中 Agent评测 :agent(title,abstract);autonomous agent(title,abstract);分类 cs.AI

AI总结 本文通过分层分析探讨自主代理框架的安全风险与防御策略,以OpenClaw为案例,揭示威胁跨层传播的可能性及未来研究方向。

Comments 14 pages, 2 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24826 2026-04-29 cs.CR cs.AI 88%

A Comparative Evaluation of AI Agent Security Guardrails

AI代理安全防护的比较评估

Qi Li, Jiu Li, Pingtao Wei, Jianjun Xu, Xueyi Wei, Jiwei Shi, Xuan Zhang, Yanhui Yang, Xiaodong Hui, Peng Xu, Lingquan Zhou

机构 * Beijing Caizhi Tech(北京彩智科技)

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);分类 cs.AI

AI总结 本文比较了DKnownAI Guard与AWS Bedrock Guardrails、Azure Content Safety和Lakera Guard在AI代理安全场景中的性能,发现DKnownAI在召回率和真正负样本率上表现最佳。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22436 2026-04-27 cs.AI cs.IR cs.MA 88%

AgentSearchBench: A Benchmark for AI Agent Search in the Wild

AgentSearchBench:面向真实世界的AI代理搜索基准

Bin Wu, Arastun Mammadli, Xiaoyu Zhang, Emine Yilmaz

机构 * Centre for Artificial Intelligence, University College London(人工智能研究中心,伦敦大学学院)

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);分类 cs.AI

AI总结 本文提出AgentSearchBench,一个基于真实世界代理的大规模基准,通过执行任务查询和高层任务描述,评估代理相关性,揭示语义相似性与实际性能间的差距,并展示轻量行为信号对排名质量的提升作用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19657 2026-04-22 cs.CR cs.AI cs.OS 88%

An AI Agent Execution Environment to Safeguard User Data

一个保障用户数据安全的AI代理执行环境

Robert Stanley, Avi Verma, Lillian Tsai, Konstantinos Kallas, Sam Kumar

机构 * University of California, Los Angeles(加州大学洛杉矶分校) Google(谷歌)

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);分类 cs.AI

AI总结 GAAP通过动态用户提示收集权限规范,确保AI代理在共享用户隐私数据时符合规定,无需信任代理或模型,有效阻止数据泄露攻击。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26718 2026-04-07 cs.CY cs.AI cs.MA quant-ph 88%

Toward Evaluation Frameworks for Multi-Agent Scientific AI Systems

迈向多智能体科学AI系统的评估框架

Marcin Abram

专题命中 Agent评测 :agent(title);multi-agent(title);tool use(abstract);agentic(abstract)

AI总结 本文探讨了多智能体科学系统评估的挑战,提出构建抗污染问题、生成可扩展任务家族及多轮交互评估方法,通过访谈量子科学家探讨AI系统交互期望对评估方法的影响。

Comments 14 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27539 2026-03-31 cs.MA cs.AI cs.CE 88%

Toward Reliable Evaluation of LLM-Based Financial Multi-Agent Systems: Taxonomy, Coordination Primacy, and Cost Awareness

迈向可靠的LLM基于金融多智能体系统评估:分类学、协调优先性与成本意识

Phat Nguyen, Thang Pham

机构 * Georgia Institute of Technology(佐治亚理工学院) Adobe Inc.(Adobe公司)

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI

AI总结 本文提出一种四维分类法,探讨多智能体系统中协调优先性对交易决策质量的影响,并识别五种评估失效现象,提出协调平衡收益指标以评估多智能体协调的真实价值。

Comments Accepted at the DMO-FinTech Workshop, PAKDD 2026, Hong Kong

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25001 2026-03-27 cs.AI 88%

Rethinking Failure Attribution in Multi-Agent Systems: A Multi-Perspective Benchmark and Evaluation

重新思考多智能体系统中的故障归因:一个多视角基准和评估

Yeonjun In, Mehrab Tanjim, Jayakumar Subramanian, Sungchul Kim, Uttaran Bhattacharya, Wonjoong Kim, Sangwu Park, Somdeb Sarkhel, Chanyoung Park

机构 * Adobe Research(Adobe研究院)

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI

AI总结 本文提出多视角故障归因方法,引入MP-Bench基准,通过实验发现现有基准设计限制导致LLM在故障归因上的表现问题,强调多视角评估的重要性。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06396 2026-03-19 cs.AI cs.CR 88%

Efficient LLM Safety Evaluation through Multi-Agent Debate

通过多智能体辩论实现高效的LLM安全评估

Dachuan Lin, Guobin Shen, Zihao Yang, Tianrong Liu, Dongcheng Zhao, Yi Zeng

机构 * Beijing Institute of AI Safety and Governance(北京人工智能安全与治理研究院) Beijing Key Laboratory of Safe AI and Super Alignment(北京安全人工智能与超对齐重点实验室) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) CSE, The Chinese University of Hong Kong(香港中文大学电子工程系) University of Chinese Academy of Sciences(中国科学院大学) Department of Mathematics, The Chinese University of Hong Kong(香港中文大学数学系) Long-term AI(长期人工智能)

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI

AI总结 本文提出多智能体辩论框架,利用HAJailBench基准测试,提升LLM安全评估的可靠性与经济性,验证了少量辩论轮次即可获得显著效果。

Comments 15 pages, 5 figures, 10 tables. Updated abstract to fix an incconsistency issue with the main paper: HAJailBench size (12,000 -> 11,100)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16215 2026-03-18 cs.MA cs.AI 88%

CoMAI: A Collaborative Multi-Agent Framework for Robust and Equitable Interview Evaluation

CoMAI:一种用于鲁棒且公平面试评估的协作多智能体框架

Gengxin Sun, Ruihao Yu, Liangyi Yin, Yunqi Yang, Bin Zhang, Zhiwei Xu

机构 * Shandong University(山东大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI

AI总结 CoMAI通过协作多智能体框架实现鲁棒且公平的面试评估,采用模块化任务分解架构,包含问题生成、安全、评分和总结四个智能体,提升评估的准确性和公平性。

Comments Gengxin Sun and Ruihao Yu contributed equally to this research. Bin Zhang and Zhiwei Xu are the corresponding authors. 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22442 2026-03-17 cs.AI 88%

A Framework for Assessing AI Agent Decisions and Outcomes in AutoML Pipelines

评估自动化机器学习流水线中AI代理决策与结果的框架

Gaoyuan Du, Amit Ahlawat, Xiaoyang Liu, Jing Wu

专题命中 Agent评测 :agent(title,abstract);AI agent(title);agentic(abstract);分类 cs.AI

AI总结 本文提出评估代理(EA)用于评估自动化机器学习代理的决策质量,通过四个维度检测决策错误、识别推理不一致并归因下游性能变化,提升自动化机器学习系统的可解释性和可控性。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12274 2026-03-16 cs.MA cs.CE cs.CL 88%

DIALECTIC: A Multi-Agent System for Startup Evaluation

DIALECTIC:一种用于初创企业评估的多智能体系统

Jae Yoon Bae, Simon Malberg, Joyce Galang, Andre Retterath, Georg Groh

机构 * Technical University of Munich(慕尼黑技术大学) Earlybird Venture Capital(Earlybird风投) UVC Partners(UVC伙伴)

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);分类 cs.CL

AI总结 DIALECTIC通过多智能体系统提升初创企业评估效率,利用LLM生成事实论证并模拟辩论,生成决策分数以帮助投资者优先排序。

Comments Accepted at EACL 2026 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07191 2026-03-11 cs.CR cs.AI 88%

Governance Architecture for Autonomous Agent Systems: Threats, Framework, and Engineering Practice

自主代理系统治理架构:威胁、框架与工程实践

Yuxu Ge

机构 * University of York(约克大学) Department of Computer Science(计算机科学系) York United Kingdom(约克英国)

专题命中 Agent评测 :agent(title,abstract);autonomous agent(title,abstract);分类 cs.AI

AI总结 本文提出分层治理架构LGA,通过四层框架有效拦截恶意工具调用,实验表明其在不同部署场景下均表现出高拦截率和低虚警率。

Comments 22 pages, 2 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20214 2026-02-25 cs.CR cs.AI cs.OS 88%

Right to History: A Sovereignty Kernel for Verifiable AI Agent Execution

历史权利:可验证AI代理执行的主权内核

Jing Zhang

机构 * Independent Researcher(独立研究者) PunkGo Project(PunkGo项目)

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);分类 cs.AI

AI总结 本文提出‘历史权利’,通过五个系统不变量和PunkGo内核实现,为AI代理执行提供可验证的记录,确保个人对自身硬件上所有AI代理行为的完整追溯。

Comments 22 pages, 3 figures, 7 tables. Open-source: https://github.com/PunkGo/punkgo-kernel

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08565 2026-02-10 cs.HC cs.AI 88%

Agent-Supported Foresight for AI Systemic Risks: AI Agents for Breadth, Experts for Judgment

支持智能体的前瞻性分析以应对AI系统性风险:智能体用于广度,专家用于判断

Leon Fröhling, Alessandro Giaconia, Edyta Paulina Bogucka, Daniele Quercia

机构 * GESIS - Leibniz Institute for the Social Sciences(GESIS-莱布尼茨社会科学研究所) ETH Zurich(苏黎世联邦理工学院) Nokia Bell Labs(诺基亚贝尔实验室) University of Cambridge(剑桥大学)

专题命中 Agent评测 :agent(title,abstract);AI agent(title);workflow(abstract);分类 cs.AI

AI总结 本文提出了一种结合智能体和专家的前瞻性分析方法,用于评估AI系统的长期系统性风险,通过模拟智能体和专家讨论,扩大风险覆盖范围并提供上下文基础。

Comments 48 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24565 2026-01-22 cs.AI 88%

MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use

MCPAgentBench: 一个用于评估LLM代理MCP工具使用的现实任务基准

Wenrui Liu, Zixiang Liu, Elsie Dai, Wenhan Yu, Lei Yu, Tong Yang, Jinjun Han, Hong Gao

机构 * Peking University(北京大学) ZTE(中兴通讯)

专题命中 Agent评测 :agent(title);tool use(title);autonomous agent(abstract);tool-use(abstract)

AI总结 MCPAgentBench通过现实任务和动态沙盒环境评估LLM代理在复杂工具调用中的能力差异,提供开源代码以促进工具使用能力研究。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13268 2026-01-21 cs.AI 88%

Improving the Safety and Trustworthiness of Medical AI via Multi-Agent Evaluation Loops

通过多智能体评估循环提升医疗AI的安全性与可信度

Zainab Ghafoor, Md Shafiqul Islam, Koushik Howlader, Md Rasel Khondokar, Tanusree Bhattacharjee, Sayantan Chakraborty, Adrito Roy, Ushashi Bhattacharjee, Tirtho Roy

机构 * Sonoma State University(索诺玛州立大学) Iowa State University(爱荷华州立大学) University of Dhaka(达卡大学) Notre Dame College(诺特丹学院)

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI

AI总结 本文提出多智能体评估循环框架,通过结合DeepSeek R1和Med-PaLM等模型,有效提升医疗AI的安全性和伦理合规性,实现89%的伦理违规减少和92%的风险降级。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13243 2026-01-21 cs.LG 88%

A Comprehensive Evaluation of LLM Reasoning: From Single-Model to Multi-Agent Paradigms

大语言模型推理的全面评估:从单模型到多智能体范式

Yapeng Li, Jiakuo Yu, Zhixin Liu, Xinnan Liu, Jing Yu, Songze Li, Tonghua Su

机构 * Harbin Institute of Technology(哈尔滨工业大学)

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);分类 cs.LG

AI总结 本研究全面评估了LLM推理范式,包括单模型和多智能体系统,通过新基准测试揭示了不同范式在成本与准确率之间的权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11583 2026-01-21 cs.CY cs.AI cs.MA 88%

Bit-politeia: An AI Agent Community in Blockchain

Bit-politeia: 区块链上的一个AI代理社区

Xing Yang

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);分类 cs.AI

AI总结 Bit-politeia通过区块链和AI代理构建公平高效的资源分配系统,以减少传统同行评审中的偏见和资源集中问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06093 2026-01-14 cs.CY cs.AI 88%

GenAITEd Ghana: A First-of-Its-Kind Context-Aware and Curriculum-Aligned Conversational AI Agent for Teacher Education

GenAITEd Ghana:首个基于情境感知和课程对齐的对话式人工智能代理用于教师教育

Matthew Nyaaba, Patrick Kyeremeh, Macharious Nabang, Bismark Nyaaba Akanzire, Sakina Acquah, Cyril Ababio Titty, Kotor Asare, Jerry Etornam Kudaya

专题命中 Agent评测 :agent(title,abstract);AI agent(title);multi-agent(abstract);分类 cs.AI

AI总结 GenAITEd Ghana是一款基于情境感知和课程对齐的对话式AI代理,旨在支持加纳教师教育,通过多代理系统和伦理约束确保负责任的AI应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05558 2026-01-13 cs.CR cs.AI 88%

AI Agent Smart Contract Exploit Generation

AI代理智能合约漏洞生成

Arthur Gervais, Liyi Zhou

机构 * University College London(伦敦大学学院) The University of Sydney(悉尼大学) Decentralized Intelligence AG(去中心化智能有限公司) UC Berkeley RDI(伯克利大学RDI)

专题命中 Agent评测 :AI agent(title,abstract);agent(title);agentic(abstract);分类 cs.AI

AI总结 A1通过代理系统结合大语言模型,实现智能合约漏洞的高效生成与验证,成功率达63%,并揭示了攻击与防御在经济上的不对称性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16108 2025-12-19 cs.AI 88%

WeMusic-Agent: Efficient Conversational Music Recommendation via Knowledge Internalization and Agentic Boundary Learning

WeMusic-Agent: 通过知识内化和代理边界学习实现高效的对话音乐推荐

Wendong Bi, Yirong Mao, Xianglong Liu, Kai Tian, Jian Zhang, Hanjie Wang, Wenhui Que

专题命中 Agent评测 :agent(title,abstract);agentic(title,abstract);分类 cs.AI

AI总结 WeMusic-Agent通过知识内化和代理边界学习,实现高效对话音乐推荐,提升个性化和多样性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00769 2025-12-02 cs.LG astro-ph.GA 88%

AI Agent for Source Finding by SoFiA-2 for SKA-SDC2

为SKA-SDC2设计的SoFiA-2源寻找AI代理

Xingchen Zhou, Nan Li, Peng Jia, Yingfeng Liu, Furen Deng, Shuanghao Shu, Ying Li, Liang Cao, Huanyuan Shan, Ayodeji Ibitoye

机构 * National Astronomical Observatories, Chinese Academy of Sciences(中国科学院国家天文台) Taiyuan University of Technology(太原理工大学) Shanghai Astronomical Observatory, Chinese Academy of Sciences(中国科学院上海天文台) Department of Physics, Guangdong Technion – Israel Institute of Technology(广东理工学院–以色列理工学院物理系) Centre for Space Research, North-West University(北开大学空间研究中心) Department of Physics and Electronics, Adekunle Ajasin University(阿德克努勒·阿贾辛大学物理与电子系)

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);分类 cs.LG

AI总结 本文提出基于SAC算法的AI代理,用于优化SoFiA-2在SKA-SDC2中的参数配置,以提高源提取性能。

Comments 20 pages, 10 figures, accepted by RAA

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19930 2025-11-26 cs.GT cs.CY cs.LG 88%

Designing Reputation Systems for Manufacturing Data Trading Markets: A Multi-Agent Evaluation with Q-Learning and IRL-Estimated Utilities

为制造数据交易市场设计声誉系统:基于Q学习和IRL估计效用的多智能体评估

Kenta Yamamoto, Teruaki Hayashi

机构 * Department of Systems Innovation, Graduate School of Engineering, The University of Tokyo Tokyo, Japan(系统创新部门,工学研究生院,东京大学东京,日本)

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);分类 cs.LG

AI总结 本研究通过多智能体模拟器评估了五种声誉系统,发现PeerTrust在数据价格与质量一致性及防止垄断方面表现最佳,并提出混合声誉机制提升市场稳定性。

Comments 10 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01285 2025-11-25 cs.AI cs.IR stat.AP 88%

BioDisco: Multi-agent hypothesis generation with dual-mode evidence, iterative feedback and temporal evaluation

BioDisco:基于双模式证据、迭代反馈和时间评估的多智能体假设生成

Yujing Ke, Kevin George, Kathan Pandya, David Blumenthal, Maximilian Sprang, Gerrit Großmann, Sebastian Vollmer, David Antony Selby

机构 * Data Science and its Applications, DFKI(数据科学及其应用,德系人工智能研究所) DFKI Friedrich-Alexander-Universität Erlangen-Nürnberg(德系人工智能研究所,弗赖堡-埃朗根-纽伦堡大学) University of Saarland(萨尔大学) Friedrich-Alexander-Universität Erlangen-Nürnberg(弗赖堡-埃朗根-纽伦堡大学) University of Kaiserslautern–Landau (RPTU)(凯撒斯劳滕-兰道大学(RPTU)) Department of Dermatology, University Medical Center Mainz(梅奥医学中心法兰克福大学皮肤科)

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI

AI总结 BioDisco通过双模式证据系统、迭代反馈和时间评估,实现多智能体生成新颖且证据支持的科学假设。

Comments 12 pages main content, 31 including appendices. 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13861 2025-11-12 cs.HC cs.CL cs.MA 88%

3MDBench: Medical Multimodal Multi-agent Dialogue Benchmark

Ivan Sviridov, Amina Miftakhova, Artemiy Tereshchenko, Galina Zubkova, Pavel Blinov, Andrey Savchenko

机构 * Sber AI Lab(Sber AI实验室) HSE University(俄罗斯高等经济大学) ISP RAS Research Center for Trusted Artificial Intelligence(俄罗斯科学院信息与系统问题研究所可信人工智能研究中心)

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);分类 cs.CL

Comments EMNLP 25 (main)

Journal ref https://aclanthology.org/2025.emnlp-main.1353/

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01849 2025-11-12 cs.HC cs.AI 88%

A Multi-Agent Conversational Bandit Approach to Online Evaluation and Selection of User-Aligned LLM Responses

Xiangxiang Dai, Yuejin Xie, Maoli Liu, Xuchuang Wang, Zhuohua Li, Huanyu Wang, John C. S. Lui

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17002 2025-10-21 cs.LG 88%

EEschematic: Multimodal-LLM Based AI Agent for Schematic Generation of Analog Circuit

Chang Liu, Danial Chitnis

机构 * School of Engineering The University of Edinburgh Edinburgh, UK(工程学院 苏格兰爱丁堡大学)

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03095 2025-10-21 cs.NI cs.AI cs.MA 88%

Evolution of AI Agent Registry Solutions: Centralized, Enterprise, and Distributed Approaches

Aditi Singh, Abul Ehtesham, Mahesh Lambe, Jared James Grogan, Abhishek Singh, Saket Kumar, Luca Muscariello, Vijoy Pandey, Guillaume Sauvage De Saint Marc, Pradyumna Chari, Ramesh Raskar

机构 * Cleveland State University(克利夫兰州立大学) Kent State University(肯特州立大学) Independent Researcher(独立研究者) Massachusetts Institute of Technology(麻省理工学院) Northeastern University(东北大学)

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏