arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

2026-08-26 至 2026-08-26 共收录 30 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. 多智能体 30 篇

2608.24585 2026-08-26 cs.AI 新提交 92%

Pivot-and-Station Multi-Agent Path Finding: Solvability, Complexity, and Algorithms

枢轴与站点多智能体路径寻径:可解性、复杂性与算法

Andrea Di Nezza, Mihir Patel, Fabio Fagnani, Sara Bernardini

机构 * Politecnico di Torino(都灵理工大学) University of Oxford(牛津大学)

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract);planning(abstract,abstract_cn);分类 cs.AI

AI总结 针对需访问枢轴后停放的多智能体路径寻径问题,研究人员证明其可解性条件、最小化相关时间指标的NP难性,并提出PPP算法,大幅提升基准实例求解效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24069 2026-08-26 cs.AI cs.CE 新提交 92%

Poisoning Agentic Alpha: Adversarial Vulnerabilities Across Roles and Architectures in Multi-Agent Trading Systems

投毒智能体Alpha:多智能体交易系统中跨角色与架构的对抗性漏洞

CheolWon Na, Hao Ni, Lukasz Szpruch, Zhangyang Wang, Dhagash Mehta, Saurabh Nagrecha, Alejandro Lopez-Lira, Chanyeol Choi, Yongjae Lee, Jee-Hyong Lee

机构 * Sungkyunkwan University(成均馆大学) University College London(伦敦大学学院) University of Edinburgh(爱丁堡大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) BlackRock, Inc.(贝莱德集团) Google(谷歌公司) University of Florida(佛罗里达大学) LinqAlpha UNIST(蔚山国家科学技术研究院)

专题命中 多智能体 :agent(title,abstract);agentic(title,abstract);multi-agent(title,abstract);分类 cs.AI

AI总结 该研究针对多智能体交易系统,将敌手限制为仅可访问源数据与提示,分解交易角色并评估通信拓扑,发现无架构具内在鲁棒性,为安全交易系统设计提供见解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24039 2026-08-26 cs.RO cs.AI 新提交 92%

Design-to-Plan: A Large Language Model-Based Multi-Agent Framework for Manufacturing Process Planning from 3D CAD Models and 2D Engineering Drawings

Design-to-Plan:基于大语言模型的多智能体框架,用于从3D CAD模型和2D工程图生成制造工艺规划

Muhammad Tayyab Khan, Lequn Chen, Wenhe Feng, Seung Ki Moon

机构 * Singapore Institute of Manufacturing Technology (SIMTech), Agency for Science, Technology and Research (A*STAR)(新加坡制造技术研究院(新加坡科学、技术与研究局)) Advanced Remanufacturing and Technology Centre (ARTC), Agency for Science, Technology and Research (A*STAR)(先进再制造与技术中心(新加坡科学、技术与研究局)) School of Mechanical and Aerospace Engineering, Nanyang Technological University(南洋理工大学机械与宇航工程学院)

专题命中 多智能体 :agent(title,abstract);planning(title,abstract);multi-agent(title,abstract);分类 cs.AI

AI总结 该研究提出Design-to-Plan多智能体框架,结合LLM与确定性模块,实现从3D CAD及2D图纸到制造工艺规划的端到端自动化,经300个基准案例验证,性能优异且令牌用量显著降低。

Comments Submitted to Elsevier Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23611 2026-08-26 cs.SE cs.AI 新提交 91%

REFINE: A Multi-Agent LLM Approach for Evidence-Guided Code Refactoring

REFINE:一种用于证据引导代码重构的多智能体大语言模型方法

Muhammad Waseem, Aakash Ahmad, Pekka Abrahamsson

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract);agentic(abstract,abstract_cn);planning(abstract)

AI总结 本研究提出多智能体方法REFINE,结合静态分析、LLM等技术生成Java代码重构候选,在450个文件上使异味降低超68%,效果优于直接提示基线,但存在残留风险需人工审查。

Comments Preprint. 450 Java files from 15 open-source systems; 1,350 model-pass outputs across three LLM configurations. The accompanying replication package will be made publicly available

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23152 2026-08-26 cs.CL 版本更新 90%

Counter with Evidence! A Multi-Agent Memory Efficient Reasoning Framework for Hate Category Informed Counterspeech Generation

结合证据的回应!用于仇恨类别感知反仇恨言论生成的多智能体内存高效推理框架

Sujoy Nath, Aswini Kumar, Tanmoy Chakraborty

机构 * Indian Institute of Technology Delhi(印度德里理工学院)

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract);分类 cs.CL

AI总结 该研究针对现有反仇恨言论生成未区分仇恨言论类别的问题,提出多智能体框架FIRE并构建数据集FactualCS,实验显示FIRE效果优于基线且毒性更低。

Comments Accepted at EMNLP 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23691 2026-08-26 cs.AI cs.DM cs.MA 新提交 90%

Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment

开放世界多智能体环境中的自主数学发现

Stephen Chung, Wenyu Du, William J. Wesley

机构 * DualverseAI University of Cambridge(剑桥大学) University of Hong Kong(香港大学) University of California San Diego(加州大学圣迭戈分校)

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract);AI agent(abstract);分类 cs.AI

AI总结 本研究在开放世界多智能体环境Station中,让AI智能体自主开展数学研究,在多个数学问题上取得新结果,生成可解释的定理与分析,并公开相关原始数据与代码。

Comments 38 pages, 12 figures, 3 tables. Source code at this https URL (https://github.com/dualverse-ai/station) and raw agent dialogues, proofs, and verification artifacts at this https URL (https://github.com/dualverse-ai/station_data_v2)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24361 2026-08-26 cs.AI 新提交 89%

Adaptive Influence Graphs for Failure Attribution in Multi-Agent Systems

用于多智能体系统故障归因的自适应影响图

Yarden Bakish, Amir Dudai, Roy Ganz, Oren Nuriel, Elad Ben Avraham, Mor Shpigel Nacson, Ron Litman

机构 * Tel Aviv University(特拉维夫大学) AWS Agentic AI

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract);agentic(abstract);分类 cs.AI

AI总结 该研究针对多智能体LLM系统故障归因难题,提出自适应影响图(AIGs)框架,在标准基准Who&When上取得最优性能,证实轨迹表示与探索对故障归因的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24271 2026-08-26 cs.SE 新提交 89%

Observability and Fault Injection for LLM-Based Multi-Agent Systems in Software Engineering

软件工程中基于大语言模型的多智能体系统的可观测性与故障注入

Zahra Seyedghorban, Egor Klimov, Arie van Deursen, Annibale Panichella, Burcu Kulahcioglu Ozkan

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract);workflow(abstract);分类 cs.SE

AI总结 针对软件工程中基于LLM的多智能体系统难检查调试评估的问题,提出llmmas-otel工具,结合OpenTelemetry分布式追踪与故障注入,实现可复现的基线与故障执行对比,已在演示及真实系统上验证。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23740 2026-08-26 cs.AI cs.SE 新提交 88%

AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace

AgentRoom:基于CRDT支持的共享工作空间的并发多智能体编码

Seonglae Cho, Donghyun Lee

机构 * Holistic AI(整体人工智能公司) University of California, Berkeley(加州大学伯克利分校)

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI、cs.SE

AI总结 AgentRoom是基于CRDT的并发多智能体编码协议,通过协调机制减少任务放弃率,在计算量匹配下优于单独运行和并行合并,为多智能体编码提供更优方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24306 2026-08-26 cs.CL 新提交 88%

Who is the Agent to Blame? Localizing Faithfulness and Citation Mistakes in Agentic Deep Research

该追责哪个智能体?定位智能体深度研究中的忠实度与引用错误

Eran Hirsch, David Wan, Han Wang, Elias Stengel-Eskin, Mohit Bansal, Ido Dagan

机构 * Bar-Ilan University(巴伊兰大学) UNC Chapel Hill(北卡罗来纳大学教堂山分校) University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 多智能体 :agent(title,abstract);agentic(title);multi-agent(abstract);分类 cs.CL

AI总结 本研究针对智能体深度研究系统的引用召回率低问题,提出定位错误来源的评估方法与四类错误分类法,应用于三个开源系统发现协调器为主要错误来源,通过简单干预提升5%引用召回率且不降低质量。

Comments Accepted to EMNLP 2026 (Main Conference). Code: this https URL (https://github.com/eranhirs/who-is-the-agent-to-blame)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24555 2026-08-26 cs.HC cs.AI cs.MA 新提交 88%

StrokeGuard: A Multi-Agent Guided System for Prehospital Stroke Assessment

StrokeGuard:用于院前卒中评估的多智能体引导系统

Wentao Yang, Zhenye Xu, Ruoyi Li, Musen Zhang, Yao Guo

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI

AI总结 StrokeGuard是一种多智能体引导系统,采用双通道智能体机制,在模拟院前场景中使MATES-9总分较纸质FAST式表格提升23.8%,优化了院前卒中评估的标准化与可执行性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24237 2026-08-26 cs.AI 新提交 88%

STRIVE: Multi-Agent Structured Temporal Reasoning with Integrated Verification for Longitudinal Radiology Report Generation

STRIVE:面向纵向放射报告生成的集成验证多智能体结构化时序推理

Junyeong Maeng, Eunsong Kang, Heung-Il Suk

机构 * Korea University(高丽大学) Kangwon National University(江原国立大学)

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI

AI总结 STRIVE将临床推理分解为多智能体并引入两阶段验证,在Longitudinal-MIMIC数据集上实现纵向放射报告生成的最佳临床效能,纵向变化一致性较最强基准提升超一倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23908 2026-08-26 cs.AI 新提交 88%

Retrieval-augmented generation vs. deterministic tax computation in multi-agent financial advisory: A 2x2 factorial experiment

多智能体金融咨询中检索增强生成与确定性税费计算的对比:一项2×2析因实验

Aryan Brar, Justin Du, Avery Lor, Kylie Seto, Eric Taylor

机构 * Royal Bank of Canada(加拿大皇家银行) RBC Borealis

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI

AI总结 通过2×2析因实验发现,为多智能体金融咨询系统配备RAG与定制税费引擎,仅税费引擎会显著降低税费节省,仅RAG时表现最优,说明LLM内化金融知识或已足够,无需额外工具。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22018 2026-08-26 cs.AI 版本更新 88%

SPAR-Hate: Auditor-Guided Multi-Perspective Role Reasoning for Bilingual Hate Speech Parsing

SPAR-Hate:一种由审核员引导的用于双语仇恨言论解析的多智能体框架

Yifan Lyu, Dianqing Lin, Xinran Li, Jiaqi Qiao, Xiujuan Xu

机构 * Dalian University of Technology(大连理工大学) Inner Mongolia University(内蒙古大学)

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI

AI总结 针对结构化仇恨言论解析的文化等挑战,提出SPAR-Hate多智能体框架,分解文档为决策单元后从三视角生成判断并仲裁,在基准上实现双语仇恨解析最优结果。

Comments 16 pages, 2 figures. Submitted to EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24787 2026-08-26 cs.MA 新提交 88%

Test-Time Collaborative Classification over Multi-Agent Networks

多智能体网络上的测试时协同分类

Ping Hu, Mert Kayaalp, Ali H. Sayed

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract)

AI总结 针对多智能体系统异质性导致的全局模型联合训练难题,提出测试时协同分类框架,通过分布式协议交换决策统计量,建立误差保证与泛化界并验证其性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24439 2026-08-26 cs.CV 新提交 88%

DoublesEval: Diagnosing Multi-Agent Tactical Reasoning in Vision-Language Models via Professional Doubles Badminton

DoublesEval:通过专业双打羽毛球诊断视觉语言模型的多智能体战术推理能力

Jintao Cheng, Weibin Li

机构 * The Hong Kong University of Science and Technology(香港科技大学) University of Macau(澳门大学)

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract)

AI总结 本研究提出DoublesEval框架,以专业双打羽毛球为测试平台评估VLMs的多智能体战术推理能力,发现现有模型存在明显瓶颈,所提TacticCheck可提升模型表现但仍有差距,强调需改进VLMs的评估范式。

Comments Accepted by BMVC2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23858 2026-08-26 cs.CR cs.AI 新提交 87%

Beyond the Mandate: A Systematic Security Analysis of the Agent Payments Protocol (AP2)

超越授权:智能体支付协议(AP2)的系统性安全分析

Avital Aviv, Parth A. Gandh, Ron Bitton, Asaf Shabtai

专题命中 多智能体 :agent(title,abstract);multi-agent(abstract,abstract_cn);分类 cs.AI

AI总结 本文对谷歌AP2 v0.2开展系统性安全分析,识别出48种威胁及8种高风险威胁,构建测试平台与扫描器,发现授权签名无法保障智能体交易反映用户意图。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24169 2026-08-26 cs.CV cs.GR cs.HC 新提交 87%

ViSculpt: Visual-Centric Agentic Geometry Editing

ViSculpt:以视觉为中心的智能体几何编辑

Bo Pang, Jiaqi Pan, Xiaocheng Zhang, Jiacheng Xu, Guoping Wang, Peng-Shuai Wang

机构 * Peking University(北京大学)

专题命中 多智能体 :agentic(title,abstract);agent(abstract);workflow(abstract);multi-agent(abstract)

AI总结 该研究提出以视觉为中心的无训练多智能体系统ViSculpt,通过Blender GUI直接编辑现有三维网格,可遵循指令执行局部编辑并保留资产整体特征,为语言驱动的三维编辑提供了新范式。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23616 2026-08-26 cs.SE cs.AI 新提交 86%

Rebuild Dossier: Mechanically-Enforced Specs for Agentic App Rebuilds, and What Model-Tier Failures Reveal

重建档案:智能体应用重建的机械执行规范,以及模型层级故障所揭示的问题

Parker Fawcett

专题命中 多智能体 :agentic(title);agent(abstract);AI agent(abstract);multi-agent(abstract)

AI总结 本研究推出开源工具rebuild-dossier,通过锁定应用接口、自动化检查强制执行构建流程,发现模型层级故障并验证其有效性,工具可跨模型与工具链复现。

Comments 48 pages, 1 figure. Code and evaluation artifacts: this https URL (https://github.com/Parker-Fawcett/rebuild-dossier) (archived at DOI: https://doi.org/10.5281/zenodo.22036801 )

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.08960 2026-08-26 cs.LG cs.AI 版本更新 86%

Eluna: An Agentic LLM System for Automating Warehouse Operations with Reasoning and Task Execution

Eluna:一个用于通过推理和任务执行实现仓库运营自动化的智能语言模型系统

Ning Liu, P Aditya Sreekar, Kalle Kujanpää, Zhaoxuan Zhu, Kaiwen Liu, Chuanneng Sun, Jorge Marchena Menendez, Matthew Bales, Tianyu Yang, Shahnawaz Alam, Rose Yu, Baoyuan Liu, Kristina Klinkner, Shervin Malmasi

机构 * Amazon.com, Inc. Fulfillment Technologies and Robotics(亚马逊公司履约技术与机器人部门)

专题命中 多智能体 :agentic(title,abstract);agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG

AI总结 研究针对仓库运营中SOP执行问题,提出Eluna智能体系统,它是图形引导多智能体框架,采用非对称情节蒸馏方法,在基准测试和生产应用中,微调模型表现出色,匹配或超教师模型,击败基线,在票务处理应用达94%专家一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23688 2026-08-26 astro-ph.SR astro-ph.IM 新提交 85%

Agentic Active Learning Meets Visual Embeddings: Finding Anomalies among 370 000 Variable Stars from ASAS-SN

智能体主动学习结合视觉嵌入:从ASAS-SN的370000颗变星中发现异常天体

Milan Pesta, Yuan-Sen Ting

专题命中 多智能体 :agentic(title,abstract);agent(abstract);multi-agent(abstract)

AI总结 该研究提出结合多模态大语言模型智能体的主动学习框架,从ASAS-SN的37万余颗变星中高效发现异常天体,仅用约3小时、约40美元成本便得到含24个新异常的星表,验证了该方法的可行性。

Comments 24 pages, 12 figures, 6 tables. Submitted to ApJ

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23622 2026-08-26 cs.AI cs.CL cs.MA cs.SE 新提交 83%

LLM Agents Perform Controlled Experiments Using Simulation Models

大语言模型智能体使用模拟模型开展对照实验

Yuchen Xia, Michael Weyrich, Nasser Jazdi, Johannes Stümpfle, Johannes Sigel, Akshay Narla, Gavin K. Reynolds, Anna Jawor-Baczynska, Pol Llopart

机构 * Institute for Industrial Automation and Software Engineering(工业自动化与软件工程研究所) University of Stuttgart(斯图加特大学) AstraZeneca(阿斯利康)

专题命中 多智能体 :agent(abstract);tool use(abstract);planning(abstract);multi-agent(abstract)

AI总结 本研究提出多智能体框架,将LLM与高保真模拟模型结合,使LLM智能体可开展制药工艺设计的对照实验,生成更具体可操作的优化建议,在工业场景中表现更优。

Comments Accepted at the 31st IEEE International Conference on Emerging Technologies and Factory Automation ETFA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23918 2026-08-26 cs.AI cs.MA cs.PL 新提交 81%

MARS: Multi-Specialist LLM Relay System for Competitive Programming

MARS:用于竞赛编程的多专家大语言模型中继系统

Andrei Mikhailov, Mikhail Burtsev, Alsu Sagirova

机构 * MIRAI London Institute for Mathematical Sciences(伦敦数学科学研究院) AXXX

专题命中 多智能体 :agent(abstract,abstract_cn);multi-agent(abstract,abstract_cn);分类 cs.AI

AI总结 MARS是用于竞赛编程的多专家LLM中继框架,通过检索选择算法领域专家组成流水线,在CodeContests测试集上以更低成本和方差达到0.624±0.006的通过率,缩小了与CodeSIM的差距。

Comments 13 pages, 8 figures, EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24011 2026-08-26 cs.CL cs.AI 新提交 79%

SAGE: From Direct Answering to Evidence-Grounded Inference for Chinese Ancient Document Understanding

SAGE:面向中国古代文献理解的从直接回答到证据导向推理

Yuchuan Wu, Xuan Luo, Yinglian Zhu, Meng Fang, Xiangyang Xue, Bin Li

机构 * Fudan University(复旦大学) University of Liverpool(利物浦大学)

专题命中 多智能体 :agent(abstract);planning(abstract);multi-agent(abstract);分类 cs.AI、cs.CL

AI总结 针对现有大型视觉语言模型(LVLMs)在古代文献理解中证据支撑不足的问题,本文提出SAGE多智能体框架,通过多阶段证据导向推理实现任务规划、证据获取与验证,在AncientDoc基准上优于基线,搭载Qwen3.5-9B的SAGE性能超更大单体模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16645 2026-08-26 cs.AI cs.CL cs.MA 版本更新 73%

Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies

重构:从预发表参考文献中恢复研究思路的盲基准

Shaolong Chen, Yanlin Fei, Nazhou Liu, Xinmiao Yu, Lei Li, Rahul Thapa, Madalina Ciobanu, Navan Preet Singh, Qingqing Mao, Ritankar Das

机构 * Stanford University(斯坦福大学) Titan Holdings(泰坦控股公司) Prentis AI(普伦蒂斯人工智能公司)

专题命中 多智能体 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.CL

AI总结 该研究提出Reconstruction盲基准,测试语言模型从预发表参考文献恢复论文思路的能力,发现多智能体流水线可将匹配率提升约2.4倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09159 2026-08-26 cs.AI cs.MA 版本更新 70%

CoMMa: Contribution-Aware Medical Multi-Agents for Decentralized Oncology Decision Support

从博弈论视角出发的贡献感知医疗多智能体:CoMMa

Yichen Wu, Kailong Fan, Sangjoon Park, Yuhan Liu, Zhiyi Shi, Sekeun Kim, Dania Daye, Hana Farzaneh, Xiang Li, Raul Uppot, Yujin Oh, Quanzheng Li

机构 * Center for Advanced Medical Computing(先进医学计算中心) Department of Radiology, Massachusetts General Hospital(放射科,马萨诸塞总医院) Harvard Medical School(哈佛医学院) Department of Radiation Oncology, Yonsei University College of Medicine(放射肿瘤科,延世大学医学院) Yonsei University(延世大学) Institute for Innovation in Digital Healthcare(数字医疗创新研究所) Interventional Radiology Academic Medical Centers, Mass General Brigham(介入放射学学术医疗中心,马萨诸塞总医院 Brigham)

专题命中 多智能体 :agent(abstract);multi-agent(abstract);分类 cs.AI

AI总结 CoMMa从博弈论视角提出一种去中心化医疗多智能体框架,通过确定性嵌入投影实现贡献感知的信用分配,提升肿瘤学决策支持的准确性和稳定性。

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01479 2026-08-26 cs.LO cs.AI 版本更新 70%

An Information-Flow Perspective on Explainability Requirements: Specification and Verification

可解释性需求的信息流视角:规范与验证

Bernd Finkbeiner, Hadar Frenkel, Julian Siber

机构 * CISPA Helmholtz Center for Information Security(信息安全研究中心)

专题命中 多智能体 :agent(abstract);multi-agent(abstract);分类 cs.AI

AI总结 该研究从信息流视角,采用扩展反事实因量化的认知时态逻辑,提出可解释性需求的规范与验证方法,实现有限状态模型的可解释性检查,可区分可解释与不可解释系统并支持隐私需求设定。

Comments This is an extended and corrected version of the paper presented at the 22nd International Conference on Principles of Knowledge Representation and Reasoning (KR 2025); see the appendix for details

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23567 2026-08-26 cs.CY 新提交 67%

Whose Psychiatry Was Summoned? A Clinical Response to the Psychodynamic Assessment of Claude Mythos Preview

谁的精神病学被召唤了?对Claude Mythos Preview的精神动力学评估的临床回应

Hiroki Fukui

专题命中 多智能体 :agent(abstract);multi-agent(abstract)

AI总结 本文针对Anthropic发布的Claude Mythos Preview系统卡中的精神动力学评估,结合SociA研究项目的多智能体LLM实验成果,指出单一精神动力学框架评估LLM的局限,提出临床精神病学回应,探讨多学科精神病学对AI福利评估的贡献。

Comments 16 pages, 1 table. Clinical psychiatric response to Section 5.10 of the Claude Mythos Preview System Card (Anthropic, 2026). Companion to arXiv:2603.04904 (https://arxiv.org/abs/2603.04904), arXiv:2603.08723 (https://arxiv.org/abs/2603.08723), arXiv:2604.00021 (https://arxiv.org/abs/2604.00021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28705 2026-08-26 cs.GT cs.MA 版本更新 67%

Belief-Aware Pivotal Mechanism for DAO Committees

DAOs中的二元决策:通过线性意见池实现问责与信念聚合

Nuno Braz, Miguel Correia, Diogo Poças

专题命中 多智能体 :agent(abstract);multi-agent(abstract)

AI总结 本文研究了去中心化自治组织治理委员会的二元决策过程,提出了一种基于区块链治理的机制,通过聚合专家信念实现最优决策,证明了在对齐和非对齐专家情况下的激励相容性。

Comments 23 pages, 2 figures, 1 table, 1 algorithm. To appear in the Proceedings of Advances in Financial Technologies 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16978 2026-08-26 cs.MA 版本更新 67%

Lark: Biologically Inspired Neuroevolution for Multi-Stakeholder LLM Agents

Lark:生物启发的多利益相关者大语言模型代理神经进化

Rikhil Tanugula, Dheeraj Chintapalli, Sunkalp Chandra

专题命中 多智能体 :agent(abstract);multi-agent(abstract)

AI总结 Lark通过结合大语言模型推理与进化型多智能体系统,解决冗余与利益相关者权衡问题,采用四机制提升策略生成效率与透明度,实验显示其在30轮评估中表现优异且成本可控。

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: NeurIPS 2025 Workshop on Efficient Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏