arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 584 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. Agent评测 584 篇

2511.09882 2026-06-16 cs.GT cs.MA 版本更新 50%

Truth, Justice, and Secrecy: Cake Cutting Under Privacy Constraints

真理、公正与保密:隐私约束下的蛋糕分割

Yaron Salman, Tamir Tassa, Omer Lev, Roie Zivan

专题命中 Agent评测 :agent(abstract)

AI总结 本文提出首个隐私保护的蛋糕分割协议,在保证无嫉妒和策略证明性的同时,通过密码学技术保护代理的偏好隐私。

Comments This is the full version of our paper published in the Proceedings of AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24177 2026-06-15 cond-mat.mes-hall cond-mat.dis-nn cond-mat.stat-mech 版本更新 50%

Optimized control protocols for stable skyrmion creation using deep reinforcement learning

基于深度强化学习的稳定斯格明子生成优化控制协议

Ji Seok Song, Se Kwon Kim, Kyoung-Min Kim

专题命中 Agent评测 :agent(abstract)

AI总结 提出深度强化学习方法,优化动态磁场-温度路径以生成热稳定性增强的斯格明子,在Fe3GeTe2单层中实现更高成功率和更长寿命,归因于最小化耗散功使状态接近平衡分布。

Comments Supplemental Material and Supplemental Vidoes will be provided with the published manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19652 2026-06-12 cs.CV 版本更新 50%

Navigating Gigapixel Pathology Images with Large Multimodal Models

利用大型多模态模型导航千兆像素病理图像

Thomas A. Buckley, Kian R. Weihrauch, Katherine Latham, Andrew Z. Zhou, Padmini A. Manrai, Arjun K. Manrai

机构 * Department of Biomedical Informatics, Harvard Medical School(哈佛医学院生物医学信息学系) Department of Pathology, Massachusetts General Hospital(麻省总医院病理学系) Department of Pathology and Laboratory Medicine, Brown University(布朗大学病理学与实验室医学系)

专题命中 Agent评测 :agent(abstract)

AI总结 提出GIANT方法,无需训练即可让通用多模态模型自主导航WSI,通过迭代选择多放大倍数裁剪并聚合证据,在MultiPathQA基准上实现SOTA。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21561 2026-06-12 cs.CV 版本更新 50%

Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning

通过逐步偏好调优的多模态智能体迭代工具使用探索

Pengxiang Li, Zhi Gao, Bofei Zhang, Yapeng Mi, Xiaojian Ma, Chenrui Shi, Tao Yuan, Yuwei Wu, Yunde Jia, Song-Chun Zhu, Qing Li

机构 * Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science & Technology, Beijing Institute of Technology(北京智能信息科技重点实验室,计算机科学与技术学院,北京理工大学) State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI) State Key Laboratory of General Artificial Intelligence, Peking University(通用人工智能国家重点实验室,北京大学) Harbin Institute of Technology(哈尔滨工业大学) Guangdong Laboratory of Machine Perception and Intelligent Computing, Shenzhen MSU-BIT University(广东机器感知与智能计算实验室,深圳MSU-BIT大学) Department of Automation, Tsinghua University(自动化系,清华大学)

专题命中 Agent评测 :agent(abstract)

AI总结 提出SPORT方法,通过任务合成、步骤采样、步骤验证和偏好调优的迭代循环,使多模态智能体无需预收集数据即可自主探索和优化工具使用策略,在GTA和GAIA基准上分别提升6.41%和3.64%。

Comments 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05835 2026-06-11 cs.GT 版本更新 50%

Bandit Social Learning with Exploration Episodes

带有探索回合的匪徒社会学习

Kiarash Banihashem, Natalie Collina, Aleksandrs Slivkins

专题命中 Agent评测 :agent(abstract)

AI总结 研究自利代理通过多臂匪徒协议进行社会学习时,尽管个体有探索动机,但集体探索失败导致贝叶斯遗憾线性增长,表明即使存在有机探索也需要外部驱动。

Comments Appears in ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15580 2026-06-11 econ.TH cs.GT 版本更新 50%

Screening for Choice Sets

选择集的筛选

Tan Gan, Yingkai Li

专题命中 Agent评测 :agent(abstract)

AI总结 研究代理人私下知道可行行动或技术集,仅向委托人披露子集的筛选问题,通过包含序假设刻画最优机制,并应用于说服管理、行动激励和生产技术激励。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12210 2026-06-11 eess.SY cs.SY math.OC 版本更新 50%

Solvability of the Output Corridor Control Problem by Pulse-Modulated Feedback

脉冲调制反馈下输出走廊控制问题的可解性

Alexander Medvedev, Anton V. Proskurnikov

专题命中 Agent评测 :agent(abstract)

AI总结 针对具有特定结构的三阶正系统,证明脉冲调制反馈在稳态下总能解决输出走廊控制问题,并用于评估药代动力学-药效学模型的患者安全性可行性。

Comments shortened version will be presented at IFAC World Congress 2026, Busan, Korea

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13452 2026-06-11 physics.soc-ph cs.MA 版本更新 50%

Collective decision-making with higher-order interactions on $d$-uniform hypergraphs

在$d$-一致超图上的高阶交互集体决策

Thierry Njougouo, Timoteo Carletti, Elio Tuci

专题命中 Agent评测 :agent(abstract)

AI总结 研究在$d$-一致超图上基于群体交互的舆论动力学模型,通过平均场分岔分析识别两个临界阈值,揭示交互组大小和品质比决定共识稳定性,且大组规模可能导致采纳劣质选项。

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.03999 2026-06-11 math.PR cs.DC cs.DM 版本更新 50%

Consensus on Dynamic Stochastic Block Models: Fast Convergence and Phase Transitions

动态随机块模型上的共识:快速收敛与相变

Haoyu Wang, Jiaheng Wei, Zhenyuan Zhang

专题命中 Agent评测 :agent(abstract)

AI总结 研究动态随机块模型上多数规则共识的收敛性,证明马尔可夫模型中任意初始偏差导致最终获胜优势,并刻画非马尔可夫模型的相变阈值。

Comments 34 pages, 2 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07129 2026-06-10 stat.AP 版本更新 50%

Collaborative estimation and evaluation of SARS-CoV-2 variant nowcasting in the United States

美国SARS-CoV-2变异株实时预测的协作估计与评估

Isaac MacArthur, Thomas Robacker, Bren Case, Spencer J. Fox, Dylan H. Morris, Evan L. Ray, Benjamin Rogers, Becky Sweger, Natalie M. Linton, John Huddleston, Andrew Magee, Zachary Susswein, Jover Lee, Trevor Bedford, Marlin D. Figgins, Ehsan Suez, Rajath Prabhakar, Tomas Leon, Brent Siegel, Mugdha Thakur, Christopher M. Hoover, Rahil Ryder, Jesse Elder, Michael Kupperman, Ruian Ke, Emma Goldberg, Sebastian Funk, Maryclare Griffin, Nicholas G. Reich, Kaitlyn E. Johnson

专题命中 Agent评测 :planning(abstract)

AI总结 本文介绍美国SARS-CoV-2变异株实时预测中心的构建,评估五种模型和基线模型在2024-2025年流感季的表现,发现基线模型整体表现良好,测序量低的地区模型性能波动更大。

Comments 32 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09319 2026-06-10 cs.CR 版本更新 50%

Benchmarking Knowledge-Extraction Attack and Defense on Retrieval-Augmented Generation

检索增强生成的知识提取攻击与防御基准测试

Zhisheng Qi, Utkarsh Sahu, Li Ma, Haoyu Han, Ryan Rossi, Franck Dernoncourt, Mahantesh Halappanavar, Nesreen Ahmed, Yushun Dong, Yue Zhao, Yu Zhang, Yu Wang

专题命中 Agent评测 :agentic(abstract)

AI总结 提出首个针对RAG系统知识提取攻击的系统性基准,涵盖多种攻击/防御策略、检索嵌入模型、生成器及数据集,在统一框架下评估,为隐私保护RAG系统提供实用基础。

Comments 12 pages. Accepted at the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2026), Dataset and Benchmark Track, Oral Presentation

Journal ref In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 26), August 09-13, 2026, Jeju Island, Republic of Korea. ACM, New York, NY, USA, 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02720 2026-06-09 cs.CY 版本更新 50%

Cognitive Comparability and the Limits of Governance: Evaluating Authority Under Radical Capability Asymmetry

认知可比性与治理的局限:在极端能力不对称下评估权威

Tony Rost

专题命中 Agent评测 :agent(abstract)

AI总结 本文通过六维框架探讨治理中权威的评估问题,发现极端能力不对称下存在结构性失效,部分维度需新规范理论而非制度设计。

Comments 20 pages, 2 tables. Interdisciplinary paper on AI governance and political theory

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06613 2026-06-09 econ.GN q-fin.EC 版本更新 50%

Some economics of artificial superintelligence

超级人工智能的经济学

Henry A. Thompson

专题命中 Agent评测 :agent(abstract)

AI总结 本文运用经济学逻辑分析超级人工智能(ASI)的威胁,指出在竞争、全面利益和信用交易条件下,ASI可能不会完全掠夺人类,为超级智能未来提供乐观视角。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04401 2026-06-09 quant-ph hep-ph 版本更新 50%

Cross-platform hardware benchmark of style-based quantum GANs for data augmentation on superconducting and trapped-ion processors

基于风格的量子生成对抗网络在超导和离子阱处理器上的跨平台硬件基准测试

Julien Baglio

专题命中 Agent评测 :workflow(abstract)

AI总结 在超导(IBM)和离子阱(IonQ)量子处理器上,对基于风格的qGAN进行高能物理数据增强任务的基准测试,比较质量与运行时,发现IonQ质量略优但IBM速度更快。

Comments 28 pages, 11 figures, 2 tables. v2: Major revision including change of title to reflect better scope; added affiliation; matches published version

Journal ref AIP Advances 16, 065008 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏