arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 15866 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. Agent评测 15866 篇

2602.14322 2026-02-24 cs.LG cs.LO 57%

Conformal Signal Temporal Logic for Robust Reinforcement Learning Control: A Case Study

符合信号时间逻辑用于鲁棒强化学习控制:一个案例研究

Hani Beirami, M M Manjurul Islam

机构 * ICASSP 2026

专题命中 Agent评测 :agent(abstract);分类 cs.LG

AI总结 本文提出一种符合STL的屏蔽,用于增强强化学习在航空航天中的鲁棒性和安全性,通过实验验证其在复杂环境下的有效性。

Comments 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21730 2026-02-24 cs.CL 57%

ProPerSim: Developing Proactive and Personalized AI Assistants through User-Assistant Simulation

通过用户-助手模拟开发前瞻性与个性化的人工智能助手

Jiho Kim, Junseong Choi, Woosog Chay, Daeun Kyung, Yeonsu Kwon, Yohan Jo, Edward Choi

机构 * KAIST(韩国科学技术院)

专题命中 Agent评测 :agent(abstract);分类 cs.CL

AI总结 ProPerSim通过用户-助手模拟框架开发了能够主动和个性化推荐的AI助手,实验显示其在多样化的用户场景中有效提升了用户满意度。

Comments Accepted at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18971 2026-02-24 cs.AI 57%

When Do LLM Preferences Predict Downstream Behavior?

当LLM偏好预测下游行为时是什么情况?

Katarina Slama, Alexandra Souly, Dishank Bansal, Henry Davidson, Christopher Summerfield, Lennart Luettgau

机构 * UK AI Security Institute(英国人工智能安全研究所)

专题命中 Agent评测 :agentic(abstract);分类 cs.AI

AI总结 研究发现LLM的偏好能预测捐赠建议,但对任务表现的影响不一致。

Comments 31 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18947 2026-02-24 cs.AI 57%

(Perlin) Noise as AI coordinator

Perlin噪声作为AI协调器

Kaijie Xu, Clark Verbrugge

机构 * McGill University(麦吉尔大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI

AI总结 本文提出将Perlin噪声应用于大规模AI控制,通过连续噪声场实现稳定、协调的行为生成,提升游戏AI的效率与可控性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18891 2026-02-24 cs.CY cs.AI cs.HC 57%

Orchestrating LLM Agents for Scientific Research: A Pilot Study of Multiple Choice Question (MCQ) Generation and Evaluation

协调大型语言模型代理进行科学研究:多选题生成与评估的试点研究

Yuan An

机构 * College of Computing and Informatics(计算与信息学院)

专题命中 Agent评测 :workflow(abstract);分类 cs.AI

AI总结 本文研究了通过协调多个LLM代理生成和评估多选题,发现生成的题目在质量上存在与专家审核题目不完全可比的问题,同时揭示了科学研究中劳动力向AI研究操作技能转变的趋势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18716 2026-02-24 cs.RO cs.AI 57%

Temporal Action Representation Learning for Tactical Resource Control and Subsequent Maneuver Generation

战术资源控制与后续机动生成的时序动作表示学习

Hoseong Jung, Sungil Son, Daesol Cho, Jonghae Park, Changhyun Choi, H. Jin Kim

机构 * Seoul National University(首尔国立大学) Georgia Institute of Technology(佐治亚理工学院)

专题命中 Agent评测 :agent(abstract);分类 cs.AI

AI总结 TART通过时序动作表示学习框架,有效整合资源使用与机动生成,提升有限资源下的战术决策能力。

Comments ICRA 2026, 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18462 2026-02-24 cs.CY cs.AI 57%

Assessing the Reliability of Persona-Conditioned LLMs as Synthetic Survey Respondents

评估基于人设的大型语言模型作为合成调查受访者可靠性

Erika Elizabeth Taday Morocho, Lorenzo Cima, Tiziano Fagni, Marco Avvenuti, Stefano Cresci

机构 * IIT-CNR, University of Pisa(IIT-CNR与比萨大学) University of Florence(佛罗伦萨大学) University of Pisa(比萨大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI

AI总结 本文评估了基于人设的LLM作为合成调查受访者可靠性,发现多属性提示可能引入偏差,影响子群体的忠实性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17594 2026-02-20 cs.AI 57%

AI Gamestore: Scalable, Open-Ended Evaluation of Machine General Intelligence with Human Games

AI 游戏商店:通过人类游戏评估机器通用智能的可扩展性和开放性

Lance Ying, Ryan Truong, Prafull Sharma, Kaiya Ivy Zhao, Nathan Cloos, Kelsey R. Allen, Thomas L. Griffiths, Katherine M. Collins, José Hernández-Orallo, Phillip Isola, Samuel J. Gershman, Joshua B. Tenenbaum

专题命中 Agent评测 :planning(abstract);分类 cs.AI

AI总结 通过AI GameStore评估机器通用智能,利用人类游戏测试AI在游戏中的表现,发现其在复杂游戏任务中表现不佳。

Comments 29 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17217 2026-02-20 cs.AI 57%

Continual learning and refinement of causal models through dynamic predicate invention

通过动态谓词发明实现因果模型的持续学习与完善

Enrique Crespo-Fernandez, Oliver Ray, Telmo de Menezes e Silva Filho, Peter Flach

机构 * University of Bristol(布里斯托大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI

AI总结 本文提出通过动态谓词发明实现因果模型的持续学习与完善,利用元解释学习提升推理效率,显著提高样本效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16938 2026-02-20 cs.CL 57%

ConvApparel: A Benchmark Dataset and Validation Framework for User Simulators in Conversational Recommenders

ConvApparel:一种用于对话推荐系统中用户模拟器的基准数据集和验证框架

Ofer Meshi, Krisztian Balog, Sally Goldman, Avi Caciularu, Guy Tennenholtz, Jihwan Jeong, Amir Globerson, Craig Boutilier

机构 * Google(谷歌)

专题命中 Agent评测 :agent(abstract);分类 cs.CL

AI总结 ConvApparel通过双代理数据收集和反事实验证框架,评估对话推荐系统中用户模拟器的现实差距和泛化能力。

Comments EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16756 2026-02-20 cs.CR cs.SE 57%

NESSiE: The Necessary Safety Benchmark -- Identifying Errors that should not Exist

NESSiE: 必要安全基准 -- 识别不应存在的错误

Johannes Bertram, Jonas Geiping

专题命中 Agent评测 :autonomous agent(abstract);分类 cs.SE

AI总结 NESSiE提出了一种轻量级安全基准,用于检测大型语言模型中不应存在的安全故障,揭示模型在帮助与安全之间的偏差。

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12207 2026-02-19 cs.HC cs.AI cs.SI 57%

VIRENA: Virtual Arena for Research, Education, and Democratic Innovation

VIRENA:虚拟竞技场用于研究、教育和民主创新

Emma Hoes, K. Jonathan Klueser, Fabrizio Gilardi

机构 * University of Zurich(苏黎世大学)

专题命中 Agent评测 :AI agent(abstract);分类 cs.AI

AI总结 VIRENA是一个基于开源技术的无代码平台,用于在现实社交环境中进行受控实验,研究人类与AI的互动、审核干预比较及群体 deliberation。

Comments VIRENA is under active development and currently in use at the University of Zurich. This preprint will be updated as new features are released. For the latest version and to inquire about demos or pilot collaborations, contact the authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08968 2026-02-18 cs.AI 57%

stable-worldmodel-v1: Reproducible World Modeling Research and Evaluation

stable-worldmodel-v1: 可复现的世界建模研究与评估

Lucas Maes, Quentin Le Lidec, Dan Haramati, Nassim Massaudi, Damien Scieur, Yann LeCun, Randall Balestriero

机构 * Mila & Université de Montréal(Mila与蒙特利尔大学) New York University(纽约大学) Brown University(布朗大学) Samsung SAIL(三星SAIL)

专题命中 Agent评测 :planning(abstract);分类 cs.AI

AI总结 stable-worldmodel-v1 提供了一个模块化、可复现的世界模型研究生态系统,支持标准化环境和持续学习研究,并用于评估DINO-WM的零样本鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15273 2026-02-18 cs.CY cs.CL 57%

FrameRef: A Framing Dataset and Simulation Testbed for Modeling Bounded Rational Information Health

FrameRef: 一个用于建模有限理性信息健康的框架数据集和仿真测试平台

Victor De Lima, Jiqun Liu, Grace Hui Yang

机构 * University of Oklahoma(俄克拉荷马大学)

专题命中 Agent评测 :agent(abstract);分类 cs.CL

AI总结 FrameRef通过仿真框架研究有限理性信息健康的动态,提供系统性数据集和方法论,影响人类判断和信息健康轨迹。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06855 2026-02-17 cs.AI 57%

AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents

AIRS-Bench: 一个面向前沿人工智能研究科学代理的任务集

Alisia Lupidi, Bhavul Gauri, Thomas Simon Foster, Bassel Al Omari, Despoina Magka, Alberto Pepe, Alexis Audran-Reiss, Muna Aghamelu, Nicolas Baldwin, Lucia Cipolina-Kun, Jean-Christophe Gagnon-Audet, Chee Hau Leow, Sandra Lefdal, Hossam Mossalam, Abhinav Moudgil, Saba Nazir, Emanuel Tewolde, Isabel Urrego, Jordi Armengol Estape, Amar Budhiraja, Gaurav Chaurasia, Abhishek Charnalia, Derek Dunfield, Karen Hambardzumyan, Daniel Izcovich, Martin Josifoski, Ishita Mediratta, Kelvin Niu, Parth Pathak, Michael Shvartsman, Edan Toledo, Anton Protopopov, Roberta Raileanu, Alexander Miller, Tatiana Shavrina, Jakob Foerster, Yoram Bachrach

机构 * FAIR at Meta(Meta 的 FAIR 部门) University of Oxford(牛津大学) University College London(伦敦大学学院)

专题命中 Agent评测 :agentic(abstract);分类 cs.AI

AI总结 AIRS-Bench通过20个任务评估代理在科研全生命周期中的能力,发现代理在部分任务中超越人类但未达理论上限,推动自主科研发展。

Comments 49 pages, 14 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05888 2026-02-17 cs.GT cs.AI 57%

Metric Hedonic Games on the Line

度量 Hedonic 游戏在直线上

Merlin de la Haye, Pascal Lenzner, Farehe Soheil, Marcus Wunderlich

机构 * Hasso Plattner Institute(哈索·普拉特纳研究所) University of Augsburg(奥格斯堡大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI

AI总结 本文研究了基于度量距离的Hedonic游戏模型,探讨了稳定联盟结构的存在性、性质及其在无序价格和稳定性价格方面的表现。

Comments accepted at AAMAS 2026, full version

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13855 2026-02-17 cs.AI cs.IR cs.MA 57%

From Fluent to Verifiable: Claim-Level Auditability for Deep Research Agents

从流畅到可验证:深度研究代理的声明级可审计性

Razeen A Rasheed, Somnath Banerjee, Animesh Mukherjee, Rima Hazra

机构 * Indian Institute of Science(印度科学研究所)

专题命中 Agent评测 :agent(abstract);分类 cs.AI

AI总结 本文提出深度研究代理的声明级可审计性作为关键设计目标,通过溯源覆盖、正确性、矛盾透明度和审计努力来评估可审计性,强调在合成过程中持续验证以提升科学输出的可信度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02356 2026-02-17 cs.CR cs.AI 57%

Measuring Physical-World Privacy Awareness of Large Language Models: An Evaluation Benchmark

评估大型语言模型的物理世界隐私意识:一种评估基准

Xinjie Shen, Mufei Li, Pan Li

机构 * Georgia Tech(佐治亚理工学院)

专题命中 Agent评测 :agent(abstract);分类 cs.AI

AI总结 EAPrivacy评估基准揭示大型语言模型在物理世界隐私意识方面存在严重缺陷,需更稳健的物理感知对齐。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12456 2026-02-17 q-fin.TR cs.AI 57%

Reinforcement Learning-Based Market Making as a Stochastic Control on Non-Stationary Limit Order Book Dynamics

基于强化学习的市场做市作为非平稳限价单动态的随机控制

Rafael Zimmer, Oswaldo Luiz do Valle Costa

机构 * Institute of Mathematics and Computer Sciences University of São Paulo(数学与计算机科学学院 São Paulo 大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI

AI总结 本文提出基于PPO算法的强化学习做市代理,通过模拟器环境评估其在非平稳市场中的表现,验证其有效性及模拟器的实用性。

Comments 9 pages, 8 figures, 3 tables, 31 equations

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07272 2026-02-17 cs.LG 57%

A Cramér-von Mises Approach to Incentivizing Truthful Data Sharing

基于Cramér-von Mises统计的促进真实数据共享方法

Alex Clinton, Thomas Zeng, Yiding Chen, Xiaojin Zhu, Kirthevasan Kandasamy

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Cornell University(康奈尔大学)

专题命中 Agent评测 :agent(abstract);分类 cs.LG

AI总结 本文提出基于Cramér-von Mises统计的激励机制,促进真实数据共享,通过理论分析和实验验证其有效性。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02873 2026-02-17 cs.AI 57%

It's the Thought that Counts: Evaluating the Attempts of Frontier LLMs to Persuade on Harmful Topics

想法才是关键:评估前沿大语言模型在有害话题上的说服尝试

Matthew Kowal, Jasper Timm, Jean-Francois Godbout, Thomas Costello, Antonio A. Arechar, Gordon Pennycook, David Rand, Adam Gleave, Kellin Pelrine

机构 * Université de Montréal, MILA(蒙特利尔大学,MILA) Carnegie Mellon University(卡内基梅隆大学) MIT, Center for Research and Teaching in Economics(麻省理工学院,经济研究与教学中心) Cornell University, University of Regina(康奈尔大学, Regina大学) Cornell University, MIT(康奈尔大学,麻省理工学院)

专题命中 Agent评测 :agentic(abstract);分类 cs.AI

AI总结 本文提出APE基准测试,评估前沿大语言模型在有害话题上的说服意愿,揭示模型在有害情境下尝试说服的倾向及风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12763 2026-02-16 cs.HC cs.AI 57%

"Not Human, Funnier": How Machine Identity Shapes Humor Perception in Online AI Stand-up Comedy

不是人类,更有趣:机器身份如何塑造在线AI单口喜剧的幽默感知

Xuehan Huang, Canwen Wang, Yifei Hao, Daijin Yang, Ray LC

机构 * The University of Hong Kong Hong Kong, SAR China Carnegie Mellon University\ -Computer Interaction Institute Pittsburgh United States East China Normal University Shanghai China Northeastern University\ of Art, Media City University of Hong Kong\ for Narrative Spaces Hong Kong, SAR China The University of Hong Kong Carnegie Mellon University\ -Computer Interaction Institute East China Normal University City University of Hong Kong\ for Narrative Spaces

专题命中 Agent评测 :agent(abstract);分类 cs.AI

AI总结 本研究探讨了AI身份如何影响幽默感知,通过设计基于机器身份的代理,发现其在单口喜剧表演中比基线GPT代理更有趣,提出人机集成系统应明确利用AI的独特身份。

Comments 27 pages, 5 figures. Conditionally Accepted to CHI '26

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11964 2026-02-13 cs.AI 57%

Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments

Gaia2:在动态和异步环境中评估大语言模型代理的基准测试

Romain Froger, Pierre Andrews, Matteo Bettini, Amar Budhiraja, Ricardo Silveira Cabral, Virginie Do, Emilien Garreau, Jean-Baptiste Gaya, Hugo Laurençon, Maxime Lecanu, Kunal Malkan, Dheeraj Mekala, Pierre Ménard, Gerard Moreno-Torres Bertran, Ulyana Piterbarg, Mikhail Plekhanov, Mathieu Rita, Andrey Rusakov, Vladislav Vorotilov, Mengjue Wang, Ian Yu, Amine Benhalloum, Grégoire Mialon, Thomas Scialom

专题命中 Agent评测 :agent(abstract);分类 cs.AI

AI总结 Gaia2通过动态和异步环境评估大语言模型代理,揭示了推理、效率和鲁棒性之间的权衡,为代理系统的发展提供灵活的基础设施。

Comments Accepted as Oral at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13220 2026-02-13 cs.CR cs.AI 57%

MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols

MCPSecBench: 一个系统化的安全基准和测试平台用于测试模型上下文协议

Yixuan Yang, Cuifeng Gao, Daoyuan Wu, Yufan Chen, Yingjiu Li, Shuai Wang

机构 * Lingnan University(岭南大学) University of Oregon(俄勒冈大学)

专题命中 Agent评测 :AI agent(abstract);分类 cs.AI

AI总结 MCPSecBench是一个系统化的安全基准和测试平台,用于评估模型上下文协议的安全性,揭示了不同攻击面和模型中的核心漏洞及保护机制的不足。

Comments This is a technical report from Lingnan University, Hong Kong. Code is available at https://github.com/AIS2Lab/MCPSecBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11643 2026-02-13 cs.RO cs.AI cs.CV 57%

ViTaS: Visual Tactile Soft Fusion Contrastive Learning for Visuomotor Learning

ViTaS: 用于视觉-运动学习的视觉触觉软融合对比学习

Yufeng Tian, Shuiqi Cheng, Tianming Wei, Tianxing Zhou, Yuanhang Zhang, Zixian Liu, Qianwei Han, Zhecheng Yuan, Huazhe Xu

专题命中 Agent评测 :agent(abstract);分类 cs.AI

AI总结 ViTaS通过融合视觉和触觉信息,利用软融合对比学习提升视觉-运动学习的性能。

Comments Published to ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01273 2026-02-12 cs.LG 57%

Learning-based agricultural management in partially observable environments subject to climate variability

基于气候变化的农业管理学习方法

Zhaoan Wang, Shaoping Xiao, Junchao Li, Jun Wang

机构 * Department of Mechanical Engineering, Iowa Technology Institute, University of Iowa(机械工程系、爱荷华技术研究所、爱荷华大学) Department of Chemical and Biochemical Engineering, Iowa Technology Institute, University of Iowa(化学与生物化学工程系、爱荷华技术研究所、爱荷华大学)

专题命中 Agent评测 :agent(abstract);分类 cs.LG

AI总结 本研究提出基于深度强化学习和RNN的农业管理框架,通过模拟实验展示其在应对气候变化和极端天气下的适应性与有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10017 2026-02-11 cs.CL 57%

SCORE: Specificity, Context Utilization, Robustness, and Relevance for Reference-Free LLM Evaluation

SCORE:特定性、上下文利用、鲁棒性与相关性用于无参考LLM评估

Homaira Huda Shomee, Rochana Chaturvedi, Yangxinyu Xie, Tanwi Mallick

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) Argonne National Laboratory(阿贡国家实验室) University of Pennsylvania(宾夕法尼亚大学)

专题命中 Agent评测 :planning(abstract);分类 cs.CL

AI总结 本文提出了一种无参考评估框架,用于评估LLM在高风险领域任务中的特定性、鲁棒性、相关性和上下文利用,通过精心编纂的数据集和人工评估验证了多指标评估的必要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09757 2026-02-11 cs.LG 57%

Towards Poisoning Robustness Certification for Natural Language Generation

面向自然语言生成的中毒鲁棒性认证

Mihnea Ghitu, Matthew Wicker

机构 * Department of Computing, Imperial College London, London, UK(计算系,伦敦帝国学院,伦敦,英国)

专题命中 Agent评测 :agent(abstract);分类 cs.LG

AI总结 本文提出TPA算法,通过计算最小中毒预算认证自然语言生成的有效性和稳定性,为安全关键应用提供可证明的鲁棒性保障。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08872 2026-02-10 cs.CL cs.IR 57%

Large Language Models for Geolocation Extraction in Humanitarian Crisis Response

大型语言模型在人道主义危机响应中的地理信息提取

G. Cafferata, T. Demarco, K. Kalimeri, Y. Mejova, M. G. Beiró

机构 * Universidad de San Andrés Victoria(圣安德烈斯大学)

专题命中 Agent评测 :agent(abstract);分类 cs.CL

AI总结 本文提出基于LLM的两步框架,通过结合命名实体识别和上下文地理编码模块,提升人道主义文档中地理信息提取的精度与公平性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08675 2026-02-10 cs.NI cs.AI 57%

6G-Bench: An Open Benchmark for Semantic Communication and Network-Level Reasoning with Foundation Models in AI-Native 6G Networks

6G-Bench:面向AI原生6G网络的语义通信与网络级推理开放基准

Mohamed Amine Ferrag, Abderrahmane Lakas, Merouane Debbah

机构 * Department of Computer and Network Engineering, United Arab Emirates University, UAE(计算机与网络工程系,阿拉伯联合酋长国大学) G Research Center (6GRC), Khalifa University, UAE(6G研究中心(6GRC),哈利法大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI

AI总结 6G-Bench通过开放基准评估AI原生6G网络中的语义通信与网络级推理能力,涵盖22种基础模型,揭示了不同模型在语义推理上的显著差异。

详情

展开后加载摘要…

URL PDF HTML 收藏