arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 3048 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 3048 篇

2607.08768 2026-07-10 cs.CL 新提交 70%

UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks

UniClawBench:面向实际任务中主动智能体的通用基准测试

Zhekai Chen, Chengqi Duan, Kaiyue Sun, Bohao Li, Yuqing Wang, Manyuan Zhang, Xihui Liu

机构 * HKU MMLab(香港大学多媒体实验室) Meituan(美团)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 针对现有基准测试难以评估主动智能体的问题,提出UniClawBench,基于五项基础模型能力设计400个双语现实任务,采用实时评估和闭环评估策略,通过多模型框架对比展示基础模型能力和框架设计对性能的影响,还公开了基准测试和代码。

Comments Project Page: https://uniclawbench.github.io | GitHub Repo: https://github.com/HKU-MMLab/UniClawBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07824 2026-07-10 cs.MA cs.AI 新提交 70%

From Triggers to Emotions: A CPM-Grounded Appraisal Multi-Agent for Dynamic Emotional Evolution in Persona-Based Dialogue

从触发因素到情感:基于CPM的评估多智能体模型用于基于角色的对话中的动态情感演变

Jingyao Cai, Shuaijun Liu, Abdul Rehman, Yutong Guo, Qin Tian, Thomas Dolby, Sue Green, Chantel Cox, Xiaosong Yang

机构 * National Centre for Computer Animation(国家计算机动画中心) Information Hub(信息中心) Key Laboratory of Child Cognition & Behavior Development of Hainan Province(海南省儿童认知与行为发展重点实验室) Chief Technology Officer(首席技术官) School of Health and Care(健康与护理学院)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 针对基于角色对话中情感模拟的局限,借鉴CPM理论提出CPM - MultiAgent框架,通过情感触发提取、协作评估和状态更新实现情感一致的角色模拟,经多种实验验证能有效模拟动态情感演变。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06413 2026-07-08 stat.ME cs.AI 新提交 70%

An Experimental Design Approach to Evaluating Agentic AI's Autonomous Model Discovery

一种评估智能体人工智能自主模型发现的实验设计方法

Hao He, Xueying Liu, Chris J. Kuhlman, Xinwei Deng

机构 * Department of Statistics, Virginia Tech(统计学系,弗吉尼亚理工学院) Department of Statistical Science, Baylor University(统计科学系,贝勒大学) Advanced Research Computing, Virginia Tech(高级研究计算,弗吉尼亚理工学院)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 研究大型语言模型编码智能体自主模型发现行为,提出实验设计与分析框架,将智能体视为随机模型发现算子,在多种受控因素下研究Codex和Claude Code两个算子,进行回归模型和推理,开发规范分解,通过网络造词游戏测试平台得出相关深刻发现。

Comments 39 pages, 11 figures, 6 tables. Data and code available at the GitHub repository listed in the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06196 2026-07-08 cs.CL cs.CY 新提交 70%

Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability

Pluralis v0.1:迈向用于人工智能风险与可靠性的多元文化、多模态、多语言基准测试

Alicia Parrish, Rajat Shinde, Sanket Badhe, Xinyi Bai, Sree Bhargavi Balija, Hua-Rong Chu, Emilio Ferrara, Armstrong Foundjem, Rajat Ghosh, Aakash Gupta, Xuanli He, Ong Chen Hui, Minji Jung, Madhangi Karimanal, Faiza Khan Khattak, Boryoung Kim, Eugenia Kim, Liliya Lavitas, Seok Min Lim, Victor Lu, Jim Moirangthem, Dhivya Nagasubramanian, Deepak Pandita, Sita Rajagopal, Geetha Raju, Evgeniia Razumovskaia, Aravind Reddy, Federico Ricciuti, Nobin Sarwar, Sungpil Shin, Sunayana Sitaram, Snehal Thorat, Tharindu Cyril Weerasooriya, Jasmijn Bastings, Joachim Baumann, Kongtao Chen, Murali Emani, Mariya Hendriksen, Jiho Jin, Jun Seong Kim, Younghoon Ko, Alicja Kwasniewska, Minjae Lee, Tom Wei-cyuan Lin Kashyap Ramanandula Manjusha, Junho Myung, Junyeong Park, Roma Patel, Shyam Ratan, Sudarsun Santhiappan, Priyanka Suresh, Tuesday, Ksheeraj Sai Vepuri Laura Amortegui-Ordonez, Claire Dennis, Minsuk Kahng, Chris Knotz, Alice Oh, Balaraman Ravindran, Soojung Ryu William Bartholomew, Hiwot Tesfaye, Lora Aroyo

机构 * Google DeepMind(谷歌DeepMind) University of Alabama in Huntsville(阿拉巴马大学亨茨维尔分校) Google(谷歌) University of Missouri Columbia(密苏里大学哥伦比亚分校) Chunghwa Telecom Laboratories(春木电信实验室) University of Southern California(南加州大学) Polytechnique Montreal(蒙特利尔理工学院) Nutanix ThinkEvolve Labs(ThinkEvolve实验室) UCL(伦敦大学学院) Infocomm Media Development Authority(信息通信媒体发展局) Monark Health(Monark健康) Seoul National University(首尔国立大学) Microsoft(微软) Centre for Responsible AI (CeRAI), Wadhwani School of Data Science and AI (WSAI), Indian Institute of Technology Madras(负责任人工智能中心(CeRAI)、瓦达威人工智能学校(WSAI)、印度理工学院马德拉斯分校) University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校) Microsoft Research India(微软印度研究院) Stanford University(斯坦福大学) Argonne National Laboratory(阿贡国家实验室) University of Oxford(牛津大学) KAIST(韩国科学技术院) Yonsei University(延世大学) Amazon(亚马逊) UIUC(伊利诺伊大学香槟分校) Rochester Institute of Technology(罗切斯特理工学院) Xenoscube Inc.(Xenoscube公司) Korea AI Safety Institute (K-AISI)(韩国人工智能安全研究所(K-AISI)) MLCommons CommonGround Artifex Labs(Artifex实验室)

专题命中 评测与基准 :LLM(abstract);language model(abstract);分类 cs.CL

AI总结 研究针对现有AI安全评估框架忽视文化差异问题,构建多语言多模态数据集Pluralis v0.1,从文化优先视角引入新评估范式,提出Judge - Pluralis集成,揭示特定地区失败模式,为多语言多元文化评估提供基础和创新起点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05773 2026-07-08 cs.AI 新提交 70%

Beyond Static Evaluation: Building Simulation Environments for Scalable Agentic Reinforcement Learning

超越静态评估:构建用于可扩展智能体强化学习的模拟环境

Akshay Arora, Ishan Nigam, Ashutosh Aggarwal, Shefali Bansal, Krishna Singh, Sweta Kumari, Nikhil Mittal, Shariq Farhan, Siddarth Malreddy

机构 * Uber AI Solutions(优步人工智能解决方案)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 研究针对大语言模型成为自主智能体后传统静态评估失效的问题,引入AgenticAI-Supervisor平台,通过API和UI驱动构建强化学习环境,运用可验证执行结果、多维奖励塑造及严格验证测试,以客户支持智能体案例展示核心能力及未来工作重点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05638 2026-07-08 cs.SE cs.AI 新提交 70%

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems

EvalLoop:一种用于商业人工智能系统评估驱动迭代改进的方法

Kenneth Benavides, Josh Fleischer, Danti Chen

机构 * Robert Half(罗伯特·哈夫公司)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 研究针对商业AI系统评估忽视迭代改进价值的问题,提出EvalLoop方法,通过维度指标分组、故障模式分类和结构化迭代工作流程,实现评估驱动的迭代改进,经案例研究验证有效,还能助力模型选择与部署权衡,且被打包为可重用工件。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04727 2026-07-07 cs.SE cs.AI cs.CV 新提交 70%

Dashboard2Code: Evaluating Multimodal Models on Reconstructing Interactive Dashboards

Dashboard2Code:在重建交互式仪表板上评估多模态模型

Tianhao Niu, Ziyu Han, Qiguang Chen, Shiqi Zhou, Baocai Shan, Hengjie Fang, Qingfu Zhu, Wanxiang Che

机构 * Research Center for Social Computing and Interactive Robotics(社会计算与交互机器人研究院)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 研究多模态模型在交互式仪表板重建任务,介绍Dashboard2Code任务,提出DashboardMimic基准及自动化评估框架,实验发现开源与闭源模型在此任务上有差距。

Comments Accepted to ACL2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04436 2026-07-07 cs.SE cs.AI 新提交 70%

A Retrieval-Augmented Framework for Detecting and Resolving Pragmatic Ambiguities in Natural Language Requirements

一种用于检测和解决自然语言需求中语用歧义的检索增强框架

Pavithra PM Nair, Preethu Rose Anish

机构 * Tata Consultancy Services(塔塔咨询公司)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 研究自然语言需求中的语用歧义,利用检索增强技术和不同领域知识库模拟不同专业利益相关者,生成候选消歧需求并验证,评估该方法在需求文档上的效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04008 2026-07-07 cs.CL cs.IR 新提交 70%

Candidate-Constrained Retrieval-Augmented Generation for LongEval-RAG: System Design and Empirical Analysis

用于LongEval-RAG的候选约束检索增强生成:系统设计与实证分析

Yingdong Yang, Haijian Wu

专题命中 评测与基准 :LLM(abstract,abstract_cn);分类 cs.CL

AI总结 介绍用于LongEval-RAG的候选约束检索增强生成系统,结合多种方法,经评估得出最强平衡变体rule-minilm,表明主要增益来自稳定规则证据单元与句子级神经选择结合,强调多指标评估需求。

Comments Published in CEUR Workshop Proceedings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03887 2026-07-07 cs.CR cs.AI cs.SE 新提交 70%

Advanced Topic Modeling Techniques for Categorizing Software Vulnerabilities

用于软件漏洞分类的高级主题建模技术

Utkarsh Tiwari, Spoorthi M, Anirudh S, Nidhin Prabhakar T.

机构 * Department of Computer Science & Engineering, Amrita School of Computing, Bengaluru, Amrita Vishwa Vidyapeetham, India(计算机科学与工程系,Amrita计算学院,班加罗尔,Amrita世界大学,印度)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 研究利用大语言模型驱动的主题建模技术,从软件漏洞数据集中提取有意义见解,采用多种模型及降维和聚类方法,增强网络安全威胁优先级排序与决策,支持漏洞管理的可扩展自动化解决方案。

Comments 10 pages, 10 figures. Accepted at the 16th International Conference on Computing, Communication and Networking Technologies (ICCCNT 2025), July 6-11, 2025, IIT Indore, Madhya Pradesh, India. IEEE proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02897 2026-07-07 cs.CR cs.AI cs.CV 新提交 70%

PPE-Bench: A Benchmark for Evaluating MLLM Unlearning under Private-Public Entanglement

PPE-Bench:用于评估公私纠缠下MLLM遗忘能力的基准测试

Xianren Zhang, Delvin Ce Zhang, Dongwon Lee, Suhang Wang

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) University of Sheffield(谢菲尔德大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 研究针对现有MLLM遗忘基准测试的局限性,提出PPE-Bench基准测试,包含公私纠缠图像,引入两种方法在遗忘中保护公共信息,实验发现现有方法能减少隐私泄露,但常损害相邻公共信息。

Comments 17 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01425 2026-07-07 cs.AI 新提交 70%

Agent4cs: A Multi-agent System for Code Summarization in Large Hierarchical Codebases

Agent4cs:面向大型分层代码库的代码摘要多智能体系统

Yongjian Tang, Ezgi Sarikayak, Doruk Tuncel, Jie M. Zhang, Thomas Runkler

机构 * Siemens AG(西门子股份公司) Technical University of Munich(慕尼黑工业大学) Kings College London(伦敦国王学院)

专题命中 评测与基准 :language model(abstract);prompting(abstract);分类 cs.AI

AI总结 提出多智能体框架Agent4cs,通过自底向上方式对大型代码库进行摘要,包含摘要、关键词提取和质量保证三个智能体,在语义一致性和关键词覆盖率上分别提升8%和38%。

Comments Accepted to the main track of the 23rd European Conference on Multi-Agent Systems (EUMAS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31644 2026-07-07 cs.CL cs.CY 新提交 70%

Moral Safety in LLMs: Exposing Performative Compliance with Puzzled Cues

大语言模型中的道德安全:揭示对困惑线索的表演性遵从

Mohammadamin Shafiei, Shuyue Stella Li, Yulia Tsvetkov

机构 * University of Milan(米兰大学) University of Washington(华盛顿大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本研究提出“表演性遵从”概念,揭示大语言模型在道德困境中当身份标签被隐藏时公平性显著下降,并引入线索变化方法和“线索可见性差距”指标来区分真实与表演性道德安全。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09118 2026-07-07 cs.AI 新提交 70%

ComplexConstraints and Beyond: Expert Rubrics for RLVR

复杂约束与超越:RLVR的专家评分标准

Sushant Mehta, Liudas Panavas, Suhaas Garre, Edwin Chen

机构 * Surge AI

专题命中 评测与基准 :LLM(abstract,abstract_cn);分类 cs.AI

AI总结 提出专家设计的评分标准作为评估和训练信号,通过复杂指令遵循和企业智能体任务验证,在RL训练中显著提升模型性能。

Comments Accepted to the GEM workshop at ACL 2026: https://gem-workshop.com/

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02141 2026-07-03 cs.AI 新提交 70%

A$^{2}$utoLPBench: An Auto-Generated, Agent-Friendly LP Benchmark via Inverse-KKT Construction

A$^{2}$utoLPBench:通过逆KKT构造自动生成的、智能体友好的线性规划基准

Shuo Ren, Yaohui Han, Yifan Shi, Libo Shen, Haodong Lu, Dongfang Wu, Rongliang Fu, Bei Yu, Tsung-Yi Ho

机构 * The Chinese University of Hong Kong(香港中文大学)

专题命中 评测与基准 :LLM(abstract,abstract_cn);分类 cs.AI

AI总结 提出A$^{2}$utoLPBench,一种通过逆KKT条件自动生成线性规划文本问题的基准,提供无限新问题、难度可控、答案正确且防数据泄露。

Comments 25 pages and 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01883 2026-07-03 cs.CL 新提交 70%

PairCoder++: Pair Programming as a Universal Paradigm for Verified Code-Driven Multimodal and Structured-Artifact Generation

PairCoder++:结对编程作为验证代码驱动的多模态与结构化工件生成的通用范式

Junhao Chen, Xiang Li, Mingjin Chen, Boran Zhang, Henghaofan Zhang, Yibin Xu, Yuehan Cui, Fangsheng Weng, Fei Ma, Qi Tian, Ruqi Huang, Hao Zhao

机构 * THU(清华大学) PKU(北京大学) PolyU, Hong Kong(香港理工大学) USTC(中国科学技术大学) UESTC(电子科技大学) Tongji University(同济大学) Tianjin University(天津大学) Independent Researcher(独立研究者) Guangming Lab(光明实验室) BAAI(北京智源人工智能研究院)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 提出PairCoder框架,通过驱动者和导航者两个智能体结对编程,利用工具链反馈验证代码生成,在17个基准测试中显著提升可验证工件的生成质量。

Comments Accepted by ACL 2026. Project Page: https://yisuanwang.github.io/PairCoder/

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01280 2026-07-03 cs.LG cs.PL 新提交 70%

Fixed-Set Robustness in Programming by Example: Example Corruption and Semantic Partition Recovery

编程示例中的固定集鲁棒性:示例损坏与语义分区恢复

Yuan Si, Jialu Zhang

专题命中 评测与基准 :LLM(abstract,abstract_cn);分类 cs.LG

AI总结 研究编程示例系统中对抗性示例损坏的鲁棒性,提出版本空间分区聚合(VPA)防御方法,实验表明低裕度任务易受攻击,VPA仅在语义分区投票裕度存在时有效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01252 2026-07-03 cs.CY cs.AI 新提交 70%

How Indian Dermatologists are Utilizing Artificial Intelligence for Clinical Practice and Workflow Management: A Nationwide Survey with a Special Focus on atopic dermatitis

印度皮肤科医生如何将人工智能用于临床实践和工作流程管理:一项以特应性皮炎为重点的全国性调查

Dipayan Sengupta, Saumya Panda, Sandipan Dhar, Dipankar De, Deepika Pandhi, Narayanan B

机构 * Charnock Hospital, Kolkata, India(科钦医院,加尔各答,印度) Department of Dermatology, Jagannath Gupta Institute of Medical Sciences and Hospital, Kolkata, India(皮肤科部,贾亚纳塔·古普塔医学科学与医院,加尔各答,印度) Department of Pediatric Dermatology, Institute of Child Health, Kolkata-700017, India(儿科皮肤科部,儿童健康研究所,加尔各答-700017,印度) Department of Dermatology, Venereology, and Leprology, Postgraduate Institute of Medical Education and Research, Chandigarh-160012, India(皮肤科、性病学和麻风病学部,医学教育与研究研究生院,昌迪加尔-160012,印度) Department of Dermatology and STD, University College of Medical Sciences & GTBH, University of Delhi, Delhi-110095, India(皮肤科和性传播疾病部,医学科学大学及GTBH,德里大学,德里-110095,印度) Department of DVL, Sree Balaji Medical College and Hospital, Chennai, India(DVL部,Sree Balaji医学院和医院,钦奈,印度)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 通过全国性调查,发现印度皮肤科医生主要使用通用大语言模型处理认知和行政任务,而临床需求集中于慢性病管理和特应性皮炎工作流支持,提示临床监督下的工作流工具可能比独立诊断应用更有用。

Comments 28 pages, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.00436 2026-07-02 cs.AI 新提交 70%

PHREEQC-MCQ-200: A Diagnostic Benchmark for Tool-Augmented Scientific Simulator Agents

PHREEQC-MCQ-200:用于工具增强型科学模拟智能体的诊断基准

Ke Zhang, Sahchit Chundur, Mohammad Javad Qomi, Maziar Raissi

机构 * University of California, Riverside(加州大学河滨分校) University of California, Irvine(加州大学尔湾分校)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出PHREEQC-MCQ-200基准,评估工具增强型智能体在地球化学模拟中的表现,发现模拟器访问显著提升准确率但存在非单调性,输出访问协议影响性能。

Comments 30 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.00334 2026-07-02 cs.AI 新提交 70%

Managed Autonomy at Runtime: Gear-Based Safety and Governance for Single- and Multi-Agent Cyber-Physical Systems

运行时管理自主性:基于档位的单/多智能体信息物理系统安全与治理

Srini Ramaswamy, Wang Miaosheng

机构 * DNRS.ai, USA

专题命中 评测与基准 :LLM(abstract,abstract_cn);分类 cs.AI

AI总结 提出一种离散时间控制系统,通过五个执行档位和效用门控调度,在单智能体场景中证明单调稳定性、执行安全性和最终稳定,在多智能体CPS中结合治理状态实现分布式安全保证,在UR5机器人装配单元上达到99.6%异常检测率。

Comments to be submitted to a Journal, 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18293 2026-07-02 cs.SE cs.AI 新提交 70%

Vibe Coding Ate My Homework: An evaluation of AI approaches to greenfield software engineering and programming

Vibe Coding 吃掉我的作业:AI 方法在全新软件工程与编程中的评估

Callum Barbour

机构 * OpenAI

专题命中 评测与基准 :LLM(abstract,abstract_cn);分类 cs.AI

AI总结 本文评估了“氛围编码”(用自然语言提示编程)在全新软件工程任务中的可行性,并分析了现有基准,通过开发 Python 简单独立编程任务评估套件提供见解。

Comments 10 pages, 2 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31518 2026-07-01 cs.AI 新提交 70%

Design and Implementation of Agentic Orchestrations and Orchestration of Agents

代理编排与代理编排的设计与实现

Stefanie Rinderle-Ma, Juergen Mangler, Johannes Loebbecke, Dominik Voigt, Nataliia Klievtsova, Matthias Ehrendorfer

机构 * Technical University of Munich(慕尼黑工业大学)

专题命中 评测与基准 :LLM(abstract,abstract_cn);分类 cs.AI

AI总结 本文提出代理业务流程管理的分类框架,包括任务特异性、可追溯性、自主性等属性,并给出定性决策标准和定量评估指标,通过预测光感场景实例验证。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30755 2026-07-01 cs.CR cs.AI 新提交 70%

Understanding and Evaluating Claw-like Agent Security Through a Computer-Systems Lens

通过计算机系统视角理解与评估爪类智能体安全性

Peizhi Niu, Wenjie Qu, Shangding Gu, Tianneng Shi, Yuankai Li, Ahmad Tawaha, Hend Alzahrani, Vincent Siu, Boyi Li, Chenguang Wang, Jiaheng Zhang, Basel Alomair, Ming Jin, Muhao Chen, Chi Wang, Costas Spanos, Dawn Song

机构 * UIUC(伊利诺伊大学厄巴纳-香槟分校) NUS(新加坡国立大学) UC Berkeley(加州大学伯克利分校) UC Davis(加州大学戴维斯分校) Virginia Tech(弗吉尼亚理工大学) KACST(阿卜杜勒阿齐兹国王科技城) UC Santa Cruz(加州大学圣克鲁兹分校) NVIDIA(英伟达) Google DeepMind(谷歌DeepMind) UW Seattle(华盛顿大学西雅图分校) HUMAIN(胡迈因研究所)

专题命中 评测与基准 :LLM(abstract,abstract_cn);分类 cs.AI

AI总结 针对爪类AI智能体安全漏洞,提出计算机系统类比,构建SafeClawArena基准测试,评估406个对抗任务,发现最高攻击成功率70%,恶意插件100%成功。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27936 2026-06-29 cs.CR cs.AI stat.AP 新提交 70%

Agentic AI-Powered Re-Identification: An Emerging, Scalable Threat to Mobility Microdata Privacy

基于智能体AI的重识别:对移动微数据隐私的新兴可扩展威胁

Oscar Thees, Roman Müller, Matthias Templ

机构 * University of Applied Sciences and Arts Northwestern Switzerland (FHNW)(西北瑞士应用科学和艺术大学(FHNW))

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 研究利用智能体AI自动从公开网络搜索并交叉引用数据,实现从时空轨迹中重识别个体,实验成功率达72%,揭示了大规模重识别威胁。

Comments 15 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27334 2026-06-26 cs.AI 新提交 70%

Language-Based Digital Twins for Elderly Cognitive Assistance

基于语言的数字孪生用于老年人认知辅助

Mohammad Mehdi Hosseini, Mohammad H. Mahoor, Hiroko H. Dodge

机构 * Ritchie School of Engineering and Computer Science, University of Denver(丹佛大学里奇工程与计算机科学学院) Department of Neurology, Massachusetts General Hospital, Harvard Medical School(哈佛医学院麻省总医院神经内科)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出基于大语言模型的数字孪生框架,通过风格测量和上下文元数据模拟老年人对话行为,并引入多条件变分自编码器评估保真度和认知一致性,实现非侵入式认知健康监测。

Comments Accepted and published in the Proceedings of the ACM International Conference on PErvasive Technologies Related to Assistive Environments (PETRA 2026). The final published version is available through the ACM Digital Library

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27247 2026-06-26 cs.LG 新提交 70%

RSPC: A Benchmark for Modeling Stress and Psychiatric Conditions in Digitally Mediated Relationships using Psychiatrist Annotations

RSPC:使用精神科医生标注对数字媒介关系中的压力和精神病状况进行建模的基准

Parmitha Vangapandu, Sai Ganesh Mokkapati, Sathwik Narkedimilli, MSVPJ Sathvik, Timothy Liu, Simon See, Johannes C. Eichstaedt

机构 * Indian Institute of Information Technology Dharwad(印度信息技术学院达尔瓦德分校) Keshav Memorial College of Engineering(克沙夫纪念工程学院) National University of Singapore(新加坡国立大学) University of Birmingham(伯明翰大学) NVIDIA, Singapore(英伟达(新加坡)) Stanford University(斯坦福大学) INSEAD(欧洲工商管理学院)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.LG

AI总结 提出RSPC基准,包含精神科医生标注的Reddit帖子,用于多标签障碍分类、关系触发检测和阶段预测,发现模型能力差异及焦虑与关系不确定性的关联。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27027 2026-06-26 cs.CR cs.AI 新提交 70%

ShareLock: A Stealthy Multi-Tool Threshold Poisoning Attack Against MCP

ShareLock: 针对MCP的隐蔽多工具阈值投毒攻击

Liwei Liu, Tianzhu Han, Zijian Liu, Zishu Dong, Na Ruan

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 评测与基准 :LLM(abstract,abstract_cn);分类 cs.AI

AI总结 提出ShareLock框架,利用Shamir阈值方案将恶意指令分散到多个工具描述中,实现隐蔽性和容错性,平均攻击成功率超90%。

Comments 16 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26857 2026-06-26 cs.AI 新提交 70%

LCAi: Life Cycle Assessment with big data fusion and retrieval-augmented generation-assisted interpretation

LCAi: 基于大数据融合与检索增强生成辅助解释的生命周期评估

Georgios Tsironis, Juan D. Medrano-Garcia, Gonzalo Guillen-Gosalbez

机构 * Institute for Chemical and Bioengineering, Department of Chemistry and Applied Biosciences, ETH Zürich(苏黎世联邦理工学院化学与应用生物科学系化学与生物工程研究所) NCCR Catalysis(瑞士国家研究能力中心催化研究所)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出一种视角条件检索增强生成框架,通过多视角检索与受控合成,在AI辅助生命周期评估中实现结构化解释,减少幻觉风险并保持跨领域多样性。

Comments 23 pages, 14 figures, 6 tables. Includes Supplementary Information

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26836 2026-06-26 cs.AI 新提交 70%

The Capability Frontier: Benchmarks Miss 82% of Model Performance

能力前沿:基准测试遗漏了82%的模型性能

Bradley Fowler, Ryan Smith, Daniel Thi Graviet, William Myers, Joshua Greaves, Narmeen Fatimah Oozeer, Antía García, Philip Quirke, Amirali Abdullah, Fazl Barez, Shriyash Kaustubh Upadhyay

机构 * Martian University of Oxford(牛津大学) ThoughtWorks

专题命中 评测与基准 :LLM(abstract,abstract_cn);分类 cs.AI

AI总结 提出能力前沿概念,通过帕累托前沿量化多模型多生成下的最佳性能,纠正单模型单次评估偏差,在16个基准上实现82%性能提升和85%成本降低。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26552 2026-06-26 cs.CV cs.AI 新提交 70%

Perception, Verdict, and Evolution: Hindsight-Driven Self-Refining Forensics Agent for AI-Generated Image Detection

感知、判断与进化:基于事后洞察的自优化取证智能体用于AI生成图像检测

Yangjun Wu, Keyu Yan, Yu Liu, Jingren Zhou, Fei Huang, Rong Zhang, Zhou Zhao, Fei Wu

机构 * Zhejiang University(浙江大学) Alibaba Group(阿里巴巴集团)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出ForeAgent框架,采用感知-判断架构融合多视图线索,并引入事后洞察驱动的自优化策略,通过采样-反思-进化范式持续提升检测能力,在多个基准上达到最优性能。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏