arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 12228 信号源:cs.CL, cs.AI, cs.LG

1. 其他LLM 12228 篇

2606.05339 2026-06-05 cs.SE cs.AI 77%

A Taxonomy of Runtime Faults in Model Context Protocol Servers

模型上下文协议服务器运行时故障的分类法

Joshua Owotogbe, Indika Kumara, Willem-Jan van den Heuvel, Damian Andrew Tamburri, Antonio Ken Iannillo, Roberto Natella

机构 * Jheronimus Academy of Data Science and Tilburg University(赫伦尼姆数据科学学院和蒂尔堡大学) Jheronimus Academy of Data Science(赫伦尼姆数据科学学院) University of Sannio(萨尼亚大学) University of Luxembourg(卢森堡大学) University of Naples Federico II(那不勒斯费德里科二世大学)

专题命中 其他LLM :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文通过手动分析473个MCP服务器仓库中的837个故障线程,采用自下而上的开放式编码方法,首次建立了MCP服务器运行时故障的经验分类法,包含11个顶层类别和27个子类别(73种叶子故障类型),并通过开发者调查验证了其外部有效性。

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21448 2026-06-03 cs.CL 77%

When Models Refuse: Political Steerability and Feature Richness as Measures of Ideological Depth

当模型拒绝时:政治可操控性与特征丰富度作为意识形态深度的度量

Shariar Kabir

机构 * Bangladesh University of Engineering and Technology(孟加拉工程与技术大学)

专题命中 其他LLM :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出意识形态深度概念,通过可操控性和稀疏自编码器测量的特征丰富度,研究大语言模型拒绝遵循良性指令是否源于能力缺陷而非安全规则。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02260 2026-06-03 cs.LG cs.CR 77%

Position: Adversarial ML for LLMs Is Not Making Any Progress

立场:针对LLM的对抗性机器学习并未取得任何进展

Javier Rando, Jie Zhang, Nicholas Carlini, Florian Tramèr

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 其他LLM :LLM(title_cn);language model(abstract);分类 cs.LG

AI总结 本文认为,在大语言模型时代,对抗性机器学习研究的问题定义更模糊、更难解决且更难以评估,可能导致未来十年仍无法取得有意义进展。

Comments Accepted at ICML 2026 Position Paper Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30443 2026-06-02 cs.CL 77%

Cross-Lingual Steering for Figurative Language Generation

跨语言引导的比喻语言生成

Linfeng Liu, Tiffany Zhan, Louie Hong Yao, Saptarshi Ghosh, Tianyu Jiang

机构 * Department of Computer Science, University of Cincinnati(卡内基梅隆大学计算机科学系) School of Computer Science, Carnegie Mellon University(卡内基梅隆大学计算机科学学院) Independent Researcher(独立研究者)

专题命中 其他LLM :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 通过激活引导技术,研究多语言大模型中比喻语言生成的内部信号是否跨语言可复用,发现跨语言方向可有效转移并增强目标行为。

Comments 40 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31069 2026-06-01 cs.CV cs.CL 77%

Towards Effective Long-Video Event Prediction via Multi-Level Event Semantics Mining

面向有效长视频事件预测的多级事件语义挖掘

Bo Peng, YuanJie Lyu, PengGang Qin, Tong Xu

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 其他LLM :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 提出VISTA框架,通过多级事件语义挖掘(细节级、事件级、未来级)实现长视频事件预测,解决现有模型无法精确提取事件细节和进行细粒度分析的问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16187 2026-06-01 cs.SE cs.LG 77%

MatchFixAgent: Language-Agnostic Autonomous Repository-Level Code Translation Validation and Repair

MatchFixAgent: 语言无关的自主仓库级代码翻译验证与修复

Ali Reza Ibrahimzada, Brandon Paulsen, Reyhaneh Jabbarvand, Joey Dodds, Daniel Kroening

机构 * Siebel School of Computing and Data Science, University of Illinois Urbana-Champaign, Urbana, IL, USA(伊利诺伊大学厄巴纳-香槟分校Siebel计算与数据科学学院) Amazon, Arlington, VA, USA(亚马逊公司,阿灵顿,弗吉尼亚州,美国) University of Oxford, Oxford, UK(牛津大学,牛津,英国)

专题命中 其他LLM :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 提出基于大语言模型的多智能体框架MatchFixAgent,实现语言无关的仓库级代码翻译等价性验证与修复,在验证覆盖率和修复成功率上显著优于现有方法。

Comments Published in ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22825 2026-05-25 cs.DC cs.AI cs.ET cs.PF 77%

KPI2KVI: A Multi Agent Workflow for Calculating Key Value Indicators from Service Descriptions

KPI2KVI:一种从服务描述计算关键价值指标的多智能体工作流

Masoud Shokrnezhad, Tarik Taleb, Yan Chen, Qize Guo

机构 * ICTFICIAL OY(ICTFICIAL公司) Ruhr-Universitaet Bochum(波恩鲁尔大学)

专题命中 其他LLM :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出KPI2KVI工具,通过大语言模型驱动的确定性多智能体工作流,从非结构化服务描述中自动计算关键价值指标(KVI),包括上下文获取、KVI类别选择、KPI生成、数值收集与估计,以及带解释的区间值输出。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12710 2026-05-19 cs.CL cs.CR 77%

Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs

改进方法而非提示:针对大语言模型的进化式 jailbreak 攻击合成

Yunhao Chen, Xin Wang, Juncheng Li, Yixu Wang, Jie Li, Yan Teng, Yingchun Wang, Xingjun Ma

机构 * Department of XXX, University of YYY, Location, Country(XXX系,YYY大学,地点,国家) School of ZZZ, Institute of WWW, Location, Country(ZZZ学院,WWW研究所,地点,国家)

专题命中 其他LLM :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出 EvoSynth 框架,通过在代码空间中进行搜索,而非仅在提示空间中优化,从而提高对大语言模型的 jailbreak 攻击成功率和多样性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15613 2026-05-18 cs.CL 77%

Toward LLMs Beyond English-Centric Development

迈向超越英语中心化发展的语言模型

Sho Takase, Ukyo Honda

机构 * CyberAgent

专题命中 其他LLM :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 研究发现语言模型对英语存在显著偏见,持续预训练并非优于从头训练的低成本方案,未来需加强多语言投入。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11360 2026-05-13 cs.CR cs.AI cs.SE 77%

Options, Not Clicks: Lattice Refinement for Consent-Driven MCP Authorization

选项,而非点击:基于同意的MCP授权的晶格细化

Ying Li, Yanju Chen, Peiran Wang, Issac Khabra, Faysal Hossain Shezan, Yu Feng, Yuan Tian

机构 * University of California, Los Angeles(加州大学洛杉矶分校) University of California, San Diego(加州大学圣地亚哥分校) University of Texas at Arlington(德克萨斯大学阿灵顿分校) University of California, Santa Barbara(加州大学圣芭芭拉分校)

专题命中 其他LLM :LLM(abstract,abstract_cn);prompting(abstract);分类 cs.AI

AI总结 本文提出Conleash,一种客户端中间件,通过风险晶格自动允许安全调用,提升MCP授权的安全性与用户信任。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05709 2026-05-08 cs.AI 77%

Conceal, Reconstruct, Jailbreak: Exploiting the Reconstruction-Concealment Tradeoff in MLLMs

隐藏、重建、突破:在大规模语言模型中利用重建-隐藏权衡

Md Farhamdur Reza, Richeng Jin, Tianfu Wu, Huaiyu Dai

机构 * NC State University(北卡罗来纳州立大学) Zhejiang University(浙江大学)

专题命中 其他LLM :large language model(abstract);language model(abstract);prompting(abstract);分类 cs.AI

AI总结 本文探讨了在多模态大语言模型中利用重建与隐藏的权衡进行意图混淆攻击,提出了一种基于字符移除的变体构造方法,并引入关键词相关的干扰图像以提高攻击效果。

Comments 39 pages, including appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12924 2026-04-30 cs.HC cs.AI 77%

Never say never: Exploring the effects of available knowledge on agent persuasiveness in controlled physiotherapy motivation dialogues

不要说永远:探索可用知识对代理说服力的影响在受控的物理治疗动机对话中

Stephan Vonschallen, Rahel Häusler, Theresa Schmiedel, Friederike Eyssel

机构 * Institute of Business Information Technology, Zurich University of Applied Sciences, Switzerland(瑞士应用科学大学商业信息技术学院) Institute for Information Systems, University of Applied Sciences and Arts Northwestern Switzerland(西北瑞士应用科学大学信息系统研究所) Center for Cognitive Interaction Technology, Bielefeld University, Germany(比勒菲尔德大学认知交互技术中心)

专题命中 其他LLM :LLM(abstract,abstract_cn);prompting(abstract);分类 cs.AI

AI总结 本文研究了可用知识如何影响代理在物理治疗动机对话中的说服力,通过比较ChatGPT生成的回应,发现知识配置显著提升了代理的说服力和表达力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07591 2026-04-21 cs.CL 77%

Creating ConLangs to Probe the Metalinguistic Grammatical Knowledge of LLMs

创建ConLangs以探测LLMs的元语言语法知识

Chihiro Taguchi, Richard Sproat

机构 * University of Notre Dame(诺特达美大学) Sakana AI(萨卡纳人工智能) Notre Dame, IN, USA(印第安纳州诺特达美) Tokyo, Japan(日本东京)

专题命中 其他LLM :LLM(summary_cn,abstract_cn);分类 cs.CL

AI总结 本文提出IASC系统,通过LLMs生成构造语言,探讨LLMs对语言和语言学概念的理解能力,实验显示不同LLM和语言规范在形态语法能力上有显著差异。

Comments 53 pages, 18 tables, 3 figures. Accepted at ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14422 2026-04-17 cs.AI 77%

Demonstration of Pneuma-Seeker: Agentic System for Reifying and Fulfilling Information Needs on Tabular Data

肺气追寻者:用于表格数据上信息需求具象化与实现的代理系统

Muhammad Imam Luthfi Balaka, Raul Castro Fernandez

机构 * The University of Chicago(芝加哥大学)

专题命中 其他LLM :LLM(summary_cn,abstract_cn);分类 cs.AI

AI总结 本文提出Pneuma-Seeker系统,通过将用户信息需求转化为显式关系规范,支持迭代细化、目标数据发现和溯源执行,利用LLM作为透明交互分析伙伴。

Comments ACM CAIS 2026 (Demo)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10530 2026-04-14 cs.SE cs.AI cs.HC 77%

Towards an Appropriate Level of Reliance on AI: A Preliminary Reliance-Control Framework for AI in Software Engineering

迈向适当的AI依赖水平:软件工程中AI的初步依赖控制框架

Samuel Ferino, Rashina Hoda, John Grundy, Christoph Treude

机构 * Monash University(莫纳什大学) Singapore Management University(新加坡管理大学)

专题命中 其他LLM :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文通过二十次访谈提出初步依赖控制框架,探讨软件开发者与AI工具的交互,分析过度依赖和不足依赖的影响,并为未来研究提供方向。

Comments Accepted for publication at the 2nd Workshop on Human-Centered AI for SE (HumanAISE) held at the 34th ACM International Conference on the Foundations of Software Engineering (FSE Companion '26), July 5-9, 2026, Montreal, Quebec, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09502 2026-04-14 cs.AI cs.GT cs.MA econ.TH 77%

Strategic Algorithmic Monoculture: Experimental Evidence from Coordination Games

战略算法单一性:协调博弈中的实验证据

Gonzalo Ballestero, Hadi Hosseini, Samarth Khanna, Ran I. Shorrer

机构 * Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 其他LLM :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 研究通过实验设计区分算法单一性与战略算法单一性,发现LLM在协调任务中表现优异,但在奖励分歧时难以维持多样性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10343 2026-04-14 cs.LG 77%

WaterAdmin: Orchestrating Community Water Distribution Optimization via AI Agents

WaterAdmin: 通过AI代理协调社区供水优化

Jiaqi Wen, Pingbo Tang, Shaolei Ren, Jianyi Yang

机构 * University of Houston(休斯顿大学) Carnegie Mellon University(卡内基梅隆大学) University of California, Riverside(加州大学河滨分校)

专题命中 其他LLM :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文提出WaterAdmin框架,结合LLM进行社区上下文抽象与优化控制,以实现动态环境下供水系统的可靠性和节能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13388 2026-04-13 cs.AI 77%

Reflection of Episodes: Learning to Play Game from Expert and Self Experiences

回溯事件:从专家和自我经验中学习玩游戏

Xiaojie Xu, Zongyuan Li, Chang Lu, Runnan Qi, Yanan Ni, Lumin Jiang, Xiangbei Liu, Xuebo Zhang, Yongchun Fang, Kuihua Huang, Xian Guo, Zhanghua Wu, Zhenya Li

机构 * College of Artificial Intelligence, Nankai University(南开大学人工智能学院) Laboratory for Big Data and Decision, National University of Defense Technology(国防科技大学大数据与决策实验室) Jiangsu Automation Research Institute(江苏自动化研究所) Nanjing Research Institute of Electronic Engineering(南京电子工程研究所)

专题命中 其他LLM :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出ROE框架,通过专家和自我经验学习复杂游戏,实验显示在TextStarCraft II中击败高难度机器人。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11548 2026-04-10 cs.CR cs.AI 77%

One Shot Dominance: Knowledge Poisoning Attack on Retrieval-Augmented Generation Systems

单文档支配:面向检索增强生成系统的知识污染攻击

Zhiyuan Chang, Mingyang Li, Xiaojun Jia, Junjie Wang, Yuekai Huang, Ziyou Jiang, Yang Liu, Qing Wang

机构 * State Key Laboratory of Intelligent Game(智能游戏国家重点实验室) Science and Technology on Integrated Information System Laboratory, Institute of Software Chinese Academy of Sciences(中国科学院软件研究所综合信息系统技术国家级重点实验室) University of Chinese Academy of Sciences(中国科学院大学) Nanyang Technological University(南洋理工大学)

专题命中 其他LLM :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出AuthChain攻击方法,通过单文档污染实现对检索增强生成系统更有效的知识污染攻击,提升攻击成功率并保持隐蔽性。

Comments 15pages, 4 figures; accepted by EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07030 2026-04-09 cs.LG 77%

MoE Routing Testbed: Studying Expert Specialization and Routing Behavior at Small Scale

MoE路由测试平台:研究专家专业化和路由行为的小规模情况

Tobias Falke, Nicolas Anastassacos, Samson Tan, Chankrisna Richy Meas, Chandana Satya Prakash, Nitesh Sekhar, M Saiful Bari, Krishna Kompella, Gamaleldin F. Elsayed

机构 * Amazon AGI(亚马逊AGI)

专题命中 其他LLM :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文提出MoE路由测试平台,通过对比不同路由方法,验证了在小规模下平衡范围是实现专家专业化和高利用的关键因素。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06284 2026-04-09 cs.CR cs.AI 77%

ClawLess: A Security Model of AI Agents

ClawLess:AI代理的安全模型

Hongyi Lu, Nian Liu, Shuai Wang, Fengwei Zhang

机构 * Department of Computer Science and Engineering, Southern University of Science and Technology(南方科技大学计算机科学与工程系) Department of Computer Science and Engineering, Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系) Southern University of Science and Technology(南方科技大学) Hong Kong University of Science and Technology(香港科技大学)

专题命中 其他LLM :large language model(abstract);language model(abstract);prompting(abstract);分类 cs.AI

AI总结 ClawLess提出一种形式化安全框架,通过在最坏情况威胁模型下强制验证策略,确保AI代理运行时的安全性,结合BPF技术实现动态策略的高效执行。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05589 2026-04-08 cs.CR cs.AI 77%

Foundations for Agentic AI Investigations from the Forensic Analysis of OpenClaw

从OpenClaw的取证分析为基础的代理AI研究基础

Jan Gruber, Jan-Niclas Hilgert

机构 * KASTEL Security Research Labs, Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院 KASTEL 安全研究实验室) Fraunhofer Institute for Communication, Information Processing and Ergonomics FKIE(弗劳恩霍夫通信、信息处理与人体工程学研究所 FKIE)

专题命中 其他LLM :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文通过分析OpenClaw的技术设计,提出了一种系统化的代理AI取证方法,揭示了代理执行带来的抽象层和非确定性挑战,为代理AI的系统研究提供了基础。

Comments Preprint. Code and experimental data available at: https://github.com/jgru/forensic-analysis-of-openclaw

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02788 2026-04-06 cs.LG 77%

Structure-Aware Commitment Reduction for Network-Constrained Unit Commitment with Solver-Preserving Guarantees

具有结构意识的承诺减少用于网络受限的机组调度与求解器保持保证

Guangwen Wang, Jiaqi Wu, Yang Weng, Baosen Zhang

专题命中 其他LLM :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文提出一种求解器兼容的维度缩减框架,通过利用机组调度决策中的结构规律,减少分支限界树的搜索负担,保持可行性并确保最优解。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01417 2026-04-03 cs.IR cs.CL 77%

ReFormeR: Learning and Applying Explicit Query Reformulation Patterns

ReFormeR:学习和应用显式查询改写模式

Amin Bigdeli, Mert Incesu, Negar Arabzadeh, Charles L. A. Clarke, Ebrahim Bagheri

机构 * University of Waterloo(滑铁卢大学) University of Toronto(多伦多大学) University of California, Berkeley(加州大学伯克利分校)

专题命中 其他LLM :LLM(abstract);language model(abstract);prompting(abstract);分类 cs.CL

AI总结 ReFormeR通过提取并整合查询改写模式,引导LLM进行受控的改写操作,提升检索效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02132 2026-04-03 cs.CL cs.CR cs.CV cs.IR 77%

One Pic is All it Takes: Poisoning Visual Document Retrieval Augmented Generation with a Single Image

一张图片足矣:通过单张图片对视觉文档检索增强生成进行污染攻击

Ezzeldin Shereen, Dan Ristea, Shae McFadden, Burak Hasircioglu, Vasilios Mavroudis, Chris Hicks

机构 * The Alan Turing Institute(艾伦·图灵研究所) University College London(伦敦大学学院)

专题命中 其他LLM :large language model(abstract);language model(abstract);prompting(abstract);分类 cs.CL

AI总结 本文研究了视觉文档检索增强生成(VD-RAG)对污染攻击的脆弱性,通过单张恶意图片实现针对性和通用性攻击,展示了VD-RAG在目标和通用设置下的漏洞,但在黑盒攻击下表现出一定的鲁棒性。

Comments Published in Transactions on Machine Learning Research (03/2026)

Journal ref Transactions on Machine Learning Research (TMLR), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00260 2026-04-02 cs.LG math.OC 77%

Learning to Shuffle: Block Reshuffling and Reversal Schemes for Stochastic Optimization

学习洗牌:用于随机优化的块重排和反转方案

Lam M. Nguyen, Dzung T. Phan, Jayant Kalagnanam

机构 * IBM Research, Thomas J. Watson Research Center(IBM研究院,托马斯·J·沃森研究中心)

专题命中 其他LLM :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文提出块重排和配对反转结构,通过分析证明其在统一洗牌框架下减少前缀梯度方差常数,并降低顺序敏感性,实验验证了其在凸和非凸基准上的优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29077 2026-04-01 cs.CL 77%

Dual Perspectives in Emotion Attribution: A Generator-Interpreter Framework for Cross-Cultural Analysis of Emotion in LLMs

情绪归因的双重视角:一种用于LLM中跨文化情绪分析的生成-解释框架

Aizirek Turdubaeva, Uichin Lee

机构 * KAIST(韩国科学技术院)

专题命中 其他LLM :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出生成-解释框架,从表达与解读双重视角分析情绪归因,评估六种LLM在15国数据上的表现,揭示文化背景对情绪类型和模型性能的影响,呼吁在LLM中实现文化敏感的情绪建模。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28386 2026-03-31 cs.AI 77%

COvolve: Adversarial Co-Evolution of Large-Language-Model-Generated Policies and Environments via Two-Player Zero-Sum Game

COvolve:通过双人零和游戏实现大语言模型生成的策略与环境对抗共演

Alkis Sygkounas, Rishi Hazra, Andreas Persson, Pedro Zuidberg Dos Martires, Amy Loutfi

机构 * Machine Perception and Interaction Lab, Örebro University(厄勒布鲁大学机器感知与交互实验室)

专题命中 其他LLM :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 COvolve通过双人零和游戏实现环境与策略的对抗共演,利用大语言模型生成可执行代码,使环境和策略共同进化,提升持续学习和泛化能力。

Comments Accepted at GECCO 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25768 2026-03-30 cs.SE cs.AI cs.AR cs.MA 77%

UCAgent: An End-to-End Agent for Block-Level Functional Verification

UCAgent:一种端到端的块级功能验证代理

Junyue Wang, Zhicheng Yao, Yan Pi, Xiaolong Li, Fangyuan Song, Jinru Wang, Yunlong Xie, Sa Wang, Yungang Bao

机构 * State Key Lab of Processors, Institute of Computing Technology, CAS(中国科学院计算技术研究所处理器国家重点实验室) University of Chinese Academy of Sciences(中国科学院大学) Beijing Institute of Open Source Chip(北京开源芯片研究院)

专题命中 其他LLM :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 UCAgent通过构建纯Python验证环境和31阶段细粒度验证流程,解决传统方法在复杂半导体设计验证中的不足,实现98.5%的代码覆盖率和100%的功能覆盖率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24218 2026-03-26 cs.IR cs.AI 77%

Who Benefits from RAG? The Role of Exposure, Utility and Attribution Bias

谁从RAG中受益?曝光、效用和归因偏差的作用

Mahdi Dehghan, Graham McDonald

机构 * University of Glasgow, Glasgow, UK(格拉斯哥大学,格拉斯哥,英国)

专题命中 其他LLM :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文研究RAG中查询组公平性的影响因素,发现RAG系统在不同组查询的平均准确率和改进上存在偏差,揭示了曝光、效用和归因对公平性的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏