arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 32271 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 32271 篇

2603.03555 2026-06-05 cs.MA cs.AI cs.SI 87%

Benchmarking Emergent Coordination in Large-Scale LLM Populations: An Evaluation Framework on the MoltBook Archive

在大规模大语言模型群体中评估涌现协调:对MoltBook档案库的评估框架

Brandon Yee, Pairie Koh

机构 * Management Lab, Yee Collins Research Group(Yee Collins研究组管理实验室)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出了一种评估框架,用于在开放代理环境中评估角色专业化、信息扩散和协作任务解决的涌现协调,通过MoltBook档案库的数据集展示了该框架,并建立了量化基准,揭示了核心-外围结构、重尾级联分布和去中心化任务解决中的严重协调开销。

Comments Substantial Revision Required

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00460 2026-06-05 cs.CL 87%

Pitfalls of Evaluating Language Models with Open Benchmarks

使用开放基准评估语言模型的陷阱

Md. Najib Hasan, Md Mahadi Hassan Sibat, Mohammad Fakhruddin Babar, Souvika Sarkar, Monowar Hasan, Santu Karmaker

专题命中 评测与基准 :language model(title,abstract);LLM(abstract,abstract_cn);large language model(abstract);分类 cs.CL

AI总结 本文探讨了使用开放基准评估语言模型时存在的数据泄露风险,并通过构建作弊模型验证了这种风险,指出开放基准可能无法反映实际应用效果,需补充私有或动态生成的基准以维持评估的完整性。

Comments After further review, we found that the core contribution and methodology substantially overlap with previously published work. As a result, the manuscript does not provide a sufficiently distinct or original contribution in its current form. To avoid repetition in the literature and prevent possible confusion for readers, we believe withdrawal is the most appropriate action

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14291 2026-06-05 cs.LG stat.ML 87%

Advances in Temporal Point Processes: Bayesian, Neural, and LLM Approaches

时间点过程的进展:贝叶斯、神经网络和大语言模型方法

Feng Zhou, Quyu Kong, Jie Qiao, Cheng Wan, Yixuan Zhang, Ruichu Cai

机构 * Center for Applied Statistics and School of Statistics, Renmin University of China(应用统计中心和中国人民大学统计学院) Independent Researcher(独立研究者) School of Computer Science, Guangdong University of Technology(广东工业大学计算机学院) School of Statistics and Data Science, Southeast University(东南大学统计与数据科学学院)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文综述了时间点过程的最新研究,从贝叶斯、深度学习和大语言模型三个角度探讨了模型设计、参数估计以及经典应用领域,并展望了未来的研究挑战和方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21011 2026-06-03 cs.HC cs.AI cs.CY 87%

Generating the Modal Worker: A Cross-Model Audit of Race and Gender in LLM-Generated Personas Across 41 Occupations

生成模态工人:跨模型审计41个职业中LLM生成人设的种族与性别

Ilona van der Linden, Sahana Kumar, Arnav Dixit, Aadi Sudan, Smruthi Danda, David C. Anastasiu, Kai Lukoff

机构 * Human-Computer Interaction Lab, Computer Science and Engineering(人机交互实验室,计算机科学与工程) Santa Clara University(圣克拉拉大学)

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本研究审计了四个大型语言模型生成的150多万个职业人设,通过与BLS数据对比,发现模型压缩了人口统计变异,系统性地扭曲了种族和性别代表性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05302 2026-06-03 cs.AI 87%

PieArena: Ranking and Profiling Language Agents in Realistic Negotiation Scenarios

PieArena:在真实谈判场景中对语言智能体进行排名与画像

Chris Zhu, Sasha Cui, Will Sanok Dufallo, Runzhi Jin, Zhen Xu, Linjun Zhang, Daylian Cain

机构 * Yale University(耶鲁大学) UC Berkeley(加州大学伯克利分校) BloomBerg(摩根大通) Rutgers University(罗格斯大学)

专题命中 评测与基准 :language agent(title,abstract);LLM(summary_cn,abstract_cn);分类 cs.AI

AI总结 本文提出PieArena基准,通过多智能体交互评估LLM的谈判能力,并开发排名模型与行为画像,发现联合意图框架对中低端模型提升显著,前沿模型(如GPT-5)在谈判中达到或超过人类基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30848 2026-06-02 cs.CR cs.CL 87%

LLM Anonymization Against Agentic Re-Identification

LLM匿名化对抗智能体重识别

Ziwen Li, Jianing Wen, Tianshi Li

机构 * Khoury College of Computer Sciences(科里学院计算机科学学院) Northeastern University(东北大学)

专题命中 评测与基准 :LLM(title,title_cn);分类 cs.CL

AI总结 提出AURA框架,通过掩码-重构方法解耦隐私定位与效用保留,并利用对抗性隐私和效用检查,以抵抗基于网络搜索的智能体重识别攻击,同时保留文本的上下文效用。

Comments 32 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14782 2026-06-02 cs.CL 87%

Lessons from the Trenches on Reproducible Evaluation of Language Models

关于语言模型可重复评估的前线经验教训

Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Leo Gao, Jonathan Tow, Baber Abbasi, Alham Fikri Aji, Pawan Sasanka Ammanamanchi, Sidney Black, Jordan Clive, Anthony DiPofi, Julen Etxaniz, Benjamin Fattori, Jessica Zosa Forde, Charles Foster, Jeffrey Hsu, Mimansa Jaiswal, Wilson Y. Lee, Haonan Li, Charles Lovering, Niklas Muennighoff, Ellie Pavlick, Jason Phang, Aviya Skowron, Samson Tan, Xiangru Tang, Kevin A. Wang, Genta Indra Winata, François Yvon, Andy Zou

机构 * MBZUAI IIIT Hyderabad EleutherAI HiTZ Center - Ixa, UPV/EHU Ivy Natal University of Michigan HubSpot LibrAI Kensho Contextual AI Brown University New York University Amazon Yale University HKUST Sorbonne University CMU

专题命中 评测与基准 :language model(title,summary_cn);分类 cs.CL

AI总结 本文基于开发Language Model Evaluation Harness框架的三年经验,总结了语言模型评估中面临的方法论挑战、缺乏可重复性和透明度等问题,并提供了改进评估严谨性和信心的建议。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30803 2026-06-01 cs.AI 87%

PReMISE: Policy Rubrics as Measurement Specifications for LLM Judges

PReMISE:作为LLM评判者测量规范的政策评分标准

Swastik Roy, Rajkumar Pujari, Tharindu Kumarage, Charith Peris, Rahul Gupta, Anna Rumshisky, Pradeep Natarajan, Venkatesh Saligrama

机构 * Amazon AGI(亚马逊人工智能研究院)

专题命中 评测与基准 :LLM(title,title_cn);分类 cs.AI

AI总结 提出PReMISE框架,从人类偏好数据中发现政策级评分标准集,并从结构充分性、可靠性、偏好拟合和对抗鲁棒性四个维度审计评分标准,通过偏好排名选择和可靠性约束修复操作提升评判准确性并降低可被利用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30568 2026-06-01 cs.CL 87%

Generating and Refining Dynamic Evaluation Rubrics for LLM-as-a-Judge

生成与精炼动态评估准则用于LLM-as-a-Judge

Zijie Wang, Eduardo Blanco

机构 * University of Arizona(亚利桑那大学) Department of Computer Science(计算机科学系)

专题命中 评测与基准 :LLM(title,title_cn);分类 cs.CL

AI总结 提出无需人工标注的自动生成细粒度评估准则方法,通过数据集级和实例级粒度生成,并利用元评判奖励信号迭代微调准则生成器,在成对和逐点评估中均超越现有基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27472 2026-05-28 cs.AR cs.AI 87%

AssertLLM2: A Comprehensive LLM Benchmark for Assertion Generation from Design Specifications

AssertLLM2: 从设计规格生成断言的全面的LLM基准测试

Yuchao Wu, Wenji Fang, Jing Wang, Wenkai Li, Ziyan Guo, Zhiyao Xie

机构 * Hong Kong University of Science and Technology(香港理工大学)

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出AssertLLM2基准,包含83个真实设计,通过结构化规格、黄金RTL和变异RTL支持缺陷预防和缺陷狩猎两种实际场景,采用语法有效性、形式可证明性、覆盖率和基于突变的缺陷检测等严格评估框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27401 2026-05-28 cs.CY cs.AI 87%

Using Zero-Shot LLM-Generated Survey Data for Geographically Explicit Population Synthesis

使用零样本大语言模型生成的调查数据进行地理显式人口合成

Taylor Anderson, Sara Von Hoene, Orhan Yagizer Cinar, Emma Von Hoene, Amira Roess, Andrew Crooks, Hamdi Kavak

机构 * Dept. of Geography and Geoinformation Science, George Mason University, Fairfax, VA, USA(地理与地理信息科学系,乔治·马歇尔大学,弗吉尼亚州 Fairfax) Dept. of Computer Science, George Mason University, Fairfax, VA, USA(计算机科学系,乔治·马歇尔大学,弗吉尼亚州 Fairfax) College of Public Health, George Mason University, Fairfax, VA, USA(公共卫生学院,乔治·马歇尔大学,弗吉尼亚州 Fairfax) Dept. of Geography, University at Buffalo, Buffalo, NY, USA(地理系,布法罗大学,纽约州 Buffalo) Dept. of Computational and Data Sciences, George Mason University, Fairfax, VA, USA(计算与数据科学系,乔治·马歇尔大学,弗吉尼亚州 Fairfax)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文评估零样本大语言模型生成的健康调查数据能否作为传统迭代比例拟合工作流的输入,用于地理显式人口合成,并发现其可作为补充输入但尚不能替代真实调查数据。

Comments 15 pages, 5 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04631 2026-05-28 cs.AI 87%

Towards automated data analysis: A guided framework for LLM-based risk estimation

迈向自动化数据分析:基于LLM的风险评估引导框架

Panteleimon Rodis

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出一个在人类指导和监督下利用大语言模型进行数据集风险评估的框架,通过识别模式、生成聚类代码并解释结果,为自动化风险分析奠定基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26440 2026-05-27 cs.CL cs.SE 87%

Conv-to-Bench: Evaluating Language Models Via User-Assistant Dialogues In Code Tasks

Conv-to-Bench: 通过代码任务中的用户-助手对话评估语言模型

Victor M. dos Santos, Andre C. Castro, Samuel L. de S. Toledo, Bruno M. L. Calura, Lisandra C. de M. Menezes, Raul C. R. Mata, Telma W. de L. Soares, Bryan L. M. de Oliveira

机构 * Institute of Mathematics and Computer Science, University of São Paulo(圣保罗大学数学与计算机科学学院) Institute of Informatics, Federal University of Goiás(戈亚斯联邦大学信息学院) HUG Labs(HUG实验室) Advanced Knowledge Center for Immersive Technologies (AKCIT)(沉浸式技术高级知识中心)

专题命中 评测与基准 :language model(title,abstract);LLM(abstract,abstract_cn);large language model(abstract);分类 cs.CL

AI总结 提出Conv-to-Bench框架,自动将多轮用户-助手对话转化为结构化需求清单,用于评估大语言模型,在编程领域与人工标准高度一致且计算开销低。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06213 2026-05-27 cs.AI 87%

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models

超越固定基准和最坏情况攻击:语言模型的动态边界评估

Haoxiang Wang, Da Yu, Huishuai Zhang

专题命中 评测与基准 :language model(title,abstract);LLM(abstract,abstract_cn);large language model(abstract);分类 cs.AI

AI总结 提出动态边界评估(DBE)方法,通过定位模型在随机采样解码下通过概率接近0.5的边界项,构建统一难度尺度的评估协议,以解决固定基准的饱和问题。

Comments This submission is being withdrawn because it was submitted without the knowledge and authorization of all co-authors. The authors need to resolve this authorship/authorization issue before any public posting

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11557 2026-05-26 cs.AI 87%

UniToolCall: Unifying Tool-Use Representation, Data, and Evaluation for LLM Agents

UniToolCall: 统一LLM智能体的工具使用表示、数据与评估

Yijuan Liang, Xinghao Chen, Yifan Ge, Ziyi Wu, Hao Wu, Changyu Zeng, Wei Xing, Xiaoyu Shen

机构 * University of Science and Technology of China(中国科学技术大学) Ningbo Institute of Digital Twin(宁波数字孪生研究所) Eastern Institute of Technology(东部技术研究所) Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系)

专题命中 评测与基准 :LLM(title,title_cn);分类 cs.AI

AI总结 提出UniToolCall框架,通过标准化工具集构建、数据集生成和评估流程,结合22k+工具和390k+训练实例,引入锚点链接机制,在混合设置下使Qwen3-8B单轮严格精度达93.0%,超越GPT、Gemini和Claude。

Comments 21 pages, 10 figures, 9 tables. Code and datasets are publicly available at: https://github.com/EIT-NLP/UniToolCall

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24006 2026-05-26 cs.DC cs.LG 87%

A Tabular Schedule Abstraction for Communication-Aware Evaluation of Pipeline-Parallel LLM Training

一种用于通信感知评估流水线并行LLM训练的表格调度抽象

Daniel Barley, Jonathan Leis, Benjamin Klenk, Holger Fröning

机构 * Hardware and Artificial Intelligence (HAWAII) Lab, Heidelberg University, Heidelberg, Germany(海德堡大学硬件与人工智能实验室) NVIDIA Corporation, Santa Clara, CA, USA(英伟达公司)

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文提出一种表格调度抽象和统一的多抽象方法,通过公式推理、理想化调度表和通信感知执行模拟,比较了GPipe、1F1B、Chimera和Hanayo等流水线调度方案,发现通信会抵消气泡分析的结构优势,调度排名依赖于执行环境。

Comments Accepted at the 25th IEEE International Symposium on Parallel and Distributed Computing (ISPDC 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24000 2026-05-26 cs.CL 87%

Toxicity in Twitch Chats: An LLM-Based Analysis Across Gaming Communities

Twitch聊天中的毒性:基于LLM的游戏社区分析

Ronja Fuchs, Florian Rupp, Timo Bertram, Kai Eckert, Alexander Dockhorn

机构 * Institute for Information Processing(信息处理研究所) Leibniz University Hannover(汉诺威莱布尼茨大学) Department of Computer Science(计算机科学系) Technische Hochschule Mannheim(曼海姆技术学院) Institute for Machine Learning(机器学习研究所) Johannes Kepler University(约翰·凯普勒大学) SDU Metaverse Lab(SDU元宇宙实验室) University of Southern Denmark(南部丹麦大学)

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 使用预训练大语言模型对Twitch平台4452个直播流约2000万条聊天消息进行零样本分类,发现2.4%的消息有毒,其中MOBA游戏毒性最高(3.2%),体育游戏最低(2%),且游戏间毒性分布差异显著,表明存在游戏特定的社区规范。

Comments 8 pages, 2 figures, 5 tables. Accepted at the IEEE Conference on Games (IEEE CoG) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22099 2026-05-22 cs.CL 87%

A Comparative Study of Language Models for Khmer Retrieval-Augmented Question Answering

一种用于柬埔寨检索增强问答的语言模型比较研究

Sereiwathna Ros, Phannet Pov, Ratanaktepi Chhor, Kimleang Ly, Wan-Sup Cho, Saksonita Khoeurn

机构 * Department of Computer Science, Chungbuk National University(Chungbuk National University 计算机科学系) Department of Big Data, Chungbuk National University(Chungbuk National University 大数据系) General Department of Information and Communication Technology, Ministry of Post and Telecommunications(邮电部信息和通信技术总局) Department of Management Information Systems, Chungbuk National University(Chungbuk National University 管理信息系统系) BigDatalabs Co., Ltd(BigDatalabs 公司)

专题命中 评测与基准 :language model(title,abstract);LLM(abstract,abstract_cn);large language model(abstract);分类 cs.CL

AI总结 本文针对低资源非拉丁语种柬埔寨语言,比较了多种语言模型在检索增强问答任务中的性能,发现检索器选择是影响效果的关键因素,生成器在不同指标上表现各异。

Comments 14 pages, 1 figure,

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15676 2026-05-22 cs.SE cs.AI 87%

Automated Self-Testing as a Quality Gate: Evidence-Driven Release Management for LLM Applications

自动化自我测试作为质量门:基于证据的LLM应用发布管理

Alexandre Cristovão Maiorano

机构 * Lumytics

专题命中 评测与基准 :LLM(title,title_cn);分类 cs.AI

AI总结 本文提出了一种自动化自我测试框架,通过五个实证基础的维度(任务成功率、研究环境保持、P95延迟、安全通过率和证据覆盖)实现基于证据的发布决策(PROMOTE/HOLD/ROLLBACK),并通过长期案例研究评估了该框架在多代理对话AI系统中的有效性。

Comments 20 pages, 6 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.21497 2026-05-22 cs.CR cs.AI 87%

Autonomous LLM Agents & CTFs: A Second Look

自主大语言模型代理与CTF:再看一次

Youness Bouchari, Matteo Boffa, Marco Mellia, Idilio Drago, Thanh Minh Bui, Dario Rossi

机构 * Politecnico di Torino(托尔托纳理工大学) Università di Torino(托尔托纳大学) Huawei Technologies France(华为法国技术)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文重新审视了大语言模型代理在自动化进攻性安全任务中的表现,通过在30个基于网络的CTF挑战中测试不同架构的代理,发现通用代理在性能上与定制架构相当,并揭示了当前代理在某些类别中的持续障碍。

Comments Accepted at DeMeSSAI Workshop @ IEEE EuroS&P 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01482 2026-05-21 cs.CL 87%

Towards Consistent Detection of Cognitive Distortions: LLM-Based Annotation and Dataset-Agnostic Evaluation

迈向认知扭曲一致检测:基于大语言模型的标注与数据集无关评估

Neha Sharma, Navneet Agarwal, Kairit Sirts

机构 * LREC-2026(LREC-2026会议)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文探讨了利用大语言模型作为一致且可靠的标注器进行认知扭曲检测的方法,并提出了一种数据集无关的评估框架,以公平比较不同数据集训练的模型,结果显示GPT-4能产生一致的标注,提升了模型在主观NLP任务中的表现。

Journal ref https://lrec.elra.info/lrec2026-main-851

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20485 2026-05-21 cs.LG 87%

ZEBRA: Zero-shot Budgeted Resource Allocation for LLM Orchestration

ZEBRA: 零样本预算化资源分配用于LLM编排

May Hamri, Inbal Talgam-Cohen

机构 * Tel Aviv University(特拉维夫大学)

专题命中 评测与基准 :LLM(title,title_cn);分类 cs.LG

AI总结 该研究提出ZEBRA框架,通过将多阶段预算分配转化为连续非线性背包问题,有效解决多智能体流水线中预算分配问题,实验显示其在多个任务上均优于传统方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23355 2026-05-19 cs.AI 87%

LEGO: An LLM Skill-Based Front-End Design Generation Platform

LEGO: 一个基于LLM技能的前端设计生成平台

Jincheng Lou, Ruohan Xu, Jiecheng Ma, Runzhe Tao, Xinyu Qu, Yibo Lin

机构 * School of IC, Peking University(北京大学集成电路学院) School of EECS, Peking University(北京大学电子信息技术学院) School of Microelectronics, Xidian University(西安电子科技大学微电子学院) Institute of EDA, Peking University(北京大学EDA研究院) Beijing Advanced Innovation Center for IC(北京集成电路先进创新中心)

专题命中 评测与基准 :LLM(title,title_cn);分类 cs.AI

AI总结 本文提出LEGO平台,通过将数字前端流程分解为六个独立步骤,并将每个代理能力表示为标准化的可组合电路技能,实现了高效的前端设计生成,显著提升了RTL设计自动化的效果。

Comments Accepted to ISEDA 2026. Best Paper Nomination. 7 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17292 2026-05-19 cs.AI cs.MA 87%

MetaCogAgent: A Metacognitive Multi-Agent LLM Framework with Self-Aware Task Delegation

MetaCogAgent: 一种具有自我意识的任务委托多智能体大语言模型框架

Chenyu Wang, Yang Shu

机构 * School of Computer and Artificial Intelligence, Zhengzhou University, Zhengzhou, China(郑州大学计算机与人工智能学院) Zhejiang University, Hangzhou, China(浙江大学)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出MetaCogAgent框架,通过引入元认知自我评估单元,使每个智能体在执行任务前评估自身能力边界,从而提升任务准确性并减少API调用次数。

Comments 6 pages, submitted to IEEE SMC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17279 2026-05-19 cs.SE cs.AI 87%

Rover: Context-aware Conflict Resolution with LLM

Rover: 基于上下文的冲突解决系统

Qingyu Zhang, Junzhe Li, Jiayi Lin, Changhua Luo, Chenxiong Qian

机构 * The University of Hong Kong(香港大学) Wuhan University(武汉大学)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出Rover,一种结合程序分析和大语言模型的冲突解决系统,通过多层代码属性图获取上下文感知提示,提升代码合并的准确性与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16264 2026-05-19 cs.HC cs.CL 87%

LLM-Based Intelligent Notification Composition: From Static Personalization to Context-Aware Persuasive Messaging

基于大语言模型的智能通知生成:从静态个性化到情境感知的说服性信息

Nilesh Agrawal

机构 * Independent Researcher(独立研究者)

专题命中 评测与基准 :LLM(title,summary_cn);分类 cs.CL

AI总结 本文提出利用大语言模型提升通知信息质量,通过六个维度评估其效果,展示LLM在提升CTR和说服性方面的贡献,并提出决策框架以指导LLM生成的应用。

Comments 17 pages, 1 figure, 7 tables. Code available at https://github.com/ndagrawal/LLMNotificationComposition

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15669 2026-05-18 cs.LG 87%

Rule2DRC: Benchmarking LLM Agents for DRC Script Synthesis with Execution-Guided Test Generation

Rule2DRC:用于DRC脚本合成的LLM代理基准测试

Jinuk Kim, Junsoo Byun, Donghwi Hwang, Seong-Jin Park, Hyun Oh Song

机构 * Department of Computer Science and Engineering, Seoul National University(首尔国立大学计算机科学与工程系) Neural Processing Research Center(神经处理研究所以) Samsung Electronics Co., Ltd(三星电子公司)

专题命中 评测与基准 :LLM(title,title_cn);分类 cs.LG

AI总结 Rule2DRC是一个大规模基准,用于评估DRC脚本生成代理,包含1000个规则到脚本任务和13921个用于执行评分的评估芯片布局。它提供了一种通过DRC执行结果衡量功能正确性的评估流程,并引入SplitTester生成区分性测试用例以提升Best-of-N选择性能。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14537 2026-05-15 cs.AI 87%

Cattle Trade: A Multi-Agent Benchmark for LLM Bluffing, Bidding, and Bargaining

牛市交易:一种多智能体基准,用于评估大语言模型在战略推理、对抗互动和资源约束下的表现

Robert Müller, Clemens Müller

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出Cattle Trade基准,用于评估大语言模型在多智能体经济游戏中整合战略推理、对抗互动和资源分配的能力,揭示了智能体在资源管理与对抗策略中的表现差异。

Comments malgai workshop at iclr 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14153 2026-05-15 cs.CR cs.AI 87%

ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents

ExploitBench: 一个用于LLM网络安全代理的能力阶梯基准

Seunghyun Lee, David Brumley

机构 * Carnegie Mellon University(卡内基梅隆大学) Bugcrowd

专题命中 评测与基准 :LLM(title,title_cn);分类 cs.AI

AI总结 ExploitBench通过16个可测量标志分解exploitation过程,评估模型在网络安全任务中的能力差异,揭示了公开部署模型与私有模型在exploitation能力上的显著差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13481 2026-05-14 cs.CL 87%

PersonalAI 2.0: Enhancing knowledge graph traversal/retrieval with planning mechanism for Personalized LLM Agents

PersonalAI 2.0:通过规划机制增强知识图谱遍历/检索以实现个性化大语言模型代理

Mikhail Menschikov, Matvey Iskornev, Alexander Kharitonov, Alina Bogdanova, Mikhail Belkin, Ekaterina Lisitsyna, Artyom Sosedka, Victoria Dochkina, Ruslan Kostoev, Ilia Perepechkin, Evgeny Burnaev

机构 * Huawei, Moscow, Russia(华为,莫斯科,俄罗斯)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 PersonalAI 2.0通过动态多阶段查询流程提升知识图谱遍历与检索能力,结合实体提取和生成线索查询,提升回答准确性,优于现有方法,且在多个基准测试中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏