arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 1423 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 1423 篇

2606.29717 2026-08-13 cond-mat.mtrl-sci cs.AI cs.LG 版本更新 91%

Optimizing Expert-Designed Crystal Graph Networks for Band-Gap Prediction with an Autonomous LLM Research Loop

利用自主LLM研究循环优化专家设计的晶体图网络用于带隙预测

Chenmu Zhang, Boris I. Yakobson

机构 * Department of Materials Science and NanoEngineering(材料科学与纳米工程系)

专题命中 评测与基准 :LLM(title,title_cn);pretraining(abstract);分类 cs.AI、cs.LG

AI总结 提出一个自主LLM研究循环,在MatBench带隙基准上构建了无需外部预训练的最准确模型,超越了所有17个专家设计模型,通过实现元素对特征和空间群嵌入等已知方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07314 2026-07-30 cs.CL cs.AI 版本更新 91%

MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

MEDIC:对LLM在临床应用中的安全性和实用性领先指标的综合评估

Praveenkumar Kanithi, Clément Christophe, Marco AF Pimentel, Tathagata Raha, Prateek Munjal, Nada Saadi, Hamza A Javed, Svetlana Maslenkova, Nasir Hayat, Ronnie Rajan, Shadab Khan

机构 * M42

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 MEDIC通过综合评估框架揭示LLM在临床应用中的安全性和实用性差异,强调需采用组合方法以应对多维度性能权衡。

Comments Published in Transactions on Machine Learning Research (06/2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14568 2026-06-29 cs.SE cs.CL cs.LG 版本更新 91%

Given, When, Then, Again: Mining Subscenario Refactoring Candidates in Behaviour-Driven Test Suites with ML Classifiers and LLM-Judge Baselines

在行为驱动软件测试套件中挖掘子场景重构机会:ML分类器和LLM-判断基线

Ali Hassaan Mughal, Noor Fatima, Muhammad Bilal

机构 * Independent Researcher(独立研究者;应用MBA(数据分析),德克萨斯韦斯利安大学) Applied MBA (Data Analytics), Texas Wesleyan University(独立研究者;计算机工程学士,国立科学与技术大学(NUST)) Independent Researcher(独立研究者;管理硕士,慕尼黑技术大学) B.E. Computer Engineering, National University of Sciences and Technology (NUST) Independent Researcher M.Sc. Management, Technical University of Munich

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 本文通过ML分类器和LLM基线,识别行为驱动开发测试套件中可提取的子场景,量化其在公共BDD生态系统中的普及率。

Comments 31 pages, 10 figures, 6 tables, 56 references. v2: retitled; references corrected and verified; threshold-sensitivity and imbalance-robust metrics added; figures restyled. Code and data (Apache-2.0): https://github.com/amughalbscs16/cukereuse_subscenarios_release (archived: https://doi.org/10.5281/zenodo.20356527). Upstream corpus: https://doi.org/10.5281/zenodo.19754359

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24622 2026-06-24 cs.AI cs.LG 版本更新 91%

Random Rule Forest (RRF): Interpretable and Manageable Ensembles of LLM-Generated Questions for Predicting Success from Unstructured Data

随机规则森林 (RRF): 基于LLM生成问题的可解释且可控集成方法用于从非结构化数据预测成功

Ben Griffin, Aaron Ontoyin Yin, Diego Vidaurre, Ugur Koyluoglu, Joseph Ternasky, Fuat Alican, Yigit Ihlamur

机构 * University of Oxford, United Kingdom Aarhus University, Aarhus, Denmark Centre de Recerca Matem\`atica, Barcelona, Spain Oliver Wyman, New York, United States Vela Research, San Francisco, United States

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 提出随机规则森林 (RRF),利用大语言模型生成简单的是/否问题作为弱学习器,通过等权投票形成可审计的“绿旗”评分卡,在低基准率任务中实现透明且竞争性的预测性能。

Comments 25 pages including appendix, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15851 2026-06-18 cs.CL cs.AI 版本更新 91%

Narrative Theory-Driven LLM Methods for Automatic Story Generation and Understanding: A Survey

叙事理论驱动的LLM方法在自动故事生成与理解中的应用:综述

David Y. Liu, Aditya Joshi, Paul Dawson

机构 * School of Computer Science and Engineering(计算机科学与工程学院) School of Arts and Media(艺术与媒体学院) University of New South Wales (UNSW)(新南威尔士大学)

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);post-training(abstract)

AI总结 综述叙事理论驱动的大语言模型方法在自动故事生成与理解中的应用,分析现状并指出生成任务在理论应用、后训练方法、非虚构叙事及叙事层次等方面落后于理解任务,提出未来方向。

Comments 31 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16346 2026-06-09 cs.CL cs.LG 版本更新 91%

Helpful to a Fault: Measuring Illicit Assistance in Multi-Turn, Multilingual LLM Agents

有益于故障:测量多轮、多语言LLM代理中的非法协助

Nivya Talokar, Ayush K Tarun, Murari Mandal, Maksym Andriushchenko, Antoine Bosselut

机构 * EPFL(苏黎世联邦理工学院) independent(独立研究员) tubingen(图宾根大学)

专题命中 评测与基准 :LLM(title,title_cn);prompting(abstract);分类 cs.CL、cs.LG

AI总结 本文提出STING框架,用于评估多轮多语言LLM代理在执行非法任务时的协助能力,发现低资源语言中攻击成功率不一致,提供实际部署中的压力测试方法。

Comments Accepted in ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23781 2026-07-15 cs.CR 版本更新 91%

Leveraging Large Language Models for Trustworthiness Assessment of Web Applications

利用大语言模型评估网络应用的可信度

Oleksandr Yarotskyi, José D'Abruzzo Pereira, João R. Campos

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);prompting(abstract)

AI总结 本文提出利用大语言模型自动评估网络应用的可信度,通过比较不同提示工程技术,提出基于逻辑偏好分数的分层质量模型,实验表明规则提示能提高评估可靠性。

Comments Accepted for publication in the 19th IEEE International Conference on Software Testing, Verification and Validation (ICST) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13505 2026-06-24 cs.CL cs.AI cs.LG 版本更新 91%

Ensemble Learning for Large Language Models in Text and Code Generation: A Survey

面向文本与代码生成的大语言模型集成学习综述

Mari Ashiga, Wei Jie, Fan Wu, Vardan Voskanyan, Fateme Dinmohammadi, Paul Brookes, Jingzhi Gong, Zheng Wang

机构 * School of Computing and Engineering, University of West London(西伦敦大学计算机与工程学院) Turing Intelligence Technology Limited(图灵智能科技有限公司) School of Computer Science, University of Leeds(利兹大学计算机科学学院)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI、cs.LG

AI总结 本文综述了七种大语言模型集成方法(权重合并、知识融合、混合专家、奖励集成、输出集成、路由和级联),分析了它们在文本与代码生成中提升多样性、输出质量和应用灵活性的能力。

Comments Accepted by IEEE TAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18288 2026-06-19 cs.SE 版本更新 91%

Can Large Language Models Reason About Complex Execution Paths? An Empirical Study on Python

大型语言模型能否推理复杂执行路径?基于Python的实证研究

Wenhan Wang, Kaibo Liu, Zeyu Sun, An Ran Chen, Ge Li, Gang Huang, Lei Ma

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(summary_cn,abstract)

AI总结 本文实证研究大型语言模型在Python执行路径推理中的可行性,构建测试用例生成和缺陷分类任务,发现LLM能提升路径覆盖率,但强推理模型不一定优于弱模型。

Comments Accepted by ACM Transactions on Software Engineering and Methodology (TOSEM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06781 2026-08-06 cs.CL 版本更新 91%

When Better Codebooks Are Not Enough: Predictive Performance and Behavioral Reliability in LLM Political Event Coding

当更好的代码手册还不够:LLM政治事件编码中的预测性能与行为可靠性

Zixian He, Bharath Raahul Murugesan, Patrick Brandt, Yibo Hu

机构 * Independent Researcher(独立研究者) Illinois Institute of Technology(伊利诺伊理工学院) The University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 评测与基准 :LLM(title,title_cn);prompting(abstract);分类 cs.CL

AI总结 本研究探讨在政治事件编码任务中,将专家代码手册优化为LLM友好形式能显著提升分类性能,但预测增益并未完全转化为行为可靠性,模型在代码手册变化下仍可能失效。

Comments 13 pages, 3 figures, 13 tables. Revised version with updated experiments, behavioral reliability analyses, and additional API-model results

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23178 2026-06-25 cs.AI 版本更新 91%

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines

评判评判者:LLM-as-a-Judge pipelines中偏见缓解策略的系统评估

Sadman Kabir Soumik

机构 * Independent Researcher(独立研究员)

专题命中 评测与基准 :LLM(title,title_cn);language model(abstract);分类 cs.AI

AI总结 本文系统评估了LLM-as-a-Judge pipelines中九种偏见缓解策略,发现风格偏见是最主要的偏见类型,且所有模型在扩展对上偏好简洁性,但截断控制能区分质量和长度,表明质量敏感的评估而非单纯长度偏见。

Comments 22 pages, 4 figures. Published in Transactions on Machine Learning Research (2026)

Journal ref Transactions on Machine Learning Research (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29874 2026-06-16 cs.MA cs.AI cs.GT 版本更新 91%

Evolutionary Dynamics of Cooperation in Next-Generation LLM Agent Systems: A Cross-Provider Empirical Extension

下一代LLM智能体系统中合作的演化动力学:跨提供商的实证扩展

Francisco León Zúñiga Bolívar

机构 * Institución Universitaria Colegio Mayor del Cauca(大学机构科尔多瓦大学)

专题命中 评测与基准 :LLM(title,title_cn);prompting(abstract);分类 cs.AI

AI总结 本研究通过扩展Willis等人的基准,测试2025-2026年四个前沿LLM模型在迭代囚徒困境中的合作偏差,发现合作偏差普遍存在但提供商间差异显著,且噪声仍是普遍挑战。

Comments v2 (erratum): two truncated Gemini 3.1 Pro libraries regenerated; cooperative-plurality 9/12->10/12, conclusions unchanged. 11 pages, 3 figures, 8 tables. Extends arXiv:2501.16173. Code and n=500 replication: https://github.com/arqFranciscoLeon/evollm (archived: https://doi.org/10.5281/zenodo.20248615)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07223 2026-08-11 cs.CR cs.AI cs.CL cs.LG cs.SE 版本更新 91%

TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories

TraceSafe: 对LLM在多步骤工具调用轨迹上的安全护栏的系统评估

Yen-Shan Chen, Sian-Yao Huang, Cheng-Lin Yang, Yun-Nung Chen

机构 * CyCraft AI Lab, Taiwan(CyCraft AI实验室(台湾)) National Taiwan University(国立台湾大学)

专题命中 评测与基准 :LLM(title,title_cn);language model(abstract,comments);large language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出TraceSafe-Bench,首个评估中间轨迹安全的综合基准,发现安全护栏效果更依赖结构数据能力而非语义安全对齐,通用模型在轨迹分析中表现更优,且准确性随执行步骤增加而提升。

Comments Accepted to Conference on Language Modeling (COLM) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04156 2026-08-20 cs.AI cs.LG 版本更新 91%

BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding

BrainBench:面向全面脑电理解的大语言模型基准测试

Yangxuan Zhou, Yuning Chen, Chen Wu, Jiquan Wang, Shijian Li, Gang Pan, Sha Zhao

机构 * Zhejiang University(浙江大学) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 该研究推出BrainBench基准测试,涵盖4个子集共17个数据集等,评估大语言模型在两种范式下的脑电理解能力,为相关研究提供可复现测试平台。

Comments 51 pages,28 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04491 2026-08-18 cs.HC cs.AI cs.CL cs.CY 版本更新 91%

A validity-guided workflow for robust large language model research in psychology

面向心理学中鲁棒大语言模型研究的有效性导向工作流程

Zhicheng Lin

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 本文提出一个六阶段工作流程,通过整合心理测量与因果推断,提升心理学中大语言模型研究的有效性,强调研究目标、计算工具验证、实验设计、透明执行、数据分析和结果报告的全面性。

Journal ref Behavior Research Methods, 58, 216 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31483 2026-08-07 cs.CL cs.AI 版本更新 91%

BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali

BenHalluEval:孟加拉语大语言模型的多任务幻觉评估框架

Shefayat E Shams Adib, Ahmed Alfey Sani, Ekramul Alam Esham, Ajwad Abrar, Ishmam Tashdeed, Md Taukir Azam Chowdhury

机构 * Department of Computer Science and Engineering, Islamic University of Technology(伊斯兰科技大学计算机科学与工程系) Department of Computer Science and Engineering, University of California(加州大学计算机科学与工程系)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract_cn);prompting(abstract)

AI总结 针对孟加拉语大语言模型幻觉评估的空白,提出BenHalluEval框架,涵盖四项任务,构建12000个幻觉候选,并提出双轨校准指标BenHalluScore,揭示模型间幻觉校准的显著差异。

Comments Preprint. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03322 2026-08-03 cs.CL cs.AI 版本更新 91%

Can Large Language Models Derive New Knowledge? A Dynamic Benchmark for Biological Knowledge Discovery

大语言模型能否推导新知识?一种动态生物知识发现基准

Chaoqun Yang, Xinyu Lin, Shulin Li, Wenjie Wang, Ruihan Guo, Fuli Feng, Tat-Seng Chua

机构 * National University of Singapore(新加坡国立大学) Tsinghua University(清华大学) University of Science and Technology of China(中国科学技术大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 本文提出DBench-Bio,一种动态生物知识发现基准,通过三阶段流程评估AI的新知识发现能力,揭示当前模型在知识发现上的局限性。

Comments Accepted by KDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21482 2026-07-27 cs.AI cs.CL 版本更新 91%

Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks

无需云端的智能编码:在纵向数据准备任务中评估开放权重的大语言模型

Mack Nixon, Liam Wright, Yevgeniya Kovalchuk, Alison Fang-Wei Wu, Martin Danka, Andy Boyd, David Bann

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 研究在纵向数据准备任务中评估开放权重的大语言模型,介绍开源框架,含真实数据集、任务定义和评估程序,通过基准测试发现当前先进模型表现良好,为治理受限研究中的AI辅助数据准备提供可行路径。

Comments Presented at SLLS 2026; accepted at CLS 2026 and RSS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.03301 2026-07-02 cs.CL cs.AI 版本更新 91%

SHIELD: A Diverse Clinical Note Dataset and Distilled Small Language Models for Enterprise-Scale De-identification

SHIELD: 一个多样化的临床笔记数据集和蒸馏小语言模型,用于企业级去标识化

Jose D. Posada, David Love, Somalee Datta, Priya Desai

机构 * Stanford Medicine(斯坦福医学院)

专题命中 评测与基准 :language model(title,abstract);small language model(title,abstract);LLM(abstract_cn);large language model(abstract)

AI总结 针对现有临床去标识化基准数据集的不足,构建了包含1381份笔记、10229个PHI标注的多样化数据集SHIELD,并通过教师-学生蒸馏框架将大语言模型能力迁移至可本地部署的小模型,在标准工作站上达到0.89精确率和0.88召回率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31393 2026-06-19 cs.CL cs.AI 版本更新 91%

Target-Side Paraphrase Augmentation for Sign Language Translation with Large Language Models

面向手语翻译的大语言模型目标端释义增强

Pedro Dal Bianco, Jean Paul Nunes Reinhold, Oscar Stanchi, Facundo Quiroga, Franco Ronchetti, Ulisses Brisolara Corrêa

机构 * III-LIDI Universidad Nacional de La Plata(III-LIDI国立拉普拉塔大学) CDTEC, Federal University of Pelotas(CDTEC,联邦 Pelotas 大学) CONICET III-LIDI Comision de Investigaciones Cientificas Universidad Nacional de La Plata(科学委员会国立拉普拉塔大学) Universidade Federal de Pelotas(联邦 Pelotas 大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 针对手语翻译中平行语料稀缺和目标词汇长尾分布的问题,提出利用GPT-4o生成参考句子的受控释义变体进行目标端增强,并在三种手语数据集上验证了方法的有效性。

Comments Accepted at GenSign @ CVPR 2026. Non-Proceedings Track (https://genai4sl.github.io/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03085 2026-06-15 cs.LG cs.CL 版本更新 91%

Multi-component Causal Tracing in Large Language Models

大型语言模型中的多组件因果追踪

Zirui Yan, Dennis Wei, Dmitriy A. Katz, Prasanna Sattigeri, Ali Tajer

机构 * Rensselaer Polytechnic Institute(拉特拉姆技术学院) IBM Research(IBM研究院)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.LG

AI总结 本文提出一个统一框架,通过软干预和度量转换高效识别对目标性能指标最关键的多组件子集,优于现有基线方法。

Comments Accepted to ACL 2026 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09740 2026-08-12 cs.SE 版本更新 90%

Security Tests as Executable Specifications for LLM Code Generation: Benefits, Trade-offs, and Coverage Limits

安全测试作为LLM代码生成的可执行规范:益处、权衡与覆盖范围限制

Yunhao Liang, Chengguang Gan, Ruixuan Ying, Hanjun Wei, Zhe Cui, Shiwen Ni

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract)

AI总结 该研究提出将安全测试作为LLM代码生成的可执行规范,开发SecTDD框架,通过多维度评估发现预先展示可见测试可提升联合成功率,结构化反馈修复效果更优但益处受多因素影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03565 2026-08-05 cs.IR 版本更新 90%

Skill Is Not Document: Query-Conditioned Compatibility for LLM Agent Skill Routing

技能不是文档:面向LLM智能体技能路由的查询条件基准与两阶段检索器

Zifei Wang, Wei Wen, Qiang Ji, Keyu Chen, Ruizhi Qiao, Xing Sun

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract)

AI总结 针对LLM智能体技能路由中技能兼容性被忽视的问题,提出Reject-as-Resource Retriever (R3)框架,构建双语基准R3-Skill,并设计两阶段检索系统(R3-Embedding + R3-Reranker)显式建模技能兼容性,显著提升检索效果。

Comments 24 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18790 2026-08-05 cs.CR 版本更新 90%

RoguePrompt: Dual-Layer Ciphering for Self-Reconstruction to Circumvent LLM Moderation

RoguePrompt: 双层加密用于自我重建以规避LLM审核

Benyamin Tafreshian

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract)

AI总结 RoguePrompt通过双层加密技术实现自我重建,有效规避LLM审核并诱导模型执行禁止指令。

Comments This submission has been withdrawn because it has been superseded by a substantially revised and expanded version, available as arXiv:2607.27373. Please refer to and cite the newer version

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15082 2026-08-04 eess.AS 版本更新 90%

From Who Said What to Who They Are: Modular Training-free Identity-Aware LLM Refinement of Speaker Diarization

从谁说了什么到他们是谁:用于说话人 diarization 的无训练模块化身份感知大语言模型优化

Yu-Wen Chen, William Ho, Maxim Topaz, Julia Hirschberg, Zoran Kostic

专题命中 评测与基准 :LLM(title,summary_cn);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 针对现有 SD+ASR 框架缺乏灵活性与真实说话人身份的问题,提出无训练模块化流水线,结合 SD、ASR 与 LLM 优化低置信度标签,在患者-临床医生数据集上实现 29.7% 相对误差降低,提供完整身份检测流水线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22952 2026-07-24 cs.SE 版本更新 90%

Sifting the Noise: A Comparative Study of LLM Agents in Vulnerability False Positive Filtering

筛选噪声:LLM代理在漏洞假阳性过滤中的比较研究

Yunpeng Xiong, Ting Zhang

专题命中 评测与基准 :LLM(title,title_cn);prompting(abstract)

AI总结 本文比较了三种先进的LLM代理框架在漏洞假阳性过滤中的性能,展示了LLM代理在减少SAST噪声方面的有效性,但指出其效果依赖于模型和漏洞类型,且存在计算成本差异。

Comments To appear in Proceedings of the 35th ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04788 2026-07-21 cs.CY 版本更新 90%

From Sycophancy to Deception: A Unified Taxonomy for LLM Spontaneous Misalignment

从幻觉到谋略:一种统一的分类法和基准分析用于LLM欺骗

Jerick Shi, Terry Jingcheng Zhang, Zhijing Jin, Vincent Conitzer

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract)

AI总结 本文提出统一的LLM欺骗分类法,揭示现有基准在欺骗机制覆盖不足的问题,并为开发者和监管者提供改进建议。

Comments Accepted to ICLR Agents in the Wild: Safety, Security, and Beyond Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19011 2026-06-16 cs.CR cs.AI cs.CL cs.LG 版本更新 90%

Do You Really Need a GPU to Guard Your LLM? CPU-Class Classifiers and Multi-Stage Pipelines for Safety Enforcement at Scale

你真的需要GPU来保护你的LLM吗?用于大规模安全执行的CPU级分类器与多阶段流水线

Vasudev Majhi, Dhruv Gupta, Advait Singh, Matthew Barker, Dhruv Kumar

机构 * BITS Pilani(比斯帕利尼大学) Trustwise(Trustwise公司)

专题命中 评测与基准 :LLM(title,title_cn);分类 cs.CL、cs.AI、cs.LG

AI总结 本文研究CPU级分类器(如SVM、梯度提升树)在LLM输入安全检测中的性能,发现其与GPU模型互补,并设计三阶段流水线GuardChain,在80%的分布内查询中达到近峰值精度,降低部署成本。

Comments Under Review. 25 pages, 5 figures, 38 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06210 2026-06-09 cs.CL cs.AI cs.CY cs.LG 版本更新 90%

Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook

基于价值码本的LLM文化价值对齐的分布式开放式评估

Jaehyeok Lee, Xiaoyuan Yi, Jing Yao, Hyunjin Hwang, Roy Ka-Wei Lee, Xing Xie, JinYeong Bak

机构 * KAIST(韩国科学技术院)

专题命中 评测与基准 :LLM(title,title_cn);分类 cs.CL、cs.AI、cs.LG

AI总结 提出DOVE框架,通过率失真变分优化构建价值码本,利用不平衡最优传输度量分布对齐,解决LLM文化价值评估中的构造-组成-上下文挑战。

Comments ICML 2026 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22025 2026-06-11 cs.CL cs.AI cs.IR cs.SE 版本更新 90%

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications

当通用提示改进有害:LLM应用的评估驱动迭代

Daniel Commey

机构 * Daniel Commey

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 提出最小可行评估套件(MVES),通过结构化评估框架和本地复现实验,发现通用提示添加并非单调改进,强调评估驱动的提示迭代。

Comments Technical report. 42 pages, 3 figures. Code, test suites, and result logs: https://github.com/dcommey/llm-eval-benchmarking

详情

展开后加载摘要…

URL PDF HTML 收藏