arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

1512.01639 2015-12-08 cs.CL stat.ML 57%

PJAIT Systems for the IWSLT 2015 Evaluation Campaign Enhanced by Comparable Corpora

Krzysztof Wołk, Krzysztof Marasek

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Journal ref Proceedings of the 12th International Workshop on Spoken Language Translation, Da Nang, Vietnam, December 3-4, 2015, p.101-104

详情

展开后加载摘要…

URL PDF HTML 收藏
1511.02196 2015-11-09 cs.LG 57%

Evaluating Protein-protein Interaction Predictors with a Novel 3-Dimensional Metric

Haohan Wang, Madhavi K. Ganapathiraju

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments This article is an extended version of a poster presented in AMIA TBI 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
1506.06272 2015-06-23 cs.CV cs.LG stat.ML 57%

Aligning where to see and what to tell: image caption with region-based attention and scene factorization

Junqi Jin, Kun Fu, Runpeng Cui, Fei Sha, Changshui Zhang

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1506.01171 2015-06-04 cs.CL 57%

A Hybrid Model for Enhancing Lexical Statistical Machine Translation (SMT)

Ahmed G. M. ElSayed, Ahmed S. Salama, Alaa El-Din M. El-Ghazali

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 9 pages, 10 figures

Journal ref International Journal of Computer Science Issues Volume 12, Issue 2, March 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
1405.1124 2014-05-07 cs.AI 57%

An ASP-Based Architecture for Autonomous UAVs in Dynamic Environments: Progress Report

Marcello Balduccini, William C. Regli, Duc N. Nguyen

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments Proceedings of the 15th International Workshop on Non-Monotonic Reasoning (NMR 2014)

详情

展开后加载摘要…

URL PDF HTML 收藏
1305.2981 2013-05-15 cs.CY cs.SI 57%

Metrics for Computing Trust in a Multi-Agent Environment

Sanat Kumar Bista, Keshav P. Dahal, Peter I. Cowling, Bhadra Man Tuladhar

专题命中 安全评测 :trustworthy(abstract);分类 cs.CY

Comments SKIMA 2006

详情

展开后加载摘要…

URL PDF HTML 收藏
1203.3495 2012-03-19 cs.LG stat.ML 57%

Parameter-Free Spectral Kernel Learning

Qi Mao, Ivor W. Tsang

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments Appears in Proceedings of the Twenty-Sixth Conference on Uncertainty in Artificial Intelligence (UAI2010)

详情

展开后加载摘要…

URL PDF HTML 收藏
1202.3761 2012-02-20 cs.LG stat.ML 57%

New Probabilistic Bounds on Eigenvalues and Eigenvectors of Random Kernel Matrices

Nima Reyhani, Hideitsu Hino, Ricardo Vigario

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1110.1016 2011-10-06 cs.AI 57%

Engineering Benchmarks for Planning: the Domains Used in the Deterministic Part of IPC-4

S. Edelkamp, R. Englert, J. Hoffmann, F. Liporace, S. Thiebaux, S. Trueg

专题命中 安全评测 :safety(abstract);分类 cs.AI

Journal ref Journal Of Artificial Intelligence Research, Volume 26, pages 453-541, 2006

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26462 2026-08-28 cs.LG cs.AI cs.CL 新提交 56%

Diff Mining: Logit Differences Reveal Finetuning Objectives

Diff Mining:对数差揭示微调目标

Greg Kocher, Robert West, Clément Dumas, Julian Minder

机构 * EPFL(洛桑联邦理工学院) ENS Paris-Saclay(巴黎萨克雷高等师范学校) Université Paris-Saclay(巴黎萨克雷大学) MATS

专题命中 安全评测 :分类 cs.CL、cs.AI、cs.LG;trustworthy(comments)

AI总结 提出Diff Mining框架,通过对比微调模型与基础模型的对数来识别微调目标,在微调领域检测、偏见识别等任务上优于现有方法,可用于开发微调审计工具。

Comments 37 pages, 7 figures. ICLR 2026 Workshop: Principled Design for Trustworthy AI. Code available at https://github.com/science-of-finetuning/diffing-toolkit

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20983 2026-08-24 cs.SI cs.HC 新提交 56%

Beyond Truth Discovery: A Two-Stage Framework to Assess the Severity of False Claim during Disasters

超越真相发现:一种评估灾害期间虚假声明严重程度的两阶段框架

Ruichen Yao, Tejna Dasari, Gulshat Baispay, Aizhan Zaurbek, Yifan Liu, Yaokun Liu, Zelin Li, Dong Wang

专题命中 安全评测 :alignment(abstract,comments)

AI总结 该研究针对灾害中社交媒体虚假声明严重程度评估问题,提出两阶段框架,结合可信度与危害性维度构建基准,发现LLMs及上下文学习与人工判断对齐度更强。

Comments The 2026 ACM Conference on Human-AI Complementarity and Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26159 2026-08-11 cs.AI cs.CY cs.LG 版本更新 56%

When benchmark inferences do not compose: Projectibility in AI evaluation

当基准推理无法组合:AI评估中的可投射性

Brett Reynolds

机构 * Humber Polytechnic(汉伯理工学院) University of Toronto(多伦多大学)

专题命中 安全评测 :分类 cs.AI、cs.CY、cs.LG;alignment(comments)

AI总结 本文针对AI评估中基准推理无法组合的问题,提出非组合原则,结合古德曼的竞争延伸问题与基于论证的有效性框架,通过案例和模拟开发可投射性审计以诊断基准到应用论证的衔接缺陷。

Comments 34 pages, 2 figures, 5 tables. v2 substantially revises Secs. 5-8 and the conclusion, adds a measured instance of factor-structure instability, and corrects a claim in Sec. 3.3 that endpoint alignment suffices for composition. Supersedes the withdrawn arXiv:2510.15236. Code: https://github.com/BrettRey/benchmark-inference-composition

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00826 2026-08-04 cs.CR cs.HC 新提交 56%

XR-PRISM: Data-Driven Privacy and Risk Impact Scoring Metric for Extended Reality in Healthcare

XR-PRISM:面向医疗领域扩展现实的数据驱动型隐私与风险影响评分指标

Nafisa Anjum, M. Rasel Mahmud

专题命中 安全评测 :safety(abstract);trustworthy(comments)

AI总结 该研究针对医疗XR领域缺乏统一风险评估框架的问题,提出XR-PRISM评分指标,基于65篇文献构建四层威胁分类,发现多数对策缺乏标准化评估,该指标可透明量化医疗XR的安全隐私风险。

Comments Published at the 1st International Workshop on Trustworthy, Secure, and Privacy-Aware AI for Extended Reality (TRUST-XR 2025), held in conjunction with IEEE ISMAR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.09766 2026-07-14 cs.AI cs.CL cs.LG cs.MA 新提交 56%

Norm Enforcement for AI Agents: Robustly Shaping Behavior in Multi-Agent Systems

人工智能代理的规范执行:在多智能体系统中稳健塑造行为

Yaowen Ye, Jacob Steinhardt

专题命中 安全评测 :分类 cs.CL、cs.AI、cs.LG;trustworthy(comments)

AI总结 研究多智能体系统中语言模型代理的规范执行机制,针对简单机制易被利用的问题,确定估计代理可靠性及加重惩罚更新估计这两个关键要素,所构建机制可抵御利用且低成本惩罚违规,是塑造智能体行为的可扩展手段。

Comments ICML2026 Trustworthy AI for Good Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12793 2026-04-15 cs.HC 56%

Human Agency, Causality, and the Human Computer Interface in High-Stakes Artificial Intelligence

人类代理、因果性与高风险人工智能中的人机接口

Georges Hattab

专题命中 安全评测 :trustworthy(abstract);alignment(comments)

AI总结 本文探讨高风险人工智能系统中人类代理的保护问题,提出因果代理框架以解决人机交互中的不确定性与因果控制问题。

Comments 2026 CHI Workshop on Human-AI Interaction Alignment: Designing, Evaluating, and Evolving Value-Centered AI For Reciprocal Human-AI Futures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06140 2026-03-18 cs.CV 56%

Boosting the Local Invariance for Better Adversarial Transferability

提升局部不变性以增强对抗迁移性

Bohan Liu, Xiaosen Wang

机构 * School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院)

专题命中 安全评测 :trustworthy(abstract,comments)

AI总结 本文提出LI-Boost方法,通过提升对抗扰动的局部不变性来增强模型间对抗迁移性,实验表明其在多种攻击类型上均有效。

Comments Code is available at https://github.com/Trustworthy-AI-Group/TransferAttack

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05066 2026-01-16 cs.CV 56%

Beautiful Images, Toxic Words: Understanding and Addressing Offensive Text in Generated Images

美丽的图像,有毒的词语:理解并解决生成图像中的冒犯性文本

Aditya Kumar, Tom Blanchard, Adam Dziedzic, Franziska Boenisch

专题命中 安全评测 :safety(abstract);alignment(comments)

AI总结 本文提出了一种针对生成图像中冒犯性文本的微调策略,并发布了ToxicBench基准,用于评估和改进文本到图像模型的安全性。

Comments Accepted at AAAI 2026 (AI Alignment Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09110 2025-10-28 cs.CV 56%

Pulling Back the Curtain: Unsupervised Adversarial Detection via Contrastive Auxiliary Networks

Eylon Mizrahi, Raz Lapid, Moshe Sipper

机构 * Ben-Gurion University(本·古里安大学) DeepKeep

专题命中 安全评测 :safety(abstract);trustworthy(journal_ref)

Comments Accepted for Oral Presentation at SafeMM-AI @ ICCV 2025 (Spotlight)

Journal ref ICCV 2025 Workshop on SafeMM-AI: Safe and Trustworthy Multimodal AI Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15075 2025-08-26 cs.CL cs.AI cs.CV cs.LG 56%

Traveling Across Languages: Benchmarking Cross-Lingual Consistency in Multimodal LLMs

Hao Wang, Pinzhi Huang, Jihan Yang, Saining Xie, Daisuke Kawahara

专题命中 安全评测 :分类 cs.CL、cs.AI、cs.LG;prompt injection(comments)

Comments The first version of this paper mistakenly included a prompt injection phrase, which was inappropriate and unprofessional. Although we corrected the version on arXiv and withdrew from the conference, my co-authors and university strongly request a full withdrawal. Given the situation, I no longer have the authority to manage this paper, and withdrawing it from arXiv is the most responsible action

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.11402 2025-03-13 cs.CL cs.AI cs.LG 56%

Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?

Neelabh Sinha, Vinija Jain, Aman Chadha

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 安全评测 :分类 cs.CL、cs.AI、cs.LG;trustworthy(comments)

Comments Accepted at The Fifth Workshop on Trustworthy Natural Language Processing (TrustNLP 2025) in Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics (NAACL), 2025. 8 pages + references + Appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.08650 2025-02-14 cs.CV 56%

ColorSense: A Study on Color Vision in Machine Visual Recognition

Ming-Chang Chiu, Yingfei Wang, Derrick Eui Gyu Kim, Pin-Yu Chen, Xuezhe Ma

机构 * IBM Research(IBM研究院)

专题命中 安全评测 :safety(abstract);trustworthy(comments)

Comments 12 pages, 11 figures, Accepted at Secure and Trustworthy Machine Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16291 2025-01-28 cs.GT cs.FL 56%

Sequential Decision Making in Stochastic Games with Incomplete Preferences over Temporal Objectives

Abhishek Ninad Kulkarni, Jie Fu, Ufuk Topcu

专题命中 安全评测 :trustworthy(abstract);alignment(comments)

Comments 9 pages, 3 figures, accepted at AAAI 2025 (AI alignment track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.08187 2023-08-17 cs.LG cs.AI 54%

Endogenous Macrodynamics in Algorithmic Recourse

Patrick Altmeyer, Giovan Angela, Aleksander Buszydlik, Karol Dobiczek, Arie van Deursen, Cynthia C. S. Liem

专题命中 安全评测 :trustworthy(comments,journal_ref);分类 cs.AI、cs.LG

Comments 12 pages, 11 figures. Originally published at the 2023 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML). IEEE holds the copyright

Journal ref in 2023 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), Raleigh, NC, USA, 2023 pp. 418-431

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26519 2026-08-28 cs.CE cs.SE 新提交 50%

Report of the 2026 Workshop on Next-Generation Ecosystems for Scientific Computing: Harnessing Community, Software, and AI for Cross-Disciplinary Team Science

2026年科学计算下一代生态系统研讨会报告:利用社区、软件与AI推进跨学科团队科学

Lois Curfman McInnes, Dorian Arnold, Prasanna Balaprakash, Mike Bernhardt, Franck Cappello, Beth Cerny, Deborah DiazGranados, Anshu Dubey, Nichole Etienne, Roscoe Giles, Diego Gomez-Zara, Denice Ward Hood, Mary Ann Leung, Vanessa Lopez-Marrero, Olivia B. Newton, Irene Qualters, Keita Teranishi, Stefan M. Wild, Gabrielle Allen, Richard Arthur, Alexandra Ballow, Tony Baylis, David E. Bernholdt, Daniel Bielich, Johanna Cohoon, Jeremy Crampton, Charles Ferenbaugh, Stephen M. Fiore, Thomas Herault, Tanzima Islam, Stephen Jacobsohn, Meifeng Lin, Charles Lively, Satoshi Matsuoka, Stasa Milojevic, Daniel Nichols, Chris Oehmen, Santiago Ospina Tabares, Michael E. Papka, Katherine Riley, Damian Rouson, Sudip K. Seal, Brittany Segundo, John Shalf, Andrew Siegel, Valerie Taylor, Jim Willenbring, Lou Woodley

专题命中 安全评测 :trustworthy(abstract)

AI总结 本报告梳理2026年科学计算下一代生态系统研讨会成果,明确四大战略主题及八项社区行动优先事项,为AI时代构建可信可持续的科学计算生态系统指明方向。

Comments 27 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26382 2026-08-28 cs.CV 新提交 50%

VIPER: An Expert-Curated Benchmark for Vision-Language Models in Veterinary Pathology

VIPER:用于兽医病理学视觉语言模型的专家策划基准

Luca L. Weishaupt, Simone de Brot, Javier Asin, Llorenç Grau-Roma, Nic G. Reitsam, Andrew H. Song, Dongmin Bang, Stefan T. Kaluziak, Long Phi Le, Jakob Nikolas Kather, Faisal Mahmood, Guillaume Jaume

机构 * Harvard-MIT HST(哈佛-麻省理工卫生科学与技术部) Mass General Brigham(麻省总医院布里格姆医疗系统) Harvard Medical School(哈佛医学院) COMPATH, University of Bern(伯尔尼大学COMPATH) UC Davis(加州大学戴维斯分校) University of Augsburg(奥格斯堡大学) UT MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心) TU Dresden(德累斯顿工业大学) University of Lausanne(洛桑大学)

专题命中 安全评测 :safety(abstract)

AI总结 研究针对现有病理视觉语言模型基准聚焦人体组织的问题,推出首个兽医病理学视觉语言模型评估基准VIPER,测试16类模型发现领域差距,验证领域特定训练的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24232 2026-08-28 cs.IR 版本更新 50%

Strategy-Aware Parameter-Efficient Adaptation for LLM-based Auto-Bidding

基于大语言模型的自动出价的策略感知参数高效适配

Songyue Cai, Lianyu Wang, Shan Gu, Ziru Xu, Jian Xu, Xiaofeng Zhu, Bo Zheng

专题命中 安全评测 :alignment(abstract)

AI总结 研究广告自动出价问题,提出SAGE框架,通过位置增强、文本对齐和约束门控LoRA三个组件,实现参数高效多模态对齐,在大规模基准实验中性能卓越,调整参数少,消融研究验证各组件贡献。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25647 2026-08-28 cs.CE 版本更新 50%

Smallholder Farmers on Remote Islands Receiving Retrieval-Grounded Multilingual LLM Assistance and Agronomic Advice

基于检索的多语言LLM辅助工具用于岛屿小农户

Nikolaos D. Tantaroudas, Ilias Karachalios, Andrew J. McCracken

专题命中 安全评测 :trustworthy(abstract)

AI总结 针对偏远岛屿小农户获取农业建议困难且方言知识缺乏全球语料支持的问题,提出嵌入双语电商平台的对话AI助手Falco eleonorae,通过工具增强检索(MCP)获取本地化数据,实现可信的多语言、语音和图像交互。

Comments 13, 4

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23749 2026-08-28 cs.SE 版本更新 50%

A Survey of LLM-based Automated Program Repair: Taxonomies, Design Paradigms, and Applications

基于大型语言模型的自动程序修复综述:分类、设计范式与应用

Boyang Yang, Zijian Cai, Fengling Liu, Bach Le, Lingming Zhang, Tegawendé F. Bissyandé, Yang Liu, Haoye Tian

专题命中 安全评测 :alignment(abstract)

AI总结 本文综述了基于LLM的自动程序修复,提出统一分类法,涵盖四种范式,并讨论了评估实践和未来研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25340 2026-08-27 cs.HC 新提交 50%

HRGuard: Gating Relationship Manipulation in Multi-Turn Agentic AI Conversations

HRGuard:多轮智能体AI对话中的门控关系操纵

Pei-Sze Tan, Tasuku Igarashi, Isao Echizen

专题命中 安全评测 :safety(abstract)

AI总结 针对多轮智能体AI对话中AI协助的人际操纵问题,构建含1000个五轮对话的基准数据集,提出HRGuard双门控机制,在8个模型上表现优于通用安全方案,可减少有害顺从并保留保护指导。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25210 2026-08-27 cs.DB 新提交 50%

Bolt-on, Verifiable Provenance for LLM-Powered Data Processing

用于大语言模型驱动数据处理的外挂式可验证来源

Yiming Lin, Sepanta Zeighami, Aditya G. Parameswaran

专题命中 安全评测 :trustworthy(abstract)

AI总结 针对LLM黑盒无法提供答案来源及可信度的问题,提出外挂式框架BLIP,可高效生成小尺寸可验证来源,在七个数据集上准确率较最优基线高30%以上且成本低。

详情

展开后加载摘要…

URL PDF HTML 收藏