arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9346 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9346 篇

2506.23115 2025-07-01 cs.CV cs.AI cs.CL 62%

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Haonan Chen, Hong Liu, Yuping Luo, Liang Wang, Nan Yang, Furu Wei, Zhicheng Dou

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) Stanford University(斯坦福大学) Microsoft Corporation(微软公司)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Homepage: https://haon-chen.github.io/MoCa/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21578 2025-06-30 cs.CL cs.AI 62%

HealthQA-BR: A System-Wide Benchmark Reveals Critical Knowledge Gaps in Large Language Models

Andrew Maranhão Ventura D'addario

机构 * Independent Researcher(独立研究者)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19468 2025-06-25 cs.CL cs.AI 62%

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages

Wenhan Han, Yifan Zhang, Zhixun Chen, Binbin Liu, Haobin Lin, Bingni Zhang, Taifeng Wang, Mykola Pechenizkiy, Meng Fang, Yin Zheng

机构 * Eindhoven University of Technology(埃因霍温理工大学) ByteDance(字节跳动) University of Liverpool(利物浦大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19352 2025-06-25 cs.CL cs.AI cs.HC 62%

Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation

Jisu Shin, Juhyun Oh, Eunsu Kim, Hoyun Song, Alice Oh

机构 * School of Computing(计算机学院) Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Findings of ACL 2025; github repo: https://github.com/ddindidu/atomic-persona-evaluation/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19191 2025-06-25 cs.AI cs.CL cs.GT math.LO 62%

Bayesian Evolutionary Swarm Architecture: A Formal Epistemic System Grounded in Truth-Based Competition

Craig Steven Wright

机构 * Department of Computer Science University of Exeter(计算机科学系大学埃克塞特)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 83 pages, 14 sections, 92 formal results, no prior conference publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17111 2025-06-23 cs.AI cs.CL 62%

Are Bias Evaluation Methods Biased ?

Lina Berrayana, Sean Rooney, Luis Garcés-Erice, Ioana Giurgiu

机构 * First Author Affiliation(第一作者机构) Second Author Affiliation(第二作者机构)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted to ACL 2025 Workshop GEM

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15699 2025-06-23 cs.LG cs.AI 62%

BLUR: A Benchmark for LLM Unlearning Robust to Forget-Retain Overlap

Shengyuan Hu, Neil Kale, Pratiksha Thaker, Yiwei Fu, Steven Wu, Virginia Smith

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14002 2025-06-18 cs.LG cs.AI cs.IT math.IT stat.ML 62%

Taming Polysemanticity in LLMs: Provable Feature Recovery via Sparse Autoencoders

Siyu Chen, Heejune Sheen, Xuyuan Xiong, Tianhao Wang, Zhuoran Yang

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 136 pages, 21 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13798 2025-06-18 cs.CY cs.AI 62%

Contemporary AI foundation models increase biological weapons risk

Roger Brent, T. Greg McKelvey

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY

Comments 58 pages, 10 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13313 2025-06-17 cs.CL cs.AI econ.GN q-fin.EC 62%

Large Language Models as 'Hidden Persuaders': Fake Product Reviews are Indistinguishable to Humans and Machines

Weiyao Meng, John Harvey, James Goulding, Chris James Carter, Evgeniya Lukinova, Andrew Smith, Paul Frobisher, Mina Forrest, Georgiana Nica-Avram

机构 * N/LAB, Nottingham University Business School, University of Nottingham, UK(N/LAB,诺丁汉大学商学院,诺丁汉大学,英国) Haydn Green Institute for Entrepreneurship and Innovation, University of Nottingham, UK(创业与创新研究院,诺丁汉大学,英国) Strategic Innovation Ltd, UK(战略创新有限公司,英国)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12839 2025-06-17 stat.ML cs.AI cs.LG 62%

Fair Bayesian Model-Based Clustering

Jihu Lee, Kunwoong Kim, Yongdai Kim

机构 * Department of Statistics(统计系)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05284 2025-06-17 cs.AI cs.LG 62%

Can a Bayesian Oracle Prevent Harm from an Agent?

Yoshua Bengio, Michael K. Cohen, Nikolay Malkin, Matt MacDermott, Damiano Fornasiere, Pietro Greiner, Younesse Kaddar

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted at UAI 2025 (Uncertainty in Artificial Intelligence). 20 pages, 2 figures. Code available at: https://github.com/saifh-github/conservative-bayesian-public

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11515 2025-06-16 cs.CV cs.CL cs.LG 62%

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs

Xiao Xu, Libo Qin, Wanxiang Che, Min-Yen Kan

机构 * Research Center for Social Computing and Information Retrieval, Harbin Institute of Technology(社会计算与信息检索研究中心,哈尔滨工业大学) School of Computer Science and Engineering, Central South University(计算机科学与工程学院,中南大学) School of Computing, National University of Singapore(计算学院,新加坡国立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted by IEEE Transactions on Circuits and Systems for Video Technology (TCSVT). June 2025. DOI: https://doi.org/10.1109/TCSVT.2025.3578266

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10885 2025-06-13 cs.CL cs.AI 62%

Slimming Down LLMs Without Losing Their Minds

Qingda, Mai

机构 * University of Waterloo(滑铁卢大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.06104 2025-06-12 cs.LG cs.AI 62%

Gradient Aligned Regression via Pairwise Losses

Dixian Zhu, Tianbao Yang, Livnat Jerby

机构 * Department of Genetics, Stanford University, CA, USA(遗传学系,斯坦福大学,加州,美国) Department of Computer Science, Texas A\&M University, TX, USA(计算机科学系,德克萨斯A&M大学,德克萨斯州,美国)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments ICML 2025; 23 pages, 12 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09160 2025-06-11 cs.AI cs.LG cs.RO 62%

Innate-Values-driven Reinforcement Learning based Cognitive Modeling

Qin Yang

机构 * Intelligent Social Systems and Swarm Robotics Lab (IS$^3$R)(智能社会系统与群体机器人实验室) Computer Science and Information Systems Department(计算机科学与信息系统系) Bradley University(布拉德利大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments The paper had been accepted by the 2025 IEEE Conference on Cognitive and Computational Aspects of Situation Management (CogSIMA). arXiv admin note: text overlap with arXiv:2401.05572

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06528 2025-06-10 cs.CL cs.AI cs.HC 62%

Epistemic Integrity in Large Language Models

Bijean Ghafouri, Shahrad Mohammadzadeh, James Zhou, Pratheeksha Nair, Jacob-Junqi Tian, Hikaru Tsujimura, Mayank Goel, Sukanya Krishna, Reihaneh Rabbany, Jean-François Godbout, Kellin Pelrine

机构 * University of Southern California(南加州大学) McGill University(麦吉尔大学) UC Berkeley(加州大学伯克利分校) Vector Institute(向量研究所) Cardiff University(卡迪夫大学) University College London(伦敦大学学院) IIIT Hyderabad(海得拉巴印度理工学院) Harvard University(哈佛大学) Université de Montréal(蒙特利尔大学) Mila

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06737 2025-06-10 cs.CL cs.AI 62%

C-PATH: Conversational Patient Assistance and Triage in Healthcare System

Qi Shi, Qiwei Han, Cláudia Soares

机构 * School of Business and Economics(商业与经济学院) School of Science and Technology(科学与技术学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted in IEEE ICDH 2025, 10 pages, 8 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06020 2025-06-09 cs.CL cs.AI 62%

When to Trust Context: Self-Reflective Debates for Context Reliability

Zeqi Zhou, Fang Wu, Shayan Talaei, Haokai Zhao, Cheng Meixin, Tinson Xu, Amin Saberi, Yejin Choi

机构 * Brown University(布朗大学) Stanford University(斯坦福大学) University of New South Wales(新南威尔士大学) Xi’an University of Electronic Science and Technology(西安电子科技大学) University of Chicago(芝加哥大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04474 2025-06-06 cs.LG cs.AI 62%

Classifying Dental Care Providers Through Machine Learning with Features Ranking

Mohammad Subhi Al-Batah, Mowafaq Salem Alzboon, Muhyeeddin Alqaraleh, Mohammed Hasan Abu-Arqoub, Rashiq Rafiq Marie

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Journal ref Data and Metadata. 2025; 4:755

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01782 2025-06-03 cs.CY cs.AI cs.SY eess.SY 62%

Systematic Hazard Analysis for Frontier AI using STPA

Simon Mylius

机构 * Simon Mylius

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY

Comments 29 pages, 5 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15865 2025-06-03 q-fin.GN cs.AI cs.CL 62%

Standard Benchmarks Fail -- Auditing LLM Agents in Finance Must Prioritize Risk

Zichen Chen, Jiaao Chen, Jianda Chen, Misha Sra

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments 46 pages, 2 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01206 2025-06-03 cs.CL cs.AI 62%

Mamba Drafters for Speculative Decoding

Daewon Choi, Seunghyuk Oh, Saket Dingliwal, Jihoon Tack, Kyuyoung Kim, Woomin Song, Seojin Kim, Insu Han, Jinwoo Shin, Aram Galstyan, Shubham Katiyar, Sravan Babu Bodapati

机构 * KAIST(韩国科学技术院) Amazon AGI(亚马逊人工智能实验室) Seoul National University(首尔国立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14301 2025-06-03 cs.CL cs.AI 62%

SEA-HELM: Southeast Asian Holistic Evaluation of Language Models

Yosephine Susanto, Adithya Venkatadri Hulagadri, Jann Railey Montalan, Jian Gang Ngui, Xian Bin Yong, Weiqi Leong, Hamsawardhini Rengarajan, Peerat Limkonchotiwat, Yifan Mai, William Chandra Tjhi

机构 * AI Singapore(AI新加坡) National University of Singapore(国立新加坡大学) Center for Research on Foundation Models (CRFM)(基础模型研究中心) Stanford University(斯坦福大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06560 2025-06-03 cs.CL cs.CY 62%

Position: It's Time to Act on the Risk of Efficient Personalized Text Generation

Eugenia Iofinova, Andrej Jovanovic, Dan Alistarh

机构 * Institute of Science and Technology Austria(奥地利科学与技术研究所)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00113 2025-06-03 cs.CL cs.AI 62%

Wait, that's not an option: LLMs Robustness with Incorrect Multiple-Choice Options

Gracjan Góral, Emilia Wiśnios, Piotr Sankowski, Paweł Budzianowski

机构 * University of Warsaw(华沙大学) Institute of Mathematics, Polish Academy of Sciences(波兰科学院数学研究所) MIM Solutions(MIM解决方案) K-Scale Labs(K-Scale实验室) IDEAS NCBR IDEAS Research Institute(IDEAS研究学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted for ACL 2025 Main Conference and NeurIPS 2024 FM-EduAssess Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00047 2025-06-03 cs.CY cs.AI cs.CE 62%

Risks of AI-driven product development and strategies for their mitigation

Jan Göpfert, Jann M. Weinand, Patrick Kuckertz, Noah Pflugradt, Jochen Linßen

机构 * Institute of Climate and Energy Systems, Jülich Systems Analysis (ICE-2), Forschungszentrum Jülich(气候与能源系统研究所,朱利奇系统分析(ICE-2),朱利奇研究中心) Chair for Fuel Cells, Faculty of Mechanical Engineering, RWTH Aachen University(燃料电池主任,机械工程学院,亚琛工业大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14956 2025-06-03 cs.CL cs.AI cs.IR 62%

ExPerT: Effective and Explainable Evaluation of Personalized Long-Form Text Generation

Alireza Salemi, Julian Killingback, Hamed Zamani

机构 * Center for Intelligent Information Retrieval(智能信息检索中心) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12583 2025-06-02 cs.RO cs.AI cs.LG 62%

A Comprehensive Survey on Physical Risk Control in the Era of Foundation Model-enabled Robotics

Takeshi Kojima, Yaonan Zhu, Yusuke Iwasawa, Toshinori Kitamura, Gang Yan, Shu Morikuni, Ryosuke Takanami, Alfredo Solano, Tatsuya Matsushima, Akiko Murakami, Yutaka Matsuo

机构 * The University of Tokyo(东京大学) Japan AI Safety Institute(日本人工智能安全研究所)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted to IJCAI 2025 Survey Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.09429 2025-06-02 cs.LG cs.CL cs.CV 62%

Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models

Kening Zheng, Junkai Chen, Yibo Yan, Xin Zou, Xuming Hu

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Guangxi Zhuang Autonomous Region Big Data Research Institute(广西壮族自治区大数据研究院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.LG

Comments Accepted by Findings of ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏