arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-04 至 2025-08-04 共收录 20 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 2 篇

2409.18203 2025-08-04 cs.HC cs.AI cs.CL cs.LG 75%

Policy Maps: Tools for Guiding the Unbounded Space of LLM Behaviors

Michelle S. Lam, Fred Hohman, Dominik Moritz, Jeffrey P. Bigham, Kenneth Holstein, Mary Beth Kery

机构 * Stanford University(斯坦福大学) Apple(苹果公司) Carnegie Mellon University(卡内基梅隆大学)

专题命中 安全训练 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments UIST 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24025 2025-08-04 cs.CV cs.AI 57%

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models

Chenbin Pan, Wenbin He, Zhengzhong Tu, Liu Ren

机构 * Bosch Research North America(博世北美研究部) Bosch Center for Artificial Intelligence (BCAI)(博世人工智能中心) Texas A&M University(德克萨斯A&M大学)

专题命中 安全训练 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 越狱攻击 1 篇

2506.15170 2025-08-04 cs.CR 78%

From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem

Yanxu Mao, Tiehan Cui, Peipei Liu, Datao You, Hongsong Zhu

专题命中 越狱攻击 :jailbreak(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 提示注入 1 篇

2508.00602 2025-08-04 cs.CR cs.AI cs.LG 84%

LeakSealer: A Semisupervised Defense for LLMs Against Prompt Injection and Leakage Attacks

Francesco Panebianco, Stefano Bonfanti, Francesco Trovò, Michele Carminati

机构 * Politecnico di Milano(米兰理工学院)

专题命中 提示注入 :prompt injection(title,abstract);jailbreak(abstract);分类 cs.AI、cs.LG

Comments 22 pages, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与事实性 1 篇

2505.23628 2025-08-04 cs.CL cs.AI 62%

AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale Corpora

Jiaxin Bai, Wei Fan, Qi Hu, Qing Zong, Chunyang Li, Hong Ting Tsang, Hongyu Luo, Yauwai Yim, Haoyu Huang, Xiao Zhou, Feng Qin, Tianshi Zheng, Xi Peng, Xin Yao, Huiwen Yang, Leijie Wu, Yi Ji, Gong Zhang, Renhai Chen, Yangqiu Song

机构 * CSE, HKUST(香港科技大学计算机科学与工程系) CSE, CUHK(香港城市大学计算机科学与工程系) Theory Lab, Huawei(华为理论实验室)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

Comments 9 pages, preprint, code: https://github.com/HKUST-KnowComp/AutoSchemaKG

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 隐私与版权 1 篇

2507.21170 2025-08-04 cs.CR cs.AI cs.CL 62%

OneShield -- the Next Generation of LLM Guardrails

Chad DeLuca, Anna Lisa Gentile, Shubhi Asthana, Bing Zhang, Pawan Chowdhary, Kellen Cheng, Basel Shbita, Pengyuan Li, Guang-Jie Ren, Sandeep Gopisetty

机构 * IBM Research(IBM研究院) Princeton University(普林斯顿大学)

专题命中 隐私与版权 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 4 篇

2508.00673 2025-08-04 cs.CL 74%

MELAC: Massive Evaluation of Large Language Models with Alignment of Culture in Persian Language

Farhan Farsi, Farnaz Aghababaloo, Shahriar Shariati Motlagh, Parsa Ghofrani, MohammadAli SadraeiJavaheri, Shayan Bali, Amirhossein Shabani, Farbod Bijary, Ghazal Zamaninejad, AmirMohammad Salehoof, Saeedeh Momtazi

机构 * Amirkabir University of Technology(阿姆irkabir技术大学) Part AI Research Center(Part人工智能研究中心) University of Mazandaran(马赞德兰大学) King’s College London(伦敦国王学院)

专题命中 安全评测 :alignment(title);分类 cs.CL

Comments Preprint. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00081 2025-08-04 cs.AI 57%

Rethinking Evidence Hierarchies in Medical Language Benchmarks: A Critical Evaluation of HealthBench

Fred Mutisya, Shikoh Gitau, Nasubo Ongoma, Keith Mbae, Elizabeth Wamicha

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18004 2025-08-04 cs.AI 57%

E.A.R.T.H.: Structuring Creative Evolution through Model Error in Generative AI

Yusen Peng, Shuhua Mao

机构 * University of Warwick(沃里克大学) Wuhan University of Technology(武汉理工大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 44 pages,11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17130 2025-08-04 cs.AI 57%

Chain-of-Trust: A Progressive Trust Evaluation Framework Enabled by Generative AI

Botao Zhu, Xianbin Wang, Lei Zhang, Xuemin, Shen

机构 * Department of Electrical and Computer Engineering, Western University(西方大学电气与计算机工程系) James Watt School of Engineering, University of Glasgow(格拉斯哥大学詹姆斯·瓦特工程学院) Department of Electrical and Computer Engineering, University of Waterloo(滑铁卢大学电气与计算机工程系)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Journal ref IEEE Network, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

7. AI治理与伦理 2 篇

2503.06072 2025-08-04 cs.CL cs.AI 62%

A Survey on Post-training of Large Language Models

Guiyao Tie, Zeli Zhao, Dingjie Song, Fuyang Wei, Rong Zhou, Yurou Dai, Wen Yin, Zhejian Yang, Jiangyue Yan, Yao Su, Zhenhan Dai, Yifeng Xie, Yihan Cao, Lichao Sun, Pan Zhou, Lifang He, Hechang Chen, Yu Zhang, Qingsong Wen, Tianming Liu, Neil Zhenqiang Gong, Jiliang Tang, Caiming Xiong, Heng Ji, Philip S. Yu, Jianfeng Gao

机构 * Huazhong University of Science and Technology(华中科技大学) Lehigh University(莱斯大学) The University of Hong Kong(香港大学) Jilin University(吉林大学) Southern University of Science and Technology(南方科技大学) Worcester Polytechnic Institute(沃思堡理工学院) LinkedIn Corporation(领英公司) Squirrel Ai Learning University of Georgia(佐治亚大学) Duke University(杜克大学) Michigan State University(密歇根州立大学) Salesforce Research(Salesforce研究) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Illinois at Chicago(伊利诺伊大学芝加哥分校) Microsoft Research(微软研究院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments 87 pages, 21 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23454 2025-08-04 cs.HC cs.CY cs.ET cs.GR q-bio.NC 57%

Breaking the mould of Social Mixed Reality - State-of-the-Art and Glossary

Marta Bieńkiewicz, Julia Ayache, Panayiotis Charalambous, Cristina Becchio, Marco Corragio, Bertram Taetz, Francesco De Lellis, Antonio Grotta, Anna Server, Daniel Rammer, Richard Kulpa, Franck Multon, Azucena Garcia-Palacios, Jessica Sutherland, Kathleen Bryson, Stéphane Donikian, Didier Stricker, Benoît Bardy

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 其他安全 8 篇

2508.00741 2025-08-04 cs.CL cs.AI 73%

Out-of-Context Abduction: LLMs Make Inferences About Procedural Data Leveraging Declarative Facts in Earlier Training Data

Sohaib Imran, Rob Lamb, Peter M. Atkinson

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00674 2025-08-04 cs.AI cs.HC cs.LG 62%

Context-Aware Visualization for Explainable AI Recommendations in Social Media: A Vision for User-Aligned Explanations

Banan Alkhateeb, Ellis Solaiman

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00574 2025-08-04 cs.CL cs.AI 62%

SynAdapt: Learning Adaptive Reasoning in Large Language Models via Synthetic Continuous Chain-of-Thought

Jianwei Wang, Ziming Wu, Fuming Lai, Shaobing Lian, Ziqian Zeng

机构 * Tencent Inc.(腾讯公司)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19651 2025-08-04 cs.LG cs.CL 62%

Unlocking Multi-Modal Potentials for Link Prediction on Dynamic Text-Attributed Graphs

Yuanyuan Xu, Wenjie Zhang, Ying Zhang, Xuemin Lin, Xiwei Xu

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00395 2025-08-04 cs.CV cs.AI 57%

Decouple before Align: Visual Disentanglement Enhances Prompt Tuning

Fei Zhang, Tianfei Zhou, Jiangchao Yao, Ya Zhang, Ivor W. Tsang, Yanfeng Wang

机构 * Cooperative Medianet Innovation Center, Shanghai Jiao Tong University(合作中位网创新中心,上海交通大学) School of Artificial Intelligence, Shanghai Jiao Tong University(人工智能学院,上海交通大学) Shanghai Innovation Institute(上海创新研究院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Beijing Institute of Technology(北京理工大学) A*STAR Centre for Frontier AI Research(A*STAR前沿人工智能研究中心)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 16 pages, Accepted at IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00769 2025-08-04 physics.bio-ph q-bio.TO 50%

Designing cultured tissue moulds using evolutionary strategies

Allison E. Andrews, Hugh Dickinson, James P. Hague

专题命中 其他安全 :alignment(abstract)

Comments 9 pages, 7 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00650 2025-08-04 physics.soc-ph 50%

Evac-Cast: An Interpretable Machine-Learning Framework for Evacuation Forecasts Across Hurricanes and Wildfires

Bo Li, Chenyue Liu, Ali Mostafavi

专题命中 其他安全 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23859 2025-08-04 astro-ph.IM 50%

radio-llava: Advancing Vision-Language Models for Radio Astronomical Source Analysis

S. Riggi, T. Cecconello, A. Pilzer, S. Palazzo, N. Gupta, A. M. Hopkins, C. Trigilio, G. Umana

专题命中 其他安全 :alignment(abstract)

Comments 19 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏