arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8017 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8017 篇

2604.04942 2026-04-08 cs.CL cs.AI 76%

TDA-RC: Task-Driven Alignment for Knowledge-Based Reasoning Chains in Large Language Models

TDA-RC:基于任务驱动的知识推理链对齐方法

Jiaquan Zhang, Qigan Sun, Chaoning Zhang, Xudong Wang, Zhenzhen Huang, Yitian Zhou, Pengcheng Zheng, Chi-lok Andy Tai, Sung-Ho Bae, Zeyu Ma, Caiyan Qin, Jinyu Guo, Yang Yang, Hengtao Shen

机构 * School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院) School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院) College of Professional and Continuing Education, The Hong Kong Polytechnic University(香港理工大学专业及持续教育学院) School of Robotics and Advanced Manufacture, Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)机电工程与自动化学院) School of Computing, Kyung Hee University(庆熙大学计算机学院) School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

AI总结 本文提出TDA-RC方法,通过拓扑学优化提升大语言模型推理效率与准确性,结合持久同调将不同推理范式统一到拓扑空间中,实现高效且精准的推理链优化。

Comments 14 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19463 2026-04-08 cs.CY cs.AI cs.SI 76%

Hedging and Non-Affirmation: Quantifying LLM Alignment on Questions of Human Rights

对冲与非肯定:量化大语言模型在人权问题上的对齐

Rafiya Javed, Cassandra Parent, Jackie Kay, David Yanni, Abdullah Zaini, Anushe Sheikh, Maribeth Rauh, Walter Gerych, Ramona Comanescu, Iason Gabriel, Marzyeh Ghassemi, Laura Weidinger

机构 * Google Deepmind(谷歌DeepMind) Massachusetts Institute of Technology(麻省理工学院) Independent Researcher(独立研究员) Google(谷歌) AI Accountability Lab, Trinity College Dublin(都柏林圣三一学院人工智能问责实验室)

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.CY

AI总结 研究通过系统框架量化LLM在不同群体身份上的对冲与非肯定行为,发现群体身份是主要影响因素,通过引导和正交化技术可有效缓解偏差。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16872 2026-03-19 cs.CL cs.CY 76%

Trust, Safety, and Accuracy: Assessing LLMs for Routine Maternity Advice

信任、安全与准确性:评估LLMs用于常规产科咨询

V Sai Divya, A Bhanusree, Rimjhim, K Venkata Krishna Rao

机构 * National Institute of Technology, Warangal, India(印度战争格尔国家理工学院)

专题命中 其他安全 :safety(title);分类 cs.CL、cs.CY

AI总结 本研究评估了LLMs在提供可靠且易懂的产科信息方面的表现,发现Perplexity在语义上与专家接近,而ChatGPT-4o在文本清晰度和医学术语使用上更优,为偏远地区产科教育提供了可行的AI解决方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09566 2026-03-12 cs.AI cs.LG 76%

Toward Closed-loop Molecular Discovery via Language Model, Property Alignment and Strategic Search

通过语言模型、性质对齐和战略搜索实现闭环分子发现

Junkai Ji, Zhangfan Yang, Dong Xu, Ruibin Bai, Jianqiang Li, Tingjun Hou, Zexuan Zhu

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

AI总结 Trio结合语言模型、性质对齐和战略搜索,实现高效且可解释的闭环分子设计,提升药物配体的结合亲和力、药物性和合成可及性。

Comments 30 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06123 2026-01-13 cs.LG cs.AI 76%

Latent Space Communication via K-V Cache Alignment

通过K-V缓存对齐实现潜在空间通信

Lucio M. Dery, Zohar Yahav, Henry Prior, Qixuan Feng, Jiajun Shen, Arthur Szlam

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

AI总结 本文提出通过学习共享表示空间对齐多模型的k-v缓存,实现模型间高效协作与知识共享,提升整体性能与能力转移。

Comments 15 pages, 6 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24478 2026-01-06 cs.LG cs.AI stat.ME 76%

HOLOGRAPH: Active Causal Discovery via Sheaf-Theoretic Alignment of Large Language Model Priors

HOLOGRAPH:通过sheaf理论对大型语言模型先验进行对齐以实现主动因果发现

Hyunjun Kim

机构 * Korea Advanced Institute of Science \'Ecole Polytechnique F\'ed\'erale de Lausanne (EPFL), Lausanne, Switzerland

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

AI总结 HOLOGRAPH通过sheaf理论对齐大型语言模型先验,实现主动因果发现,提供严谨的数学基础并实现竞争性性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00115 2025-12-30 cs.CL cs.AI 76%

Cognitive Alignment in Personality Reasoning: Leveraging Prototype Theory for MBTI Inference

人格推理中的认知对齐:利用原型理论进行MBTI推断

Haoyuan Li, Yuanbo Tong, Yuchen Li, Zirui Wang, Chunhou Liu, Jiamou Liu

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

AI总结 ProtoMBTI通过原型理论与认知对齐,提升文本人格推断的准确性和泛化能力。

Comments The authors have decided to withdraw this version to substantially revise and extend the work

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15545 2025-12-30 cs.CL cs.AI 76%

TokenTiming: A Dynamic Alignment Method for Universal Speculative Decoding Model Pairs

TokenTiming: 一种适用于通用推测解码模型对的动态对齐方法

Sibo Xiao, Jinyuan Fu, Zhongle Xie, Lidan Shou

机构 * Zhejiang University(浙江大学)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

AI总结 TokenTiming通过动态对齐方法实现通用推测解码,无需重新训练即可处理不同词汇模型,提升LLM推理效率1.57倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04834 2025-11-13 cs.LG cs.AI cs.CV 76%

Prompt-Based Safety Guidance Is Ineffective for Unlearned Text-to-Image Diffusion Models

Jiwoo Shin, Byeonghu Na, Mina Kang, Wonhyeok Choi, Il-Chul Moon

机构 * KAIST(韩国科学技术院)

专题命中 其他安全 :safety(title);分类 cs.AI、cs.LG

Comments Accepted at NeurIPS 2025 Workshop on Generative and Protective AI for Content Creation

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14158 2025-11-11 cs.CL cs.LG 76%

Temporal Alignment of Time Sensitive Facts with Activation Engineering

Sanjay Govindan, Maurice Pagnucco, Yang Song

机构 * University of New South Wales(新南威尔士大学)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06164 2025-11-04 cs.LG cs.AI 76%

Model Alignment Search

Satchel Grant

机构 * Stanford University(斯坦福大学)

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11018 2025-10-21 cs.CL cs.AI 76%

GRIFFIN: Effective Token Alignment for Faster Speculative Decoding

Shijing Hu, Jingyang Li, Xingyu Xie, Zhihui Lu, Kim-Chuan Toh, Pan Zhou

机构 * Fudan University(复旦大学) National University of Singapore(新加坡国立大学) Singapore Management University(新加坡管理学院)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21857 2025-10-15 cs.CV cs.AI cs.LG 76%

SPADE: Spatial Transcriptomics and Pathology Alignment Using a Mixture of Data Experts for an Expressive Latent Space

Ekaterina Redekop, Mara Pleasure, Zichen Wang, Kimberly Flores, Anthony Sisk, William Speier, Corey W. Arnold

机构 * Biomedical AI Research Lab, University of California, Los Angeles(生物医学人工智能研究实验室,加州大学洛杉矶分校) Department of Pathology, University of California, Los Angeles(病理学系,加州大学洛杉矶分校) Department of Radiology, University of California, Los Angeles(放射学系,加州大学洛杉矶分校) Department of Bioengineering, University of California, Los Angeles(生物工程系,加州大学洛杉矶分校) UCLA Medical Informatics Home Area, University of California, Los Angeles(UCLA医学信息学家庭区域,加州大学洛杉矶分校) Department of Computational Medicine, University of California, Los Angeles(计算医学系,加州大学洛杉矶分校)

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16633 2025-09-23 cs.CV cs.AI cs.CL 76%

When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs

Abhirama Subramanyam Penamakuri, Navlika Singh, Piyush Arora, Anand Mishra

机构 * Indian Institute of Technology Jodhpur(印度理工学院朱诺尔)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

Comments Accepted to EMNLP (Main) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10057 2025-08-15 q-bio.NC cs.AI cs.CL 76%

Large Language Models Show Signs of Alignment with Human Neurocognition During Abstract Reasoning

Christopher Pinier, Sonia Acuña Vargas, Mariia Steeghs-Turchina, Dora Matzke, Claire E. Stevenson, Michael D. Nunez

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

Comments Presented at the 8th Annual Conference on Cognitive Computational Neuroscience (August 12-15, 2025; Amsterdam, The Netherlands); 20 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07525 2025-07-23 cs.CV cs.AI cs.LG 76%

RadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment

Difei Gu, Yunhe Gao, Yang Zhou, Mu Zhou, Dimitris Metaxas

机构 * Rutgers University(罗格斯大学) Stanford University(斯坦福大学)

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

Comments Accepted to MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08665 2025-07-14 cs.CL cs.AI 76%

KELPS: A Framework for Verified Multi-Language Autoformalization via Semantic-Syntactic Alignment

Jiyao Zhang, Chengli Zhong, Hui Xu, Qige Li, Yi Zhou

机构 * School of Information Science and Technology, University of Science and Technology of China(信息科学与技术学院,中国科学技术大学) USTC Knowledge Computing Lab(中国科学技术大学知识计算实验室)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

Comments Accepted by the ICML 2025 AI4MATH Workshop. 22 pages, 16 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01069 2025-06-03 cs.CV cs.AI cs.LG 76%

Revolutionizing Blood Banks: AI-Driven Fingerprint-Blood Group Correlation for Enhanced Safety

Malik A. Altayar, Muhyeeddin Alqaraleh, Mowafaq Salem Alzboon, Wesam T. Almagharbeh

专题命中 其他安全 :safety(title);分类 cs.AI、cs.LG

Journal ref Data and Metadata [Internet]. 2025 Apr. 7 [cited 2025 Jun. 1];4:894

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24040 2025-06-02 cs.CL cs.AI 76%

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering

Yuexing Hao, Kumail Alhamoud, Hyewon Jeong, Haoran Zhang, Isha Puri, Philip Torr, Mike Schaekermann, Ariel D. Stern, Marzyeh Ghassemi

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14468 2025-04-22 cs.CL cs.LG eess.SP q-bio.NC 76%

sEEG-based Encoding for Sentence Retrieval: A Contrastive Learning Approach to Brain-Language Alignment

Yijun Liu

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.LG

Comments Accepted for poster presentation at the CVPR 2025 Workshop on Multimodal Foundation Models (MMFM3)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13825 2025-04-21 cs.CL cs.LG 76%

Feature Alignment and Representation Transfer in Knowledge Distillation for Large Language Models

Junjie Yang, Junhao Song, Xudong Han, Ziqian Bi, Tianyang Wang, Chia Xin Liang, Xinyuan Song, Yichao Zhang, Qian Niu, Benji Peng, Keyu Chen, Ming Liu

机构 * Pingtan Research Institute of Xiamen University(厦门大学滨海研究院) Imperial College London(伦敦帝国理工学院) University of Sussex(苏塞克斯大学) Purdue University(普渡大学) University of Liverpool(利物浦大学) JTB Technology Corp.(JTB技术公司) Emory University(埃默里大学) The University of Texas at Dallas(德克萨斯大学达拉斯分校) Kyoto University(京都大学) AppCubic Georgia Institute of Technology(佐治亚理工学院) AI Agent Lab(AI代理实验室)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05803 2025-04-18 cs.LG cs.AI cs.CV math.ST stat.TH 76%

Test-time Alignment of Diffusion Models without Reward Over-optimization

Sunwoo Kim, Minkyu Kim, Dongmin Park

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

Comments ICLR 2025 (Spotlight). The Thirteenth International Conference on Learning Representations. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12900 2025-02-19 cs.CL cs.AI cs.SD 76%

Soundwave: Less is More for Speech-Text Alignment in LLMs

Yuhao Zhang, Zhiheng Liu, Fan Bu, Ruiyu Zhang, Benyou Wang, Haizhou Li

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11205 2025-02-18 cs.LG cs.CY 76%

Deep Contrastive Learning for Feature Alignment: Insights from Housing-Household Relationship Inference

Xiao Qian, Shangjia Dong, Rachel Davidson

专题命中 其他安全 :alignment(title);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18539 2025-01-31 cs.CL cs.AI cs.IR 76%

Can we Retrieve Everything All at Once? ARM: An Alignment-Oriented LLM-based Retrieval Method

Peter Baile Chen, Yi Zhang, Michael Cafarella, Dan Roth

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12273 2025-01-22 cs.CL cs.AI 76%

Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement

Maosong Cao, Taolin Zhang, Mo Li, Chuyu Zhang, Yunxin Liu, Haodong Duan, Songyang Zhang, Kai Chen

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

Comments Tech Report. Github: https://github.com/InternLM/Condor

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06846 2024-12-11 cs.LG cs.AI 76%

Classifier-free guidance in LLMs Safety

Roman Smirnov

专题命中 其他安全 :safety(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.01101 2024-11-19 cs.LG cs.AI 76%

Feature Alignment: Rethinking Efficient Active Learning via Proxy in the Context of Pre-trained Models

Ziting Wen, Oscar Pizarro, Stefan Williams

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

Comments Accepted by Transactions on Machine Learning Research (TMLR, 2024) https://openreview.net/forum?id=PNcgJMJcdl

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10852 2024-10-16 cs.CL cs.AI 76%

SafeLLM: Domain-Specific Safety Monitoring for Large Language Models: A Case Study of Offshore Wind Maintenance

Connor Walker, Callum Rothon, Koorosh Aslansefat, Yiannis Papadopoulos, Nina Dethlefs

专题命中 其他安全 :safety(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16645 2024-09-26 cs.LG cs.AI 76%

Task Addition in Multi-Task Learning by Geometrical Alignment

Soorin Yim, Dae-Woong Jeong, Sung Moon Ko, Sumin Lee, Hyunseung Kim, Chanhui Lee, Sehui Han

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

Comments 11 pages, 5 figures, Accepted at AI for Science Workshop at 41st International Conference on Machine Learning

详情

展开后加载摘要…

URL PDF HTML 收藏