arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-14 至 2025-10-14 共收录 26 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 26 篇

2505.20333 2025-10-14 cs.CL cs.AI 81%

Multi-Scale Manifold Alignment for Interpreting Large Language Models: A Unified Information-Geometric Framework

Yukun Zhang, Qi Dong

机构 * The Chinese University of Hong Kong(香港中文大学) Fudan University(复旦大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05129 2025-10-14 cs.CL cs.LG 81%

Automated Alignment of Math Items to Content Standards in Large-Scale Assessments Using Language Models

Qingshu Xu, Hong Jiao, Tianyi Zhou, Ming Li, Nan Zhang, Sydney Peters, Yanbin Fu

机构 * University of Maryland, College Park(马里兰大学学院公园分校) Shandong Jiaotong University(山东交通大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03227 2025-10-14 cs.LG cs.AI 81%

Bridging Neural ODE and ResNet: A Formal Error Bound for Safety Verification

Abdelrahman Sayed Sayed, Pierre-Jean Meyer, Mohamed Ghazel

机构 * Univ Gustave Eiffel(法国埃菲尔大学)

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

Comments 18 pages, 5 figures, Accepted for publication in the proceedings of the 8th International Symposium on AI Verification SAIV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10732 2025-10-14 cs.CY 79%

When Openness Fails: Lessons from System Safety for Assessing Openness in AI

Tamara Paris, Shalaleh Rismani

专题命中 其他安全 :safety(title,abstract);分类 cs.CY

Comments Accepted to Symposium on Model Accountability, Sustainability and Healthcare (SMASH) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10280 2025-10-14 cs.CL 79%

On the Entity-Level Alignment in Crosslingual Consistency

Yihong Liu, Mingyang Wang, François Yvon, Hinrich Schütze

机构 * Center for Information and Language Processing, LMU Munich(信息与语言处理中心,慕尼黑大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心(MCML)) Sorbonne Université, CNRS, ISIR, France(索邦大学,CNRS,ISIR,法国)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26431 2025-10-14 cs.CL 79%

Text-Based Approaches to Item Alignment to Content Standards in Large-Scale Reading & Writing Tests

Yanbin Fu, Hong Jiao, Tianyi Zhou, Nan Zhang, Ming Li, Qingshu Xu, Sydney Peters, Robert W. Lissitz

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments need updates

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11040 2025-10-14 cs.CL 70%

Enabling Doctor-Centric Medical AI with LLMs through Workflow-Aligned Tasks and Benchmarks

Wenya Xie, Qingying Xiao, Yu Zheng, Xidong Wang, Junying Chen, Ke Ji, Anningzhe Gao, Prayag Tiwari, Xiang Wan, Feng Jiang, Benyou Wang

机构 * Shenzhen Research Institute of Big Data(大数据研究 institute) National Health Data Institute(国家健康数据研究所) Halmstad University(哈马碧大学) Shenzhen University of Advanced Technology(深圳先进技术大学)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11693 2025-10-14 cs.CL cs.AI cs.CV 62%

Scaling Language-Centric Omnimodal Representation Learning

Chenghao Xiao, Hou Pong Chan, Hao Zhang, Weiwen Xu, Mahani Aljunied, Yu Rong

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23703 2025-10-14 cs.AI cs.CL 62%

Let's Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM's Math Capability

Ruida Wang, Yuxin Li, Yi R. Fung, Tong Zhang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10398 2025-10-14 cs.CL cs.AI 62%

STEAM: A Semantic-Level Knowledge Editing Framework for Large Language Models

Geunyeong Jeong, Juoh Sun, Seonghee Lee, Harksoo Kim

机构 * Konkuk University(韩国康康大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18361 2025-10-14 q-bio.NC cs.AI cs.LG cs.RO 62%

Task-Optimized Convolutional Recurrent Networks Align with Tactile Processing in the Rodent Brain

Trinity Chung, Yuchen Shen, Nathan C. L. Kong, Aran Nayebi

机构 * Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所) Machine Learning Department, Carnegie Mellon University(卡内基梅隆大学机器学习系) Department of Psychology, University of Pennsylvania(宾夕法尼亚大学心理学系) Neuroscience Institute, Carnegie Mellon University(卡内基梅隆大学神经科学研究所)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 8 figures, 7 tables, NeurIPS 2025 Camera Ready Version (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15607 2025-10-14 cs.CL cs.AI 62%

From Problem-Solving to Teaching Problem-Solving: Aligning LLMs with Pedagogy using Reinforcement Learning

David Dinucu-Jianu, Jakub Macina, Nico Daheim, Ido Hakimi, Iryna Gurevych, Mrinmaya Sachan

机构 * Department of Computer Science, ETH Zurich(苏黎世联邦理工学院计算机科学系) ETH AI Center(ETH人工智能中心) Ubiquitous Knowledge Processing Lab (UKP Lab), Department of Computer Science, Technical University of Darmstadt and National Research Center for Applied Cybersecurity ATHENE, Germany(达姆施塔特技术大学计算机科学系、应用网络安全国家研究中心ATHENE)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Main as an oral presentation. David Dinucu-Jianu and Jakub Macina contributed equally. Code available: https://github.com/eth-lre/PedagogicalRL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09762 2025-10-14 cs.LG cs.AI 62%

PatentVision: A multimodal method for drafting patent applications

Ruo Yang, Sai Krishna Reddy Mudhiganti, Manali Sharma

机构 * Samsung Semiconductor, Inc.(三星半导体公司)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12963 2025-10-14 cs.AI cs.LG 62%

Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills

Changsheng Wang, Chongyu Fan, Yihua Zhang, Jinghan Jia, Dennis Wei, Parikshit Ram, Nathalie Baracaldo, Sijia Liu

机构 * Michigan State University(密歇根州立大学) IBM Research(IBM研究院)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11502 2025-10-14 cs.LG 57%

Learning to Make MISTAKEs: Modeling Incorrect Student Thinking And Key Errors

Alexis Ross, Jacob Andreas

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10397 2025-10-14 cs.CL 57%

AssoMem: Scalable Memory QA with Multi-Signal Associative Retrieval

Kai Zhang, Xinyuan Zhang, Ejaz Ahmed, Hongda Jiang, Caleb Kumar, Kai Sun, Zhaojiang Lin, Sanat Sharma, Shereen Oraby, Aaron Colak, Ahmed Aly, Anuj Kumar, Xiaozhong Liu, Xin Luna Dong

机构 * Meta Reality Labs(Meta现实实验室) Worcester Polytechnic Institute(沃斯特理工学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08075 2025-10-14 cs.AI 57%

Multi-Condition Conformal Selection

Qingyang Hao, Wenbo Liao, Bingyi Jing, Hongxin Wei

机构 * Southern University of Science and Technology(南方科技大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05444 2025-10-14 cs.CL 57%

PhoniTale: Phonologically Grounded Mnemonic Generation for Typologically Distant Language Pairs

Sana Kang, Myeongseok Gwon, Su Young Kwon, Jaewook Lee, Andrew Lan, Bhiksha Raj, Rita Singh

机构 * KAIST(韩国科学技术院) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Carnegie Mellon University(卡内基梅隆大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Accepted to EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12450 2025-10-14 cs.CL 57%

Language Surgery in Multilingual Large Language Models

Joanito Agili Lopo, Muhammad Ravi Shulthan Habibi, Tack Hwa Wong, Muhammad Ilham Ghozali, Fajri Koto, Genta Indra Winata, Peerat Limkonchotiwat, Alham Fikri Aji, Samuel Cahyawijaya

机构 * SEACrowd Kreasof AI Universitas Indonesia(印度尼西亚大学) MBZUAI Capital One AI Singapore(AI新加坡) Cohere

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11203 2025-10-14 cs.CR 50%

TraceAegis: Securing LLM-Based Agents via Hierarchical and Behavioral Anomaly Detection

Jiahao Liu, Bonan Ruan, Xianglin Yang, Zhiwei Lin, Yan Liu, Yang Wang, Tao Wei, Zhenkai Liang

专题命中 其他安全 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10655 2025-10-14 q-bio.OT 50%

Isotropy and Geometry of Pretrained Protein LMs

Sheikh Azizul Hakim, Kowshic Roy, M Saifur Rahman

专题命中 其他安全 :alignment(abstract)

Comments Published in the Proceedings of the ICML 2025 Workshop on Multi-modal Foun- dation Models and Large Language Models for Life Sciences, Vancouver, Canada. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10567 2025-10-14 cs.RO 50%

Reinforcement Learning-based Dynamic Adaptation for Sampling-Based Motion Planning in Agile Autonomous Driving

Alexander Langmann, Yevhenii Tokarev, Mattia Piccinini, Korbinian Moller, Johannes Betz

机构 * Professorship of Autonomous Vehicle Systems, TUM School of Engineering and Design, Technical University of Munich(自主车辆系统教授职位,TUM工程与设计学院,慕尼黑技术大学) Munich Institute of Robotics and Machine Intelligence (MIRMI)(慕尼黑机器人与机器智能研究所(MIRMI))

专题命中 其他安全 :safety(abstract)

Comments 8 pages, submitted to the IEEE ICRA 2026, Vienna, Austria

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05970 2025-10-14 cs.CV 50%

Automatic Synthesis of High-Quality Triplet Data for Composed Image Retrieval

Haiwen Li, Delong Liu, Zhaohui Hou, Zhicheng Zhao, Fei Su

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 其他安全 :alignment(abstract)

Comments This paper was originally submitted to ACM MM 2025 on April 12, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11188 2025-10-14 eess.SY cs.SY 50%

Global Attitude Synchronization for Heterogeneous Multi-agent Systems on SO(3)

Mouaad Boughellaba, Soulaimane Berkane, Abdelhamid Tayebi

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10337 2025-10-14 cs.RO 50%

Rise of the Robochemist

Jihong Zhu, Kefeng Huang, Jonathon Pipe, Chris Horbaczewsky, Andy Tyrrell, Ian J. S. Fairlamb

机构 * School of Physics, Engineering and Technology, University of York(物理、工程与技术学院,约克大学) Department of Chemistry, University of York(化学学院,约克大学)

专题命中 其他安全 :safety(abstract)

Comments This article was originally published in the IEEE Systems, Man, and Cybernetics Society eNewsletter, September 2025 issue: https://www.ieeesmc.org/wp-content/uploads/2024/10/FeatureArticle_Sept25.pdf

Journal ref https://www.ieeesmc.org/wp-content/uploads/2024/10/FeatureArticle_Sept25.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09716 2025-10-14 q-bio.QM 50%

MS2toImg: A Framework for Direct Bioactivity Prediction from Raw LC-MS/MS Data

Hansol Hong, Sangwon Lee, Jang-Ho Ha, Sung-June Chu, So-Hee An, Woo-Hyun Paek, Gyuhwa Chung, Kyoung Tai No

专题命中 其他安全 :alignment(abstract)

Comments 35 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏