arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-05 至 2025-11-05 共收录 13 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 13 篇

2509.04104 2025-11-05 cs.CL cs.HC 79%

Towards Stable and Personalised Profiles for Lexical Alignment in Spoken Human-Agent Dialogue

Keara Schaaij, Roel Boumans, Tibor Bosse, Iris Hendrickx

机构 * Centre for Language Studies, Centre for Language and Speech Technology, Radboud University,Nijmegen, The Netherlands(语言研究所以及语言与语音技术中心,拉德堡德大学,尼姆egen,荷兰) Behavioural Science Institute, Radboud University, Nijmegen, The Netherlands(行为科学研究所,拉德堡德大学,尼姆egen,荷兰)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments This preprint has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this contribution is published in TSD 2025. Lecture Notes in Computer Science, vol 16029

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09485 2025-11-05 cs.RO cs.AI cs.GR 79%

Adv-BMT: Bidirectional Motion Transformer for Safety-Critical Traffic Scenario Generation

Yuxin Liu, Zhenghao Peng, Xuanhao Cui, Bolei Zhou

机构 * University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18835 2025-11-05 cs.SE 71%

AUCAD: Automated Construction of Alignment Dataset from Log-Related Issues for Enhancing LLM-based Log Generation

Hao Zhang, Dongjun Yu, Lei Zhang, Guoping Rong, Yongda Yu, Haifeng Shen, He Zhang, Dong Shao, Hongyu Kuang

专题命中 其他安全 :alignment(title)

Comments In the 16th International Conference on Internetware 2025. 13 pages

Journal ref Proceedings of the 16th International Conference on Internetware (2025) 413-425

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01894 2025-11-05 cs.GR cs.AI cs.LG 62%

LGCC: Enhancing Flow Matching Based Text-Guided Image Editing with Local Gaussian Coupling and Context Consistency

Fangbing Liu, Pengfei Duan, Wen Li, Yi He

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05862 2025-11-05 cs.CL cs.AI 62%

Revisiting Long-context Modeling from Context Denoising Perspective

Zecheng Tang, Baibei Ji, Juntao Li, Lijun Wu, Haijia Gui, Min Zhang

机构 * Soochow University(苏州大学) LCM Laboratory(长文实验室) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15090 2025-11-05 cs.CL cs.AI 62%

ExpertLens: Activation steering features are highly interpretable

Masha Fedzechkina, Eleonora Gualdoni, Sinead Williamson, Katherine Metcalf, Skyler Seto, Barry-John Theobald

机构 * Apple(苹果公司)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02690 2025-11-05 cs.LG 57%

Curriculum Design for Trajectory-Constrained Agent: Compressing Chain-of-Thought Tokens in LLMs

Georgios Tzannetos, Parameswaran Kamalaruban, Adish Singla

机构 * MPI-SWS(马克斯·普朗克所际研究所)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments NeurIPS'25 paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02606 2025-11-05 cs.AI cs.HC 57%

A Multi-Agent Psychological Simulation System for Human Behavior Modeling

Xiangen Hu, Jiarui Tong, Sheng Xu

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21971 2025-11-05 cs.LG 57%

GRAM-DTI: adaptive multimodal representation learning for drug target interaction prediction

Feng Jiang, Amina Mollaysa, Hehuan Ma, Tommaso Mansi, Junzhou Huang, Mangal Prakash, Rui Liao

机构 * University of Texas at Arlington(德克萨斯理工大学) Johnson & Johnson Innovative Medicine(强生创新医学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Journal ref NeurIPS 2025 2nd Workshop on Multi-modal Foundation Models and Large Language Models for Life Sciences

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06910 2025-11-05 cs.CL 57%

Identifying Aspects in Peer Reviews

Sheng Lu, Ilia Kuznetsov, Iryna Gurevych

机构 * Ubiquitous Knowledge Processing Lab (UKP Lab)(通用知识处理实验室) Department of Computer Science(计算机科学系) Hessian Center for AI (hessian.AI)(黑森人工智能中心)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02711 2025-11-05 cs.DB cs.IR 50%

Relational Deep Dive: Error-Aware Queries Over Unstructured Data

Daren Chao, Kaiwen Chen, Naiqing Guan, Nick Koudas

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02423 2025-11-05 eess.SP 50%

LLM4PG: Adapting Large Language Model for Pathloss Map Generation via Synesthesia of Machines

Mingran Sun, Lu Bai, Xiang Cheng, Jianjun Wu

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02203 2025-11-05 cs.SE 50%

LLMs as Judges: Toward The Automatic Review of GSN-compliant Assurance Cases

Gerhard Yu, Mithila Sivakumar, Alvine B. Belle, Soude Ghari, Song Wang, Timothy C. Lethbridge

专题命中 其他安全 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏