arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-10 至 2025-10-10 共收录 25 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 25 篇

2510.07709 2025-10-10 cs.AI cs.CL cs.CY cs.MA 82%

Multimodal Safety Evaluation in Generative Agent Social Simulations

Alhim Vera, Karen Sanchez, Carlos Hinojosa, Haidar Bin Hamid, Donghoon Kim, Bernard Ghanem

机构 * University of Cincinnati(辛辛那提大学) King Abdullah University of Science and Technology(国王阿卜杜勒·阿齐兹大学科学与技术)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11885 2025-10-10 cs.CL 79%

Med-R$^2$: Crafting Trustworthy LLM Physicians via Retrieval and Reasoning of Evidence-Based Medicine

Keer Lu, Zheng Liang, Da Pan, Shusen Zhang, Guosheng Dong, Zhonghai Wu, Huang Leng, Bin Cui, Wentao Zhang

机构 * Center for Data Science, AAIS Peking University Beijing China(数据科学中心,北京大学北京中国) Baichuan Inc. Beijing China(北川公司北京中国) School of Computer Science Peking University Beijing China(计算机科学学院,北京大学北京中国) Peking University(北京大学) Baichuan Inc.(北川公司)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07652 2025-10-10 cs.CV 78%

Dual-Stream Alignment for Action Segmentation

Harshala Gammulle, Clinton Fookes, Sridha Sridharan, Simon Denman

机构 * Signal Processing, Artificial Intelligence and Vision Technologies (SAIVT) Lab(信号处理、人工智能与视觉技术实验室) Queensland University of Technology(昆士兰理工大学)

专题命中 安全评测 :alignment(title,abstract)

Comments Journal Submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08429 2025-10-10 cs.LG cs.AI stat.ML 76%

ClauseLens: Clause-Grounded, CVaR-Constrained Reinforcement Learning for Trustworthy Reinsurance Pricing

Stella C. Dong, James R. Finlay

机构 * Department of Applied Mathematics, University of California, Davis, CA, USA(加州大学戴维斯分校应用数学系) Wharton School of Business, University of Pennsylvania, Philadelphia, PA, USA(宾夕法尼亚大学沃顿商学院)

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

Comments Accepted for publication at the 6th ACM International Conference on AI in Finance (ICAIF 2025), Singapore. Author-accepted version (October 2025). 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14756 2025-10-10 cs.LG cs.AI 76%

LLINBO: Trustworthy LLM-in-the-Loop Bayesian Optimization

Chih-Yu Chang, Milad Azvar, Chinedum Okwudire, Raed Al Kontar

机构 * University of Michigan(密歇根大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08569 2025-10-10 cs.CL cs.AI cs.LG 75%

ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation

Qin Liu, Jacob Dineen, Yuxi Huang, Sheng Zhang, Hoifung Poon, Ben Zhou, Muhao Chen

机构 * University of California, Davis(加州大学戴维斯分校) Arizona State University(亚利桑那州立大学) Microsoft Research(微软研究院)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07491 2025-10-10 cs.AI 70%

Optimizing Ethical Risk Reduction for Medical Intelligent Systems with Constraint Programming

Clotilde Brayé, Aurélien Bricout, Arnaud Gotlieb, Nadjib Lazaar, Quentin Vallet

机构 * Enovacom Simula Research Laboratory LISN, Université Paris-Saclay(LISN,巴黎萨克雷大学)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01493 2025-10-10 cs.CL 70%

Sherkala-Chat: Building a State-of-the-Art LLM for Kazakh in a Moderately Resourced Setting

Fajri Koto, Rituraj Joshi, Nurdaulet Mukhituly, Yuxia Wang, Zhuohan Xie, Rahul Pal, Daniil Orel, Parvez Mullah, Diana Turmakhan, Maiya Goloburda, Mohammed Kamran, Samujjwal Ghosh, Bokang Jia, Jonibek Mansurov, Mukhammed Togmanov, Debopriyo Banerjee, Nurkhan Laiyk, Akhmed Sakip, Xudong Han, Ekaterina Kochmar, Alham Fikri Aji, Aaryamonvikram Singh, Alok Anil Jadhav, Satheesh Katipomu, Samta Kamboj, Monojit Choudhury, Gurpreet Gosal, Gokulakrishnan Ramakrishnan, Biswajit Mishra, Sarath Chandran, Avraham Sheinin, Natalia Vassilieva, Neha Sengupta, Preslav Nakov

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Inception Cerebras Systems(Cerebras系统)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL

Comments Accepted at COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19195 2025-10-10 cs.CL cs.AI cs.LG 67%

Can Small-Scale Data Poisoning Exacerbate Dialect-Linked Biases in Large Language Models?

Chaymaa Abbas, Mariette Awad, Razane Tajeddine

机构 * Department of Electrical and Computer Engineering, Maroun Semaan Faculty of Engineering and Architecture(电气与计算机工程系,马鲁恩·塞马安工程与建筑学院) American University of Beirut(贝鲁特美国大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08114 2025-10-10 cs.AI cs.CL 62%

Can Risk-taking AI-Assistants suitably represent entities

Ali Mazyaki, Mohammad Naghizadeh, Samaneh Ranjkhah Zonouzaghi, Amirhossein Farshi Sotoudeh

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07626 2025-10-10 cs.LG cs.CL 62%

LLM Unlearning Under the Microscope: A Full-Stack View on Methods and Metrics

Chongyu Fan, Changsheng Wang, Yancheng Huang, Soumyadeep Pal, Sijia Liu

机构 * Michigan State University(密歇根州立大学) IBM Research(IBM研究院)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04363 2025-10-10 cs.SE cs.AI cs.CL 62%

MacroBench: A Novel Testbed for Web Automation Scripts via Large Language Models

Hyunjun Kim, Sejong Kim

机构 * KAIST(韩国科学技术院)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments NeurIPS 2025 Workshop on Lock-LLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25256 2025-10-10 cs.CY cs.AI 62%

The Sandbox Configurator: A Framework to Support Technical Assessment in AI Regulatory Sandboxes

Alessio Buscemi, Thibault Simonetto, Daniele Pagani, German Castignani, Maxime Cordy, Jordi Cabot

机构 * Luxembourg Institute of Science and Technology (LIST)(卢森堡科学与技术研究院) University of Luxembourg(卢森堡大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17694 2025-10-10 cs.CL cs.AI 62%

Evaluating LLM-Generated Versus Human-Authored Responses in Role-Play Dialogues

Dongxu Lu, Johan Jeuring, Albert Gatt

机构 * Utrecht University(乌得勒支大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted for publication at the 18th International Natural Language Generation Conference (INLG 2025). Revised version: improved image quality and minor corrections. No change to conclusions

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00529 2025-10-10 cs.CL cs.CY 62%

Modeling Motivated Reasoning in Law: Evaluating Strategic Role Conditioning in LLM Summarization

Eunjung Cho, Alexander Hoyle, Yoan Hermstrüwer

机构 * ETH Zurich(苏黎世联邦理工学院) University of Zurich(苏黎世大学) Max Planck Institute for Research on Collective Goods(集体利益研究所)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.CY

Comments Accepted at NLLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17717 2025-10-10 cs.CL cs.AI 62%

From Feedback to Checklists: Grounded Evaluation of AI-Generated Clinical Notes

Karen Zhou, John Giorgi, Pranav Mani, Peng Xu, Davis Liang, Chenhao Tan

机构 * University of Chicago(芝加哥大学) Abridge

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08713 2025-10-10 cs.LG cs.AI 62%

ProtoECGNet: Case-Based Interpretable Deep Learning for Multi-Label ECG Classification with Contrastive Learning

Sahil Sethi, David Chen, Thomas Statchen, Michael C. Burkhart, Nipun Bhandari, Bashar Ramadan, Brett Beaulieu-Jones

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Accepted to PMLR 298, 10th Machine Learning for Healthcare Conference (MLHC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.02944 2025-10-10 cs.LG cs.AI cs.SY eess.SY 62%

Foundation Models for Structural Health Monitoring

Luca Benfenati, Daniele Jahier Pagliari, Luca Zanatta, Yhorman Alexander Bedoya Velez, Andrea Acquaviva, Massimo Poncino, Enrico Macii, Luca Benini, Alessio Burrello

机构 * DAUIN, Politecnico di Torino(达乌因,托斯卡纳理工学院) DEI, University of Bologna(电子工程学院,博洛尼亚大学) DIST, Politecnico di Torino(信息与通信技术学院,托斯卡纳理工学院) D-ITET, ETH Zurich(信息与通信技术系,苏黎世联邦理工学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments 17 pages, 6 tables, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07881 2025-10-10 cs.CL 57%

CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-Switching

Heyang Liu, Yuhao Wang, Ziyang Cheng, Ronghua Wu, Qunshan Gu, Yanfeng Wang, Yu Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Ant Group(蚂蚁集团)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07829 2025-10-10 cs.HC cs.AI 57%

The Rise of the Knowledge Sculptor: A New Archetype for Knowledge Work in the Age of Generative AI

Cathal Doyle

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 23 pages, 11 figures, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07748 2025-10-10 cs.AI 57%

Haibu Mathematical-Medical Intelligent Agent:Enhancing Large Language Model Reliability in Medical Tasks via Verifiable Reasoning Chains

Yilun Zhang, Dexing Kong

机构 * Zhejiang Qiushi Institute of Mathematical Medicine(浙江启思数学医学研究院) School of Mathematical Sciences, Zhejiang University(浙江大学数学科学学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07697 2025-10-10 cs.CR cs.AI 57%

Rethinking Reasoning: A Survey on Reasoning-based Backdoors in LLMs

Man Hu, Xinyi Wu, Zuofeng Suo, Jinbo Feng, Linghui Meng, Yanhao Jia, Anh Tuan Luu, Shuai Zhao

机构 * Beijing Electronic Science and Technology Institute, China(北京电子科技研究所) Nanyang Technological University, Singapore(南洋理工大学) Hainan University, China(海南大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07350 2025-10-10 cs.LG 57%

Out-of-Distribution Generalization in Climate-Aware Yield Prediction with Earth Observation Data

Aditya Chakravarty

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Journal ref ICCV 2025 Workshop on Sustainability with Earth observation and AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18709 2025-10-10 cs.CL 57%

Adaptive Originality Filtering: Rejection Based Prompting and RiddleScore for Culturally Grounded Multilingual Riddle Generation

Duy Le, Kent Ziti, Evan Girard-Sun, Bakr Bouhaya, Sean O'Brien, Vasu Sharma, Kevin Zhu

机构 * Algoverse AI Research(Algoverse AI研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Paper was accepted in to NeurIPS 2025 Workshop GenProCC

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08003 2025-10-10 cs.CV 50%

CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning

Weihuang Lin, Yiwei Ma, Jiayi Ji, Xiaoshuai Sun, Rongrong Ji

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(教育部多媒体可信感知与高效计算重点实验室,厦门大学)

专题命中 安全评测 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏