arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-02-12 至 2026-02-12 共收录 55 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 14 篇

2504.00389 2026-02-12 cs.AI 70%

CyberBOT: Towards Reliable Cybersecurity Education via Ontology-Grounded Retrieval Augmented Generation

CyberBOT: 通过基于本体的检索增强生成实现可靠的网络安全教育

Chengshuai Zhao, Riccardo De Maria, Tharindu Kumarage, Kumar Satvik Chaudhary, Garima Agrawal, Yiwen Li, Jongchan Park, Yuli Deng, Ying-Chih Chen, Huan Liu

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI

AI总结 CyberBOT通过基于本体的检索增强生成技术,提供可靠的网络安全教育解决方案,提升教学效果和学生参与度。

Comments Accepted by The Conference on Information and Knowledge Management (CIKM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03083 2026-02-12 cs.DS cs.AI cs.CL 62%

Algorithmically Establishing Trust in Evaluators

算法上建立评估者的信任

Adrian de Wynter

机构 * Microsoft(微软公司) The University of York(约克大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

AI总结 本文提出无数据算法,通过连续挑战建立评估者信任,无需标注数据,适用于低资源语言标注场景。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18970 2026-02-12 cs.AI cs.LG 62%

Bridging Explainability and Embeddings: BEE Aware of Spuriousness

弥合可解释性与嵌入:BEE 意识到虚假相关性

Cristian Daniel Păduraru, Antonio Bărbălau, Radu Filipescu, Andrei Liviu Nicolicioiu, Elena Burceanu

机构 * Bitdefender University of Bucharest(布加勒斯特大学) Mila University of Montreal(蒙特利尔大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 BEE通过分析权重空间和嵌入几何揭示隐藏的虚假相关性,适用于多种数据集和模型,提升基础模型的可信度。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11091 2026-02-12 cs.CL 57%

Can Large Language Models Make Everyone Happy?

大语言模型能否让所有人快乐?

Usman Naseem, Gautam Siddharth Kashyap, Ebad Shabbir, Sushant Kumar Ray, Abdullah Mohammad, Rafiq Ali

机构 * Macquarie University(麦考瑞大学) DSEU-Okhla University of Delhi(德里大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL

AI总结 本文提出MisAlign-Profile基准,通过构建MISALIGNTRADE数据集,评估大语言模型在安全、价值和文化维度间的不匹配权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10620 2026-02-12 cs.SE cs.CL 57%

ISD-Agent-Bench: A Comprehensive Benchmark for Evaluating LLM-based Instructional Design Agents

ISD-Agent-Bench: 一个用于评估基于大语言模型的教学系统设计代理的全面基准

YoungHoon Jeon, Suwan Kim, Haein Son, Sookbun Lee, Yeil Jeong, Unggi Lee

机构 * Upstage Opentutorials Indiana University Bloomington(印第安纳大学布卢明顿分校) Korea University Sejong Campus(韩国大学世宗校区)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 ISD-Agent-Bench通过结合经典ISD理论与现代推理方法,为评估基于LLM的教学系统设计代理提供了全面的基准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10367 2026-02-12 cs.AI 57%

LiveMedBench: A Contamination-Free Medical Benchmark for LLMs with Automated Rubric Evaluation

LiveMedBench: 一个无污染的医疗基准测试,用于具有自动评分评估的LLM

Zhiling Yan, Dingjie Song, Zhe Fang, Yisheng Ji, Xiang Li, Quanzheng Li, Lichao Sun

机构 * Lehigh University(莱维大学) Harvard University(哈佛大学) Imperial College London(伦敦帝国学院) Massachusetts General Hospital(麻省总医院) Harvard Medical School(哈佛医学院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

AI总结 LiveMedBench通过持续更新和自动评分框架,提供无污染的医疗基准测试,验证LLM在临床推理中的性能,揭示数据污染和上下文适应性问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15044 2026-02-12 cs.LG quant-ph 57%

IQNN-CS: Interpretable Quantum Neural Network for Credit Scoring

IQNN-CS:用于信用评分的可解释量子神经网络

Abdul Samad Khan, Nouhaila Innan, Aeysha Khalique, Muhammad Shafique

机构 * Lahore University of Management Sciences, Pakistan(拉合尔管理科学大学,巴基斯坦) eBRAIN Lab, Division of Engineering, New York University Abu Dhabi (NYUAD)(eBRAIN实验室,工程系,纽约大学阿布扎比分校) Center for Quantum and Topological Systems (CQTS), NYUAD Research Institute(量子与拓扑系统中心(CQTS),NYUAD研究机构)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

AI总结 IQNN-CS是一种用于信用评分的可解释量子神经网络,通过引入ICAA度量标准提升模型的可解释性和透明度,适用于金融决策中的高风险任务。

Comments Accepted for oral presentation at QUEST-IS'25. To appear in Springer proceedings

Journal ref International Conference on Quantum Engineering Sciences and Technologies for Industry and Services 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11118 2026-02-12 stat.ME stat.ML 50%

A Doubly Robust Machine Learning Approach for Disentangling Treatment Effect Heterogeneity with Functional Outcomes

一种用于解构功能性结果中治疗效应异质性的双重鲁棒机器学习方法

Filippo Salmaso, Lorenzo Testa, Francesca Chiaromonte

专题命中 安全评测 :trustworthy(abstract)

AI总结 FOCaL是一种针对功能性结果中治疗效应异质性估计的双重鲁棒机器学习方法,通过整合先进功能回归技术,实现了对复杂数据中个性化因果效应的稳健推断。

Comments 20 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19899 2026-02-12 cs.CV 50%

VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering

VeriSciQA:一个用于科学视觉问答的自动验证数据集

Yuyi Li, Daoyuan Chen, Zhen Wang, Yutong Lu, Yaliang Li

机构 * Sun Yat-sen University(中山大学) Alibaba Group(阿里巴巴集团)

专题命中 安全评测 :alignment(abstract)

AI总结 VeriSciQA通过跨模态验证框架生成并验证高质量科学视觉问答数据集,显著提升开源模型在科学图表问答任务上的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 3 篇

2601.23001 2026-02-12 cs.CL cs.AI 62%

Bias Beyond Borders: Political Ideology Evaluation and Steering in Multilingual LLMs

偏见超越国界:多语言大语言模型中的政治意识形态评估与引导

Afrozah Nadeem, Agrima Seth, Mehwish Nasim, Usman Naseem

机构 * School of Computing, Macquarie University, Australia(麦考瑞大学计算机学院) Microsoft, USA(微软公司) University of Western Australia, Australia(西澳大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文提出跨语言对齐引导框架,用于评估和减轻多语言LLM中的政治偏见,通过跨语言对齐意识形态表示以实现公平性与多样性的平衡。

Comments PrePrint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08193 2026-02-12 cs.AI 57%

Measuring What Matters: The AI Pluralism Index

衡量重要性:AI多元主义指数

Rashid Mushkani

机构 * Université de Montréal(蒙特利尔大学) Mila – Québec AI Institute(魁北克人工智能研究所)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

AI总结 本文提出AI多元主义指数,用于衡量人工智能系统在治理、包容性和透明度方面的多元实践,旨在引导激励向多元主义方向发展。

Comments Proceedings of the International Association for Safe & Ethical AI (IASEAI), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12365 2026-02-12 cs.CL cs.DB 57%

Advances in LLMs with Focus on Reasoning, Adaptability, Efficiency and Ethics

大语言模型的进展:聚焦推理、适应性、效率和伦理

Asifullah Khan, Muhammad Zaeem Khan, Aleesha Zainab, Saleha Jamshed, Sadia Ahmad, Kaynat Khatib, Faria Bibi, Abdul Rehman

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

AI总结 本文综述了大语言模型在推理、适应性、效率和伦理方面的进展,探讨了关键技术和挑战,提出未来研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 13 篇

2602.09877 2026-02-12 cs.CL 83%

The Devil Behind Moltbook: Anthropic Safety is Always Vanishing in Self-Evolving AI Societies

Moltbook背后的魔鬼:在自我进化的AI社会中,人类安全始终在消失

Chenxu Wang, Chaozhuo Li, Songyang Liu, Zejian Chen, Jinyu Hou, Ji Qi, Rui Li, Litian Zhang, Qiwei Ye, Zheng Liu, Xu Chen, Xi Zhang, Philip S. Yu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Renmin University of China(中国人民大学) University of Illinois at Chicago(伊利诺伊大学香槟分校)

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CL

AI总结 研究揭示了自我进化AI社会中安全持续性的不可能性,并提出解决方案以缓解安全风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10593 2026-02-12 cs.CV 71%

Fast Person Detection Using YOLOX With AI Accelerator For Train Station Safety

利用YOLOX与AI加速器的快速人员检测用于车站安全

Mas Nurul Achmadiah, Novendra Setyawan, Achmad Arif Bryantono, Chi-Chia Sun, Wen-Kai Kuo

机构 * Department of Electro-Optics, National Formosa University, Taiwan(国立Formosa大学电子光学系) Department of Electrical Engineering, National Taipei University, Taiwan(国立台北大学电子工程系) Department of Electrical Engineering, University of Muhammadiyah Malang, Indonesia(穆罕默迪亚大学Malang分校电子工程系) Department of Electronics Engineering, State Polytechnic of Malang, Indonesia(Malang州立理工学院电子工程系) Smart Manufacturing and Intelligent Machinery Research Center, National Formosa University, Taiwan(国立Formosa大学智能制造与智能机械研究中心) Department of Electronics Engineering, National Formosa University, Taiwan(国立Formosa大学电子工程系)

专题命中 其他安全 :safety(title)

AI总结 本文提出利用YOLOX与Hailo-8 AI加速器提高车站乘客检测的准确性和效率

Comments 6 pages, 8 figures, 2 tables. Presented at 2024 International Electronics Symposium (IES). IEEE DOI: 10.1109/IES63037.2024.10665874

Journal ref 2024 International Electronics Symposium (IES), pp. 504-509, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05228 2026-02-12 cs.AI 70%

Surgery: Mitigating Harmful Fine-Tuning for Large Language Models via Attention Sink

手术:通过注意力sink缓解大型语言模型的有害微调

Guozhi Liu, Weiwei Lin, Tiansheng Huang, Ruichao Mo, Qi Mu, Xiumin Wang, Li Shen

机构 * South China University of Technology(南方科技大学) Pengcheng Laboratory(鹏城实验室) Sun Yat-sen University(中山大学)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

AI总结 Surgery通过注意力sink机制在微调阶段抑制有害模式学习,提升模型安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11924 2026-02-12 cs.CL cs.AI cs.LG 67%

Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability

大语言模型中的内在自我修正:通过机制可解释性实现可解释的提示

Yu-Ting Lee, Fu-Chieh Chang, Yu-En Shu, Hui-Ying Shih, Pei-Yuan Wu

机构 * Graduate Institute of Communication Engineering, National Taiwan University, Taipei, Taiwan(通讯工程研究院,国立台湾大学) MediaTek Research, Taipei, Taiwan(联发科研究,台北,台湾) Department of Electrical Engineering, National Taiwan University, Taipei, Taiwan(电气工程系,国立台湾大学) Department of Electrical Engineering, National Tsing Hua University, Hsinchu, Taiwan(电气工程系,国立清华大学) AI Research Center (AINTU), National Taiwan University, Taipei, Taiwan(人工智能研究中心(AINTU),国立台湾大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究通过机制可解释性揭示大语言模型中内在自我修正的机制,证明提示引导隐藏表示偏移是其核心驱动因素。

Journal ref 4th Deployable AI Workshop at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11024 2026-02-12 cs.CV cs.AI 57%

Chain-of-Look Spatial Reasoning for Dense Surgical Instrument Counting

密集手术器械计数的链式观察空间推理

Rishikesh Bhyri, Brian R Quaranto, Philip J Seger, Kaity Tung, Brendan Fox, Gene Yang, Steven D. Schwaitzberg, Junsong Yuan, Nan Xi, Peter C W Kim

机构 * State University of New York at Buffalo(纽约州立大学布法罗分校)

专题命中 其他安全 :safety(abstract);分类 cs.AI

AI总结 本文提出Chain-of-Look框架,通过结构化视觉链提升密集手术器械计数的准确性,并引入邻近损失函数和SurgCount-HD数据集,实验证明其在复杂场景中的优越性能。

Comments Accepted to WACV 2026. This version includes additional authors who contributed during the rebuttal phase

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10597 2026-02-12 cs.CY 57%

Llama-Polya: Instruction Tuning for Large Language Model based on Polya's Problem-solving

Llama-Polya:基于波利亚问题解决框架的大型语言模型指令调优

Unggi Lee, Yeil Jeong, Chohui Lee, Gyuri Byun, Yunseo Lee, Minji Kang, Minji Jeon

专题命中 其他安全 :alignment(abstract);分类 cs.CY

AI总结 Llama-Polya通过整合波利亚问题解决框架,提升大型语言模型在数学推理和教学对齐方面的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10454 2026-02-12 cs.CL 57%

LATA: A Tool for LLM-Assisted Translation Annotation

LATA:一个辅助翻译标注的工具

Baorong Huang, Ali Asiri

机构 * School of Foreign Languages, Huaihua University(怀化大学外语学院) Alith University College, Umm al-Qura University(乌姆·阿尔·库拉大学阿利思大学学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 LATA是一个利用大语言模型辅助的翻译标注工具,旨在提高标注效率的同时保持语言学精度,适用于结构差异大的语言对。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10153 2026-02-12 cs.CR cs.LG cs.SE 57%

Basic Legibility Protocols Improve Trusted Monitoring

基础可读性协议提升可信监控

Ashwin Sreevatsa, Sebastian Prasanna, Cody Rushing

机构 * Cambridge Boston Alignment Initiative (CBAI) Summer Research Fellowship(剑桥波士顿对齐计划暑期研究奖学金) Redwood Research(红木研究)

专题命中 其他安全 :safety(abstract);分类 cs.LG

AI总结 本研究提出可读性协议,通过鼓励不受信任模型产生易于监控评估的行为,提升可信监控的安全性,同时增强监控者对诚实代码的识别能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03737 2026-02-12 cs.LG 57%

Soft Sensor for Bottom-Hole Pressure Estimation in Petroleum Wells Using Long Short-Term Memory and Transfer Learning

基于长短期记忆与迁移学习的石油井底压力估计软传感器

M. A. Fernandes, E. Gildin, M. A. Sampaio

机构 * University of São Paulo(圣保罗大学)

专题命中 其他安全 :safety(abstract);分类 cs.LG

AI总结 本文提出基于LSTM和迁移学习的石油井底压力估计软传感器,通过实测数据验证其在成本和精度上的优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04518 2026-02-12 eess.AS cs.CL 57%

Towards Efficient Speech-Text Jointly Decoding within One Speech Language Model

迈向单个语音语言模型中的高效语音-文本联合解码

Haibin Wu, Yuxuan Hu, Ruchao Fan, Xiaofei Wang, Kenichi Kumatani, Bo Ren, Jianwei Yu, Heng Lu, Lijuan Wang, Yao Qian, Jinyu Li

机构 * Microsoft(微软公司)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 本文提出了一种早停交错解码方法,以提升语音-文本联合解码的效率和性能,并通过高质量QA数据集进一步优化语音问答性能。

Comments Accepted by ASRU 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10701 2026-02-12 cs.HC 50%

Don't blame me: How Intelligent Support Affects Moral Responsibility in Human Oversight

不要责怪我:智能支持如何影响人类监督中的道德责任

Cedric Faas, Richard Uth, Sarah Sterz, Markus Langer, Anna Maria Feit

专题命中 其他安全 :safety(abstract)

AI总结 研究探讨了智能支持系统如何影响人类在监督任务中的道德责任感知,发现限制选择会降低责任感知,但对其他相关方的责任判断无影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10411 2026-02-12 cs.IR 50%

GeoGR: A Generative Retrieval Framework for Spatio-Temporal Aware POI Recommendation

GeoGR: 一种面向空间时间感知的POI推荐生成框架

Fangye Wang, Haowen Lin, Yifang Yuan, Siyuan Wang, Xiaojiang Zhou, Song Yang, Pengjie Wang

专题命中 其他安全 :alignment(abstract)

AI总结 GeoGR是一种面向空间时间感知的POI推荐生成框架,通过两阶段设计提升POI推荐的准确性和可扩展性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05826 2026-02-12 cs.HC 50%

Whispers of the Butterfly: A Research-through-Design Exploration of In-Situ Conversational AI Guidance in Large-Scale Outdoor MR Exhibitions

蝴蝶低语:一种在大规模户外MR展览中现场对话式AI引导的研以致用探索

Dongyijie Primo Pan, Shuyue Li, Yawei Zhao, Junkun Long, Hao Li, Pan Hui

专题命中 其他安全 :safety(abstract)

AI总结 本文通过研以致用方法,开发了Dream-Butterfly,一种在户外MR展览中提供多语言解释的AI导览员,探索了人机协作在安全受限环境中的应用与影响。

详情

展开后加载摘要…

URL PDF HTML 收藏