arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-01-23 至 2026-01-23 共收录 11 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 11 篇

2503.05609 2026-01-23 cs.CY cs.HC 83%

Decoding Safety Feedback from Diverse Raters: A Data-driven Lens on Responsiveness to Severity

解码来自多样评分者的安全反馈:一种数据驱动的响应性视角

Pushkar Mishra, Charvi Rastogi, Stephen R. Pfohl, Alicia Parrish, Tian Huey Teh, Roma Patel, Mark Diaz, Ding Wang, Michela Paganini, Vinodkumar Prabhakaran, Lora Aroyo, Verena Rieser

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CY

AI总结 本文提出了一种数据驱动的方法,用于分析多元环境下安全反馈的响应性,通过量化评分者对严重性差异的表达,提升多文化背景下AI系统的对齐质量。

Journal ref Transactions on Machine Learning Research, 2835-8856, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05882 2026-01-23 cs.CL cs.AI cs.LG 82%

Collaborate, Deliberate, Evaluate: How LLM Alignment Affects Coordinated Multi-Agent Outcomes

协作、审议、评估:LLM对齐如何影响协调的多智能体结果

Abhijnan Nath, Carine Graff, Nikhil Krishnaswamy

机构 * Natural Language (SIGNAL) Lab Colorado State University Fort Collins, CO USA Natural Language (SIGNAL) Lab Colorado State University

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文研究了LLM对齐方法如何影响多智能体协作效果,通过干预代理促进审议式决策,发现鲁棒性方法在支持正确任务结果方面表现更优。

Comments This submission is a new version of arXiv:2509.05882v1. with a substantially revised experimental pipeline and new metrics. In particular, collaborator agents are now instantiated independently via separate API calls, rather than generated autoregressively by a single agent. All experimental results are new. Accepted as an extended abstract at AAMAS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20062 2026-01-23 cs.CL 79%

Poor Alignment and Steerability of Large Language Models: Evidence from College Admission Essays

大语言模型的对齐性和操控性不足:来自大学录取文书的证据

Jinsook Lee, AJ Alvero, Thorsten Joachims, René Kizilcec

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

AI总结 本研究通过分析大学录取文书发现,大语言模型在对齐和操控性方面存在不足,提示信息未能有效改变模型输出的语言模式。

Comments 48 pages, 10 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10695 2026-01-23 cs.LG cs.AI cs.CL 67%

Introducing Verification Task of Set Consistency with Set-Consistency Energy Networks

引入集合一致性验证任务与集合一致性能量网络

Mooho Song, Hyeryung Son, Jay-Yoon Lee

机构 * Seoul National University(首尔国立大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出集合一致性验证任务及SC-Energy模型,通过对比损失框架提升多陈述逻辑一致性验证性能,并发布新数据集

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025), Long Papers

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11304 2026-01-23 cs.AI cs.CL cs.CV 62%

Leveraging Multimodal-LLMs Assisted by Instance Segmentation for Intelligent Traffic Monitoring

利用实例分割辅助的多模态大语言模型进行智能交通监控

Murat Arda Onsu, Poonam Lohan, Burak Kantarci, Aisha Syed, Matthew Andrews, Sean Kennedy

机构 * University of Ottawa(渥太华大学) Nokia Bell Labs(诺基亚贝尔实验室)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

AI总结 本文利用多模态大语言模型和实例分割技术,实现高准确率的交通监控系统,提升交通管理效率和安全性。

Comments 6 pages, 7 figures, submitted to 30th IEEE International Symposium on Computers and Communications (ISCC) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15888 2026-01-23 cs.CV cs.AI 57%

Understanding the Transfer Limits of Vision Foundation Models

理解视觉基础模型的迁移限制

Shiqi Huang, Yipei Wang, Natasha Thorley, Alexander Ng, Shaheer Saeed, Mark Emberton, Shonit Punwani, Veeru Kasivisvanathan, Dean Barratt, Daniel Alexander, Yipeng Hu

机构 * University College London(伦敦大学学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文研究了视觉基础模型在迁移时的限制,发现预训练目标与下游任务需求的匹配程度影响迁移性能,提出通过改进预训练目标以提升模型效果。

Comments accepted in ISBI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15481 2026-01-23 cs.LG math.OC 57%

Early predicting of hospital admission using machine learning algorithms: Priority queues approach

利用机器学习算法提前预测医院入院:优先队列方法

Jakub Antczak, James Montgomery, Małgorzata O'Reilly, Zbigniew Palmowski, Richard Turner

机构 * Wrocław University of Science and Technology(沃拉日大学科学与技术学院) University of Tasmania(塔斯马尼亚大学) School of Medicine, University of Tasmania(塔斯马尼亚大学医学院)

专题命中 其他安全 :safety(abstract);分类 cs.LG

AI总结 本文利用机器学习算法预测医院急诊入院情况,通过对比SARIMAX、XGBoost和LSTM模型,发现XGBoost在预测总入院人数上表现最佳,但所有模型在预测突发患者激增时均存在低估问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15338 2026-01-23 cs.CL 57%

From Quotes to Concepts: Axial Coding of Political Debates with Ensemble LMs

从引语到概念:利用集成语言模型进行政治辩论的轴向编码

Angelina Parfenova, David Graus, Juergen Pfeffer

机构 * Technical University of Munich(慕尼黑技术大学) Lucerne University of Applied Sciences and Arts(卢塞恩应用科学大学) Arts University of Amsterdam(阿姆斯特丹艺术大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 本文提出利用集成语言模型进行政治辩论的轴向编码,通过聚类和LLM分组方法生成层次化代码,提升辩论分析的结构化和语义对齐性。

Comments Accepted to ECIR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16023 2026-01-23 eess.AS cs.HC 50%

Timbre-Aware LLM-based Direct Speech-to-Speech Translation Extendable to Multiple Language Pairs

具有音色感知的基于大语言模型的直接语音到语音翻译系统,可扩展到多种语言对

Lalaram Arya, Mrinmoy Bhattacharjee, Adarsh C. R., S. R. Mahadeva Prasanna

专题命中 其他安全 :alignment(abstract)

AI总结 本文提出DS2ST-LM,一种基于多语言大语言模型的直接语音到语音翻译系统,通过整合Whisper编码器、可学习投影模块和音色感知合成器,实现多语言翻译并提升说话人相似性与语音自然度。

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15643 2026-01-23 cs.CV 50%

Evolving Without Ending: Unifying Multimodal Incremental Learning for Continual Panoptic Perception

无止境的进化:统一多模态增量学习以实现持续全景感知

Bo Yuan, Danpei Zhao, Wentao Li, Tian Li, Zhiguo Jiang

专题命中 其他安全 :alignment(abstract)

AI总结 本文提出了一种统一多模态和多任务持续学习的持续全景感知模型,通过跨模态编码器和可变知识继承模块解决灾难性遗忘问题,并在多任务增量场景中实现模型进化。

Comments arXiv admin note: substantial text overlap with arXiv:2407.14242

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15453 2026-01-23 cs.CV 50%

DevPrompt: Deviation-Based Prompt Learning for One-Normal ShotImage Anomaly Detection

DevPrompt: 基于偏差的提示学习用于单正常样本图像异常检测

Morteza Poudineh, Marc Lalonde

机构 * Department of Electrical and Computer Engineering, Concordia University(电气与计算机工程系,康科迪亚大学) Computer Research Institute of Montreal (CRIM)(蒙特利尔计算机研究 institute)

专题命中 其他安全 :alignment(abstract)

AI总结 DevPrompt通过结合视觉语言模型的语义能力和基于偏差的评分机制,提升单正常样本图像异常检测的性能和可解释性。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏