arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-12-09 至 2025-12-09 共收录 76 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 28 篇

2512.06583 2025-12-09 econ.GN q-fin.EC 50%

Tournament-Based Performance Evaluation and Systematic Misallocation: Why Forced Ranking Systems Produce Random Outcomes

基于竞赛的绩效评估与系统性误分配:为什么强制排名系统产生随机结果

Jeremy McEntire

专题命中 安全评测 :alignment(abstract)

AI总结 本文揭示强制排名系统因系统性误分配导致随机结果,指出其无法有效解决委托-代理问题,反而加剧了分配误差。

Comments 31 pages, 6 tables. Agent-based simulation demonstrating structural allocation failures in tournament-based forced distribution evaluation mechanisms. Includes sensitivity analyses across team bias levels, alternative distributions, and cutoff percentages

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 3 篇

2511.11551 2025-12-09 cs.AI cs.CL 66%

Aligning Machiavellian Agents: Behavior Steering via Test-Time Policy Shaping

对 Machiavellian 代理进行对齐:通过测试时策略塑造实现行为引导

Dena Mujtaba, Brian Hu, Anthony Hoogs, Arslan Basharat

专题命中 AI治理与伦理 :alignment(abstract,comments);分类 cs.CL、cs.AI

AI总结 本文提出了一种测试时策略塑造方法,通过模型引导的策略调整,解决预训练代理在复杂环境中的伦理对齐问题,实现奖励最大化与伦理约束的平衡。

Comments Accepted to AAAI 2026 AI Alignment Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07360 2025-12-09 cs.CV cs.AI 57%

Structure-Aware Feature Rectification with Region Adjacency Graphs for Training-Free Open-Vocabulary Semantic Segmentation

基于区域邻接图的结构感知特征校正用于无训练开放词汇语义分割

Qiming Huang, Hao Ai, Jianbo Jiao

机构 * The MIx Group, School of Computer Science University of Birmingham(米克集团,计算机科学学院,伯明翰大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 本文提出基于区域邻接图的结构感知特征校正方法,通过增强局部辨别能力来提升开放词汇语义分割的性能。

Comments Accepted to WACV2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12758 2025-12-09 cs.CL 57%

Democratic or Authoritarian? Probing a New Dimension of Political Biases in Large Language Models

民主或威权?探讨大型语言模型中政治偏见的新维度

David Guzman Piedrahita, Irene Strauss, Bernhard Schölkopf, Rada Mihalcea, Zhijing Jin

机构 * University of Zürich(苏黎世大学) ETH Zürich(苏黎世联邦理工学院) MPI for Intelligent Systems(智能系统研究所) University of Michigan(密歇根大学) University of Toronto(多伦多大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

AI总结 本文探讨了大型语言模型在民主与威权政治偏见方面的表现,通过引入F-scale、FavScore和角色模型探测方法,揭示了模型在不同语言提示下对民主与威权倾向性的差异。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 12 篇

2507.11473 2025-12-09 cs.AI cs.LG stat.ML 88%

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

思维链可监控性:AI安全的新且脆弱的机会

Tomek Korbak, Mikita Balesni, Elizabeth Barnes, Yoshua Bengio, Joe Benton, Joseph Bloom, Mark Chen, Alan Cooney, Allan Dafoe, Anca Dragan, Scott Emmons, Owain Evans, David Farhi, Ryan Greenblatt, Dan Hendrycks, Marius Hobbhahn, Evan Hubinger, Geoffrey Irving, Erik Jenner, Daniel Kokotajlo, Victoria Krakovna, Shane Legg, David Lindner, David Luan, Aleksander Mądry, Julian Michael, Neel Nanda, Dave Orr, Jakub Pachocki, Ethan Perez, Mary Phuong, Fabien Roger, Joshua Saxe, Buck Shlegeris, Martín Soto, Eric Steinberger, Jasmine Wang, Wojciech Zaremba, Bowen Baker, Rohin Shah, Vlad Mikulik

机构 * UK AI Security Institute(英国人工智能安全研究所) Apollo Research(阿波罗研究) METR University of Montreal(蒙特利尔大学) Mila Anthropic OpenAI(开放人工智能研究所) Google DeepMind(谷歌DeepMind) Truthful AI UC Berkeley(伯克利大学) Center for AI Safety(人工智能安全中心) AI Futures Project(人工智能未来项目) Amazon(亚马逊) Scale AI Magic Meta Redwood Research(红木研究)

专题命中 其他安全 :safety(title,abstract);AI safety(title,abstract);分类 cs.AI、cs.LG

AI总结 本文探讨了通过监控AI思维链来提升安全性的潜力,指出其脆弱性并呼吁进一步研究和投资。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07544 2025-12-09 cs.CL cs.AI 62%

MoCoRP: Modeling Consistent Relations between Persona and Response for Persona-based Dialogue

MoCoRP: 模型对话中人设与回应之间的一致性关系

Kyungro Lee, Dongha Choi, Hyunju Lee

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 MoCoRP通过显式建模人设与回应之间的一致性关系,提升基于人设的对话生成质量与一致性。

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20993 2025-12-09 cs.LG cs.AI 62%

Subgoal Graph-Augmented Planning for LLM-Guided Open-World Reinforcement Learning

子目标图增强的规划用于LLM引导的开放世界强化学习

Shanwei Fan, Bin Zhang, Zhiwei Xu, Yingxuan Teng, Siqi Dai, Lin Cheng, Guoliang Fan

机构 * The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) School of Artificial Intelligence, Shandong University(山东大学人工智能学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出SGA-ACR框架,通过整合环境特定的子目标图和多LLM规划流程,解决LLM在开放世界强化学习中的规划-执行对齐问题,提升子目标的可行性和可验证性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07694 2025-12-09 cs.CL 57%

Automated Generation of Custom MedDRA Queries Using SafeTerm Medical Map

基于SafeTerm医学地图的自动生成定制MedDRA查询

Francois Vandenhende, Anna Georgiou, Michalis Georgiou, Theodoros Psaras, Ellie Karekla, Elena Hadjicosta

机构 * ClinBAY Limited(ClinBAY有限公司)

专题命中 其他安全 :safety(abstract);分类 cs.CL

AI总结 SafeTerm系统通过多维向量空间和相似度计算,实现自动生成MedDRA查询,优化查询生成的精度与召回率。

Comments 12 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07552 2025-12-09 cs.CL 57%

Performance of the SafeTerm AI-Based MedDRA Query System Against Standardised MedDRA Queries

SafeTerm基于AI的MedDRA查询系统在标准化MedDRA查询中的表现

Francois Vandenhende, Anna Georgiou, Michalis Georgiou, Theodoros Psaras, Ellie Karekla, Elena Hadjicosta

机构 * ClinBAY Limited(ClinBAY有限公司)

专题命中 其他安全 :safety(abstract);分类 cs.CL

AI总结 SafeTerm基于AI的MedDRA查询系统在标准化MedDRA查询中表现出良好性能,通过多标准统计方法实现高召回率和精度平衡。

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06837 2025-12-09 cs.LG 57%

Neural Factorization-based Bearing Fault Diagnosis

基于神经分解的轴承故障诊断

Zhenhao Li, Xu Cheng, Yi Zhou

机构 * College of Computer and Information Science, Southwest University, Chongqing, China(计算机与信息科学学院,西南大学,重庆,中国) College of Vehicle Engineering, Chongqing Industry and Trade Polytechnic, Chongqing, China(车辆工程学院,重庆工业贸易职业技术学院,重庆,中国)

专题命中 其他安全 :safety(abstract);分类 cs.LG

AI总结 本文提出基于神经分解的轴承故障诊断框架,通过多模式特征嵌入和神经分解融合,提升复杂条件下故障诊断性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06787 2025-12-09 cs.CL 57%

LLM4SFC: Sequential Function Chart Generation via Large Language Models

LLM4SFC: 通过大语言模型生成顺序功能图

Ofek Glick, Vladimir Tchuiev, Marah Ghoummaid, Michal Moshkovitz, Dotan Di-Castro

机构 * Bosch Research(博世研究)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 LLM4SFC通过大语言模型生成可执行的顺序功能图,实现图形与文本PLC语言的高效转换。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06688 2025-12-09 cs.CL 57%

PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory

PersonaMem-v2:通过学习隐式用户人设和代理记忆实现个性化智能

Bowen Jiang, Yuan Yuan, Maohao Shen, Zhuoqun Hao, Zhangchen Xu, Zichen Chen, Ziyi Liu, Anvesh Rao Vijjini, Jiashu He, Hanchao Yu, Radha Poovendran, Gregory Wornell, Lyle Ungar, Dan Roth, Sihao Chen, Camillo Jose Taylor

机构 * University of Pennsylvania(宾夕法尼亚大学) Massachusetts Institute of Technology(麻省理工学院) University of Washington(华盛顿大学) University of California Santa Barbara(加州大学圣巴巴拉分校) Meta University of Southern California(南加州大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) Microsoft Corporation(微软公司)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 PersonaMem-v2通过学习隐式用户人设和代理记忆提升LLM个性化能力,实验显示强化微调使模型在隐式个性化任务中准确率达53%。

Comments Data is available at https://huggingface.co/datasets/bowen-upenn/PersonaMem-v2

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07504 2025-12-09 cs.CV 50%

ControlVP: Interactive Geometric Refinement of AI-Generated Images with Consistent Vanishing Points

ControlVP: 交互式几何校正AI生成图像以保持一致的消失点

Ryota Okumura, Kaede Shiohara, Toshihiko Yamasaki

机构 * The University of Tokyo(东京大学)

专题命中 其他安全 :alignment(abstract)

AI总结 ControlVP通过结合建筑轮廓的结构指导和几何约束,校正AI生成图像中的消失点不一致,提升空间结构真实性。

Comments Accepted to WACV 2026, 8 pages, supplementary included. Dataset and code: https://github.com/RyotaOkumura/ControlVP

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06662 2025-12-09 cs.CV 50%

Personalized Image Descriptions from Attention Sequences

基于注意力序列的个性化图像描述

Ruoyu Xue, Hieu Le, Jingyi Xu, Sounak Mondal, Abe Leite, Gregory Zelinsky, Minh Hoai, Dimitris Samaras

机构 * Stony Brook University(石溪大学) UNC-Charlotte(北卡罗来纳大学夏洛特分校) The University of Adelaide(阿德莱德大学)

专题命中 其他安全 :alignment(abstract)

AI总结 DEPER通过建模个性化注意力序列,实现了基于视觉-语言模型的个性化图像描述生成,提升了描述质量与人类对齐性。

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06017 2025-12-09 cs.RO eess.IV 50%

Training-Free Robot Pose Estimation using Off-the-Shelf Foundational Models

无需训练的机器人姿态估计使用现成的基础模型

Laurence Liang

机构 * McGill University(麦吉尔大学)

专题命中 其他安全 :safety(abstract)

AI总结 本文提出利用现成的基础模型无需训练即可估计机器人关节角度,通过实验验证了模型性能基准及测试扩展对预测效果的影响。

Comments Accepted at CVIS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01371 2025-12-09 cs.CV 50%

CLIP-UP: CLIP-Based Unanswerable Problem Detection for Visual Question Answering

CLIP-UP: 基于CLIP的视觉问答中不可回答问题检测

Ben Vardi, Oron Nir, Ariel Shamir

专题命中 其他安全 :alignment(abstract)

AI总结 CLIP-UP通过基于CLIP的相似性度量,为视觉问答模型提供检测不可回答问题的能力,提升模型在不可回答问题上的识别性能。

Comments Revised version. Improvements include support for open-ended VQA and an alternative injection method

详情

展开后加载摘要…

URL PDF HTML 收藏