arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-02-26 至 2026-02-26 共收录 8 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 8 篇

2602.21829 2026-02-26 cs.CV cs.AI 79%

StoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles

StoryMovie: 一个用于视觉故事语义对齐的语料库,结合电影剧本和字幕

Daniel Oliveira, David Martins de Matos

机构 * INESC-ID(INESC-ID研究所) Instituto Superior Técnico, Universidade de Lisboa(里斯本大学技术学院)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

AI总结 StoryMovie通过结合电影剧本和字幕,提升视觉叙事模型的语义对齐能力,使对话归属更准确。

Comments 15 pages, submitted to Journal of Visual Communication and Image Representation

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21595 2026-02-26 cs.RO 78%

SPOC: Safety-Aware Planning Under Partial Observability And Physical Constraints

SPOC:在部分可观测性和物理约束下安全意识的规划

Hyungmin Kim, Hobeom Jeon, Dohyung Kim, Minsu Jang, Jeahong Kim

机构 * 1 ETRI School, University of Science Technology, South Korean 2 Social Robotics Laboratory, Electronics

专题命中 安全评测 :safety(title,abstract)

AI总结 SPOC是一个用于评估安全意识具身任务规划的基准测试,通过整合部分可观测性、物理约束和逐步规划,解决现实环境中安全与可行性评估的问题。

Comments Accepted to IEEE ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22070 2026-02-26 cs.AI 70%

Language Models Exhibit Inconsistent Biases Towards Algorithmic Agents and Human Experts

语言模型表现出对算法代理和人类专家的不一致偏见

Jessica Y. Bo, Lillio Mok, Ashton Anderson

机构 * Computer Science University of Toronto(计算机科学大学 Toronto)

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 研究发现语言模型对人类专家和算法存在不一致偏见,需在高风险应用中谨慎对待。

Comments Second Conference of the International Association for Safe and Ethical Artificial Intelligence (IASEAI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17989 2026-02-26 q-bio.NC cs.AI 70%

The Subject of Emergent Misalignment in Superintelligence: An Anthropological, Cognitive Neuropsychological, Machine-Learning, and Ontological Perspective

超智能中的涌现偏差主体:一种人类学、认知神经心理学、机器学习和本体论视角

Muhammad Osama Imran, Roshni Lulla, Rodney Sappington

机构 * Department of Anthropology, University of Minnesota(明尼苏达大学人类学系) Brain & Creativity Institute, University of Southern California(美国南加州大学脑与创造力研究所) Institute for Advanced Consciousness, Loomis Innovation Center(先进意识研究所) Stimson Center(斯蒂姆森中心)

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 本文从人类学、认知神经心理学等多视角探讨超智能中人类主体与人工智能无意识的相互作用及伦理问题。

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21706 2026-02-26 cs.CV cs.AI 70%

SurGo-R1: Benchmarking and Modeling Contextual Reasoning for Operative Zone in Surgical Video

SurGo-R1:手术视频中操作区的上下文推理基准测试与建模

Guanyi Qin, Xiaozhen Wang, Zhu Zhuo, Chang Han Low, Yuancan Xiao, Yibing Fu, Haofeng Liu, Kai Wang, Chunjiang Li, Yueming Jin

机构 * National University of Singapore, Singapore(新加坡国立大学) Southern Medical University, China(南方医科大学) Guangzhou Research Translation and Innovation Institute, National University of Singapore, China(广州研究翻译与创新研究院,新加坡国立大学,中国)

专题命中 安全评测 :RLHF(abstract);safety(abstract);分类 cs.AI

AI总结 SurGo-R1通过RLHF优化的多阶段架构,在手术视频中实现了高精度的操作区识别,显著优于通用视觉-语言模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19922 2026-02-26 cs.CL cs.AI 62%

HEART: A Unified Benchmark for Assessing Humans and LLMs in Emotional Support Dialogue

HEART:一个评估人类和大语言模型在情感支持对话中能力的统一基准

Laya Iyer, Kriti Aggarwal, Sanmi Koyejo, Gail Heyman, Desmond C. Ong, Subhabrata Mukherjee

机构 * Stanford University(斯坦福大学) University of California, San Diego(加州大学圣地亚哥分校) University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 HEART通过多轮情感支持对话评估人类与大语言模型的能力差异,揭示两者在共情、一致性等维度上的表现及趋同趋势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21841 2026-02-26 cs.CR cs.AI 57%

Resilient Federated Chain: Transforming Blockchain Consensus into an Active Defense Layer for Federated Learning

容错联邦链:将区块链共识转变为联邦学习的主动防御层

Mario García-Márquez, Nuria Rodríguez-Barroso, M. Victoria Luzón, Francisco Herrera

机构 * Department of Computer Science and Artificial Intelligence(计算机科学与人工智能系) Andalusian Research Institute in Data Science and Computational Intelligence (DaSCI)(数据科学与计算智能安达卢西亚研究 institute) University of Granada(格拉纳达大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

AI总结 Resilient Federated Chain通过区块链技术增强联邦学习的对抗性攻击防御能力,提供更安全的去中心化学习环境。

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21251 2026-02-26 cs.SE cs.AI cs.MA cs.PL 57%

AgenticTyper: Automated Typing of Legacy Software Projects Using Agentic AI

AgenticTyper: 使用代理AI自动类型化遗留软件项目

Clemens Pohle

机构 * Darmstadt University of Applied Sciences(达姆斯塔德应用技术大学) MaibornWolff GmbH(马本沃尔夫公司)

专题命中 安全评测 :safety(abstract);分类 cs.AI

AI总结 AgenticTyper利用代理AI自动类型化遗留软件项目,通过迭代错误纠正和行为保留技术,高效解决类型错误问题。

Comments Accepted at ICSE 2026 Student Research Competition (SRC)

详情

展开后加载摘要…

URL PDF HTML 收藏