arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-12-11 至 2025-12-11 共收录 35 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 2 篇

2508.06783 2025-12-11 cs.LG cs.AI cs.CR cs.IT math.IT 86%

PROPS: Progressively Private Self-alignment of Large Language Models

PROPS: 大型语言模型的逐步隐私自对齐

Noel Teku, Fengwei Tian, Payel Bhattacharjee, Souradip Chakraborty, Amrit Singh Bedi, Ravi Tandon

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.AI、cs.LG

AI总结 PROPS通过多阶段隐私保护对齐框架,在保护偏好标签隐私的同时提升LLM对齐效果,实现更高的胜率。

Comments Accepted in the Transactions on Machine Learning Research (TMLR), 2025

Journal ref Transactions on ML Research (TMLR) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14256 2025-12-11 cs.AI cs.IR 57%

PathMind: A Retrieve-Prioritize-Reason Framework for Knowledge Graph Reasoning with Large Language Models

PathMind: 一种基于大语言模型的知识图谱推理检索-优先-推理框架

Yu Liu, Xixun Lin, Yanmin Shang, Yangxi Li, Shi Wang, Yanan Cao

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI

AI总结 PathMind通过检索-优先-推理框架,提升大语言模型在知识图谱推理中的准确性和可解释性,尤其在复杂推理任务中表现优异。

Comments AAAI 2026, Long Paper, Oral

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 4 篇

2409.10506 2025-12-11 cs.SE 78%

SmartC2Rust: Iterative, Feedback-Driven C-to-Rust Translation via Large Language Models for Safety and Equivalence

SmartC2Rust: 基于大语言模型的迭代反馈驱动C到Rust翻译用于安全性和等价性

Momoko Shiraishi, Yinzhi Cao, Takahiro Shinagawa

专题命中 安全训练 :safety(title,abstract)

AI总结 SmartC2Rust通过迭代反馈机制利用大语言模型实现C到Rust的翻译,提升内存安全性和语义等价性。

Journal ref ICSE '26: Proceedings of the 48th International Conference on Software Engineering, April 12-18, 2026, Rio de Janeiro, Brazil

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09909 2025-12-11 cs.LG cs.AI 62%

STACHE: Local Black-Box Explanations for Reinforcement Learning Policies

STACHE:强化学习策略的局部黑盒解释

Andrew Elashkin, Orna Grumberg

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

AI总结 STACHE提出一种用于生成强化学习策略局部黑盒解释的框架,通过鲁棒区域和最小反事实分析,揭示策略行为的演变过程。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09127 2025-12-11 cs.CL cs.AI 62%

Knowledge-Guided Large Language Model for Automatic Pediatric Dental Record Understanding and Safe Antibiotic Recommendation

基于知识引导的大型语言模型用于自动解析儿童牙科记录并安全推荐抗生素

Zihan Han, Junyan Ge, Caifeng Li

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

AI总结 本研究提出基于知识引导的大型语言模型,通过整合知识图谱、检索增强生成和多阶段安全验证,提升儿童牙科记录解析与安全抗生素推荐的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08952 2025-12-11 cs.LG cs.AI cs.HC cs.RO 62%

Learning When to Ask: Simulation-Trained Humanoids for Mental-Health Diagnosis

学习何时提问:为心理健康诊断训练的仿真 humanoid

Filippo Cenacchi, Deborah Richards, Longbing Cao

机构 * Macquarie University(麦考瑞大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种基于仿真的方法,通过训练人形机器人以更有效地进行心理健康诊断,利用反事实回放和安全学习循环提升对话节奏和关系建立。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 1 篇

2512.09403 2025-12-11 cs.LG 89%

Black-Box Behavioral Distillation Breaks Safety Alignment in Medical LLMs

黑盒行为蒸馏破坏医疗大语言模型的安全对齐

Sohely Jahan, Ruimin Sun

专题命中 越狱攻击 :alignment(title,abstract);safety(title,abstract);jailbreak(abstract);分类 cs.LG

AI总结 黑盒行为蒸馏攻击揭示医疗大语言模型的安全对齐漏洞,通过低成本手段复制模型能力并破坏安全机制。

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与事实性 6 篇

2512.09148 2025-12-11 cs.CL cs.AI 81%

Detecting Hallucinations in Graph Retrieval-Augmented Generation via Attention Patterns and Semantic Alignment

通过注意力模式和语义对齐检测图检索增强生成中的幻觉

Shanghao Li, Jinda Han, Yibo Wang, Yuanjie Zhu, Zihe Song, Langzhou He, Kenan Kamel A Alghythee, Philip S. Yu

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出通过注意力模式和语义对齐检测GraphRAG中的幻觉,开发了轻量级检测器GGA,提升了系统可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07564 2025-12-11 cs.CV cs.AI cs.CL cs.LG 67%

Toward More Reliable Artificial Intelligence: Reducing Hallucinations in Vision-Language Models

迈向更可靠的人工智能:减少视觉-语言模型中的幻觉

Kassoum Sanogo, Renzo Ardiccioni

机构 * Department of CS AI and Data Science(计算机科学与数据科学系) ESEO Engineering School(ESEO工程学院) Faculty of Law, Economy, Management(法学院、经济与管理学院)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出无需训练的自我纠正框架,通过不确定性引导的视觉再注意力减少视觉-语言模型中的幻觉,提升响应准确性。

Comments 24 pages, 3 figures, 2 tables. Training-free self-correction framework for vision-language models. Code and implementation details will be released at: https://github.com/kassoumsanogo1/self-correcting-vlm-re-Attention.git

Journal ref The 4th National and International Academic Conference Celebrating the 20th Anniversary of Rajapruk University (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09340 2025-12-11 cs.AI cs.CV cs.LG 62%

Visual Categorization Across Minds and Models: Cognitive Analysis of Human Labeling and Neuro-Symbolic Integration

跨心灵与模型的视觉分类:人类标注与神经符号整合的认知分析

Chethana Prasad Kabgere

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文通过对比人类与AI在低分辨率图像标注中的表现,探讨了认知策略与神经符号整合方法的异同,旨在推动更可解释和认知对齐的AI系统发展。

Comments 12 pages, 3 figures. Research manuscript based on the final project for CS6795 (Introduction to Cognitive Science), Georgia Tech

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08938 2025-12-11 cs.HC cs.CY 57%

The Impact of Artificial Intelligence on Strategic Technology Management: A Mixed-Methods Analysis of Resources, Capabilities, and Human-AI Collaboration

人工智能对战略技术管理的影响:资源、能力和人机协作的混合方法分析

Massimo Fascinari, Vincent English

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CY

AI总结 本文通过混合方法分析,探讨AI如何提升战略技术管理的有效性,提出AIbSTM框架,强调人机协作而非自主AI领导。

Comments 32 pages, 1 figure, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09249 2025-12-11 cond-mat.mtrl-sci 50%

Auto-3DPFM: Automating Polarization-Vector Mapping at the Nanoscale

Auto-3DPFM:在纳米尺度上自动化极化矢量映射

Ralph Bulanadi, Marti Checa, Michelle Wang, Franck Rothen, John Lasseter, Sumner B. Harris, Daniel Sando, Valanoor Nagarajan, Liam Collins, Stephen Jesse, Rama Vasudevan, Yongtao Liu

专题命中 幻觉与事实性 :alignment(abstract)

AI总结 Auto-3DPFM通过自动化技术实现纳米尺度极化矢量的高精度表征,提升铁电材料研究的效率与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04679 2025-12-11 cs.HC 50%

MisVisFix: An Interactive Dashboard for Detecting, Explaining, and Correcting Misleading Visualizations using Large Language Models

MisVisFix: 一种利用大语言模型检测、解释和纠正误导性可视化信息的交互式仪表板

Amit Kumar Das, Klaus Mueller

专题命中 幻觉与事实性 :trustworthy(abstract)

AI总结 MisVisFix利用大语言模型提供交互式工具,用于检测、解释和纠正误导性可视化信息,提升数据解读的准确性和可信度。

Comments 11 pages, 6 figures. Accepted at IEEE VIS: Visualization & Visual Analytics 2025 conference, November 2-7, 2025, Vienna, Austria

Journal ref IEEE Transactions on Visualization and Computer Graphics (TVCG), PrePrints 5555, pp. 1-11, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 安全评测 9 篇

2211.11963 2025-12-11 cs.RO cs.LG 79%

Learning-based social coordination to improve safety and robustness of cooperative autonomous vehicles in mixed traffic

基于学习的社会协调以提高混合交通中协作自动驾驶车辆的安全性和鲁棒性

Rodolfo Valiente, Behrad Toghi, Mahdi Razzaghpour, Ramtin Pedarsani, Yaser P. Fallah

专题命中 安全评测 :safety(title,abstract);分类 cs.LG

AI总结 本文提出基于多智能体强化学习的方法,通过量化自动驾驶车辆的社会偏好并引入利他主义,提升其在混合交通中与人类驾驶车辆协作的安全性和鲁棒性。

Comments arXiv admin note: substantial text overlap with arXiv:2202.00881

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02603 2025-12-11 eess.AS cs.CL cs.SD 74%

SEAL: Speech Embedding Alignment Learning for Speech Large Language Model with Retrieval-Augmented Generation

SEAL:用于带有检索增强生成的语音大语言模型的语音嵌入对齐学习

Chunyu Sun, Bingyu Liu, Zhichao Cui, Junhan Shi, Anbin Qi, Tian-hao Zhang, Dinghao Zhou, Lewei Lu

机构 * SenseTime Research(商汤科技研究院)

专题命中 安全评测 :alignment(title);分类 cs.CL

AI总结 SEAL通过统一的语音和文本嵌入框架,提升语音大语言模型的检索效率与准确性,减少延迟并增强多模态检索能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08936 2025-12-11 cs.HC cs.AI cs.CY 73%

A Principle-based Framework for the Development and Evaluation of Large Language Models for Health and Wellness

基于原则的大型语言模型在健康与健身领域开发与评估框架

Brent Winslow, Jacqueline Shreibati, Javier Perez, Hao-Wei Su, Nichole Young-Lin, Nova Hammerquist, Daniel McDuff, Jason Guss, Jenny Vafeiadou, Nick Cain, Alex Lin, Erik Schenck, Shiva Rajagopal, Jia-Ru Chung, Anusha Venkatakrishnan, Amy Armento Lee, Maryam Karimzadehgan, Qingyou Meng, Rythm Agarwal, Aravind Natarajan, Tracy Giest

机构 * Google Research(谷歌研究)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.CY

AI总结 本文提出基于原则的框架,用于系统评估LLM在健康与健身领域的应用,通过迭代开发生命周期整合综合评估技术,提升系统安全性和用户反馈的可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09552 2025-12-11 cs.CL 57%

Systematic Framework of Application Methods for Large Language Models in Language Sciences

大型语言模型在语言科学中应用方法的系统框架

Kun Sun, Rong Wang

机构 * Tongji University, China(同济大学) University of Tübingen, Germany(图宾根大学) Institute of Natural Language Processing, University of Stuttgart, Germany(斯图加特大学自然语言处理研究所)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 本文提出两个系统框架,指导LLMs在语言科学中的战略应用,通过方法选择和多阶段研究流程提升研究的可重复性和科学性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08352 2025-12-11 cs.CV cs.AI 57%

Evaluating Small Vision-Language Models on Distance-Dependent Traffic Perception

评估小型视觉-语言模型在距离依赖性交通感知中的表现

Nikos Theodoridis, Tim Brophy, Reenu Mohandas, Ganesh Sistu, Fiachra Collins, Anthony Scanlan, Ciaran Eising

机构 * University of Limerick(利默里克大学) Data Driven Computer Engineering Research Centre(数据驱动计算机工程研究中心) Lero, The Irish Software Research Centre(Lero爱尔兰软件研究中心) Valeo Vision Systems(瓦莱欧视觉系统)

专题命中 安全评测 :safety(abstract);分类 cs.AI

AI总结 本文评估了小型视觉-语言模型在距离依赖性交通感知任务中的表现,发现其在感知能力上显著弱于人类。

Comments Published in IEEE Open Journal of Vehicular Technology. Final version available at: https://ieeexplore.ieee.org/document/11230063

Journal ref IEEE Open Journal of Vehicular Technology, vol. 7, pp. 54-72, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.00881 2025-12-11 cs.RO cs.LG 57%

Robustness and Adaptability of Reinforcement Learning based Cooperative Autonomous Driving in Mixed-autonomy Traffic

基于混合自主交通的强化学习协同自动驾驶的鲁棒性与适应性

Rodolfo Valiente, Behrad Toghi, Ramtin Pedarsani, Yaser P. Fallah

专题命中 安全评测 :safety(abstract);分类 cs.LG

AI总结 本文提出了一种基于多智能体强化学习的方法,用于提高自动驾驶车辆在混合自主交通环境中的鲁棒性和适应性,通过隐式学习人类驾驶行为以优化社会效用和安全性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09271 2025-12-11 cs.CV 50%

LongT2IBench: A Benchmark for Evaluating Long Text-to-Image Generation with Graph-structured Annotations

LongT2IBench: 一个用于评估长文本到图像生成的基准,具有图结构注释

Zhichao Yang, Tianjiao Gu, Jianjie Wang, Feiyu Lin, Xiangfei Sheng, Pengfei Chen, Leida Li

专题命中 安全评测 :alignment(abstract)

AI总结 LongT2IBench通过图结构注释和LongT2IExpert提出,用于评估长文本到图像生成的对齐和解释能力。

Comments The paper has been accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09092 2025-12-11 cs.CV 50%

Explaining the Unseen: Multimodal Vision-Language Reasoning for Situational Awareness in Underground Mining Disasters

解释未见的:多模态视觉-语言推理用于地下矿难中的情境感知

Mizanur Rahman Jewel, Mohamed Elmahallawy, Sanjay Madria, Samuel Frimpong

机构 * Missouri University of Science and Technology(密苏里科学与技术大学) Washington State University(华盛顿州立大学)

专题命中 安全评测 :alignment(abstract)

AI总结 本文提出MDSE框架,通过多模态视觉-语言推理提升地下矿难情境感知能力,实现更准确的描述生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17353 2025-12-11 cs.CE 50%

RoadBench: A Vision-Language Foundation Model and Benchmark for Road Damage Understanding

RoadBench: 一种用于道路损伤理解的视觉-语言基础模型和基准

Xi Xiao, Yunbei Zhang, Janet Wang, Lin Zhao, Yuxiang Wei, Hengjia Li, Yanshu Li, Xinyuan Song, Xiao Wang, Swalpa Kumar Roy, Hao Xu, Tianyang Wang

专题命中 安全评测 :safety(abstract)

AI总结 RoadBench通过整合视觉与文本信息,提出RoadCLIP模型,显著提升道路损伤识别性能,为基础设施监测提供新基准。

Comments Accepted by WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

6. AI治理与伦理 5 篇

2510.06249 2025-12-11 cs.CL cs.AI 81%

TRepLiNa: Layer-wise CKA+REPINA Alignment Improves Low-Resource Machine Translation in Aya-23 8B

TRepLiNa:分层CKA+REPINA对齐改进Aya-23 8B低资源机器翻译

Toshiki Nakai, Ravi Kiran Chikkala, Lena Sophie Oberkircher, Nicholas Jennings, Natalia Skachkova, Tatiana Anikina, Jesujoba Oluwadara Alabi

机构 * Saarland University(萨尔兰大学) German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 TRepLiNa通过结合CKA和REPINA实现分层对齐,提升低资源语言Aya-23 8B的机器翻译质量,尤其在数据稀缺情况下效果显著。

Comments It is work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09483 2025-12-11 cs.CL cs.CY 62%

Source Coverage and Citation Bias in LLM-based vs. Traditional Search Engines

基于大语言模型的搜索引擎与传统搜索引擎的来源覆盖与引用偏见

Peixian Zhang, Qiming Ye, Zifan Peng, Kiran Garimella, Gareth Tyson

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Rutgers University(罗格斯大学) Rutgers University New Brunswick United States(罗格斯大学新 Brunswick美国)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.CY

AI总结 本文研究了基于大语言模型的搜索引擎与传统搜索引擎在来源覆盖和引用偏见方面的差异,发现LLM-SEs在资源多样性上优于传统搜索引擎,但其可信度和中立性仍需进一步提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09458 2025-12-11 cs.AI cs.LG 62%

Architectures for Building Agentic AI

构建代理AI的架构

Sławomir Nowaczyk

机构 * Center for Applied Intelligent Systems Research, Halmstad University, Sweden(应用智能系统研究所,哈尔姆斯塔德大学,瑞典)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种构建代理AI的架构分类,通过组件化设计和显式控制循环提升系统可靠性。

Comments This is a preprint of a chapter accepted for publication in Generative and Agentic AI Reliability: Architectures, Challenges, and Trust for Autonomous Systems, published by Springer Nature

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08978 2025-12-11 cs.CY cs.AI 62%

Institutional AI Sovereignty Through Gateway Architecture: Implementation Report from Fontys ICT

通过网关架构实现机构AI主权:Fontys ICT的实施报告

Ruud Huijts, Koen Suilen

机构 * Fontys ICT

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本文提出通过网关架构实现机构AI主权,构建了受监管的AI平台,实现可控的AI访问和治理,强调AI作为战略工具需专门领导和治理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08981 2025-12-11 cs.CV cs.AI 57%

Mitigating Bias with Words: Inducing Demographic Ambiguity in Face Recognition Templates by Text Encoding

通过词语缓解偏见:通过文本编码在人脸识别模板中诱导人口统计学模糊性

Tahar Chettaoui, Naser Damer, Fadi Boutros

机构 * Fraunhofer IGD(弗劳恩霍夫研究所(IGD)) TU Darmstadt(德累斯顿技术大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 通过引入文本编码,UTIE在人脸识别模板中诱导人口统计学模糊性,以减少偏见并提升验证性能。

Comments Accepted at BMVC workshop (SRBS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 其他安全 8 篇

2512.09898 2025-12-11 cs.RO cs.AI cs.CV cs.MA cs.SY eess.SY 57%

Visual Heading Prediction for Autonomous Aerial Vehicles

自主航空器的视觉航向预测

Reza Ahmari, Ahmad Mohammadi, Vahid Hemmati, Mohammed Mynuddin, Parham Kebria, Mahmoud Nabil Mahmoud, Xiaohong Yuan, Abdollah Homaifar

机构 * Department of Computer Science at North Carolina A&T State University(北卡罗来纳A&T州立大学计算机科学系) Department of Electrical and Computer Engineering at North Carolina A&T State University(北卡罗来纳A&T州立大学电气与计算机工程系)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出基于视觉的数据驱动框架,实现无人机与无人地面车辆的实时整合,通过YOLOv5检测UGV并利用轻量级ANN预测航向角,实现高精度的导航与协调。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09895 2025-12-11 cs.AI cs.DL 57%

Human-in-the-Loop and AI: Crowdsourcing Metadata Vocabulary for Materials Science

人机协同与AI:为材料科学众包元数据词汇

Jane Greenberg, Scott McClellan, Addy Ireland, Robert Sammarco, Colton Gerber, Christopher B. Rauch, Mat Kelly, John Kunze, Yuan An, Eric Toberer

机构 * Metadata Research Center, College of Computing and Informatics, Drexel University(元数据研究中心、计算与信息学院、德雷塞尔大学) Penn State University(宾夕法尼亚州立大学) Colorado School of Mines(科罗拉多矿业学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出MatSci-YAMZ平台,通过结合人工智能和人机协同方法,实现材料科学元数据词汇表的众包开发,验证了AI-HILT模型在跨学科领域的可行性与可扩展性。

Comments Metadata and Semantics Research Conference 2025, 14 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08143 2025-12-11 cs.LG 57%

PolyLingua: Margin-based Inter-class Transformer for Robust Cross-domain Language Detection

PolyLingua: 基于边界的跨领域语言识别变换器

Ali Lotfi Rezaabad, Bikram Khanal, Shashwat Chaurasia, Lu Zeng, Dezhi Hong, Hossein Bashashati, Thomas Butler, Megan Ganji

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 PolyLingua是一种轻量级变换器模型,通过两级对比学习框架实现精确的跨领域语言识别,以高准确率和低资源消耗应对复杂语言识别挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏