arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-12-11 至 2025-12-11 共收录 6 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 6 篇

2512.09148 2025-12-11 cs.CL cs.AI 81%

Detecting Hallucinations in Graph Retrieval-Augmented Generation via Attention Patterns and Semantic Alignment

通过注意力模式和语义对齐检测图检索增强生成中的幻觉

Shanghao Li, Jinda Han, Yibo Wang, Yuanjie Zhu, Zihe Song, Langzhou He, Kenan Kamel A Alghythee, Philip S. Yu

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出通过注意力模式和语义对齐检测GraphRAG中的幻觉,开发了轻量级检测器GGA,提升了系统可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07564 2025-12-11 cs.CV cs.AI cs.CL cs.LG 67%

Toward More Reliable Artificial Intelligence: Reducing Hallucinations in Vision-Language Models

迈向更可靠的人工智能:减少视觉-语言模型中的幻觉

Kassoum Sanogo, Renzo Ardiccioni

机构 * Department of CS AI and Data Science(计算机科学与数据科学系) ESEO Engineering School(ESEO工程学院) Faculty of Law, Economy, Management(法学院、经济与管理学院)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出无需训练的自我纠正框架,通过不确定性引导的视觉再注意力减少视觉-语言模型中的幻觉,提升响应准确性。

Comments 24 pages, 3 figures, 2 tables. Training-free self-correction framework for vision-language models. Code and implementation details will be released at: https://github.com/kassoumsanogo1/self-correcting-vlm-re-Attention.git

Journal ref The 4th National and International Academic Conference Celebrating the 20th Anniversary of Rajapruk University (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09340 2025-12-11 cs.AI cs.CV cs.LG 62%

Visual Categorization Across Minds and Models: Cognitive Analysis of Human Labeling and Neuro-Symbolic Integration

跨心灵与模型的视觉分类:人类标注与神经符号整合的认知分析

Chethana Prasad Kabgere

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文通过对比人类与AI在低分辨率图像标注中的表现,探讨了认知策略与神经符号整合方法的异同,旨在推动更可解释和认知对齐的AI系统发展。

Comments 12 pages, 3 figures. Research manuscript based on the final project for CS6795 (Introduction to Cognitive Science), Georgia Tech

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08938 2025-12-11 cs.HC cs.CY 57%

The Impact of Artificial Intelligence on Strategic Technology Management: A Mixed-Methods Analysis of Resources, Capabilities, and Human-AI Collaboration

人工智能对战略技术管理的影响:资源、能力和人机协作的混合方法分析

Massimo Fascinari, Vincent English

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CY

AI总结 本文通过混合方法分析,探讨AI如何提升战略技术管理的有效性,提出AIbSTM框架,强调人机协作而非自主AI领导。

Comments 32 pages, 1 figure, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09249 2025-12-11 cond-mat.mtrl-sci 50%

Auto-3DPFM: Automating Polarization-Vector Mapping at the Nanoscale

Auto-3DPFM:在纳米尺度上自动化极化矢量映射

Ralph Bulanadi, Marti Checa, Michelle Wang, Franck Rothen, John Lasseter, Sumner B. Harris, Daniel Sando, Valanoor Nagarajan, Liam Collins, Stephen Jesse, Rama Vasudevan, Yongtao Liu

专题命中 幻觉与事实性 :alignment(abstract)

AI总结 Auto-3DPFM通过自动化技术实现纳米尺度极化矢量的高精度表征,提升铁电材料研究的效率与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04679 2025-12-11 cs.HC 50%

MisVisFix: An Interactive Dashboard for Detecting, Explaining, and Correcting Misleading Visualizations using Large Language Models

MisVisFix: 一种利用大语言模型检测、解释和纠正误导性可视化信息的交互式仪表板

Amit Kumar Das, Klaus Mueller

专题命中 幻觉与事实性 :trustworthy(abstract)

AI总结 MisVisFix利用大语言模型提供交互式工具,用于检测、解释和纠正误导性可视化信息,提升数据解读的准确性和可信度。

Comments 11 pages, 6 figures. Accepted at IEEE VIS: Visualization & Visual Analytics 2025 conference, November 2-7, 2025, Vienna, Austria

Journal ref IEEE Transactions on Visualization and Computer Graphics (TVCG), PrePrints 5555, pp. 1-11, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏