arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-23 至 2025-09-23 共收录 7 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 7 篇

2409.17407 2025-09-23 cs.AI cs.CL 73%

Post-hoc Reward Calibration: A Case Study on Length Bias

Zeyu Huang, Zihan Qiu, Zili Wang, Edoardo M. Ponti, Ivan Titov

机构 * University of Edinburgh(爱丁堡大学) Alibaba Group(阿里巴巴集团) INF Technology(INF技术) University of Amsterdam(阿姆斯特丹大学)

专题命中 幻觉与事实性 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17671 2025-09-23 cs.CL cs.AI 62%

Turk-LettuceDetect: A Hallucination Detection Models for Turkish RAG Applications

Selva Taş, Mahmut El Huseyni, Özay Ezerceli, Reyhan Bayraktar, Fatma Betül Terzioğlu

机构 * Hidden for Review(保密)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16696 2025-09-23 cs.CL cs.LG 62%

Decoding Uncertainty: The Impact of Decoding Strategies for Uncertainty Estimation in Large Language Models

Wataru Hashimoto, Hidetaka Kamigaito, Taro Watanabe

机构 * Nara Institute of Science and Technology(奈良科学技术大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted at EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16742 2025-09-23 cs.AI 57%

Sycophancy Mitigation Through Reinforcement Learning with Uncertainty-Aware Adaptive Reasoning Trajectories

Mohammad Beigi, Ying Shen, Parshin Shojaee, Qifan Wang, Zichao Wang, Chandan Reddy, Ming Jin, Lifu Huang

机构 * University of California, Davis(加州大学戴维斯分校) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Virginia Tech(弗吉尼亚理工大学) Meta AI Adobe Research(Adobe研究)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16369 2025-09-23 cs.IR cs.AI cs.CE 57%

Enhancing Financial RAG with Agentic AI and Multi-HyDE: A Novel Approach to Knowledge Retrieval and Hallucination Reduction

Akshay Govind Srinivasan, Ryan Jacob George, Jayden Koshy Joe, Hrushikesh Kant, Harshith M R, Sachin Sundar, Sudharshan Suresh, Rahul Vimalkanth, Vijayavallabh

机构 * Indian Institute of Technology Madras(印度理工学院马德拉斯学院)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

Comments 14 Pages, 8 Tables, 2 Figures. Accepted and to be published in the proceedings of FinNLP, Empirical Methods in Natural Language Processing 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07755 2025-09-23 cs.CL cs.CR 57%

Factuality Beyond Coherence: Evaluating LLM Watermarking Methods for Medical Texts

Rochana Prih Hastuti, Rian Adam Rajagede, Mansour Al Ghanim, Mengxin Zheng, Qian Lou

机构 * University of Central Florida(中央佛罗里达大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL

Comments Accepted at EMNLP 2025 Findings. Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11669 2025-09-23 cs.CV 50%

Co-STAR: Collaborative Curriculum Self-Training with Adaptive Regularization for Source-Free Video Domain Adaptation

Amirhossein Dadashzadeh, Parsa Esmati, Majid Mirmehdi

机构 * University of Bristol, UK(布里斯托大学)

专题命中 幻觉与事实性 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏