arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9380 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9380 篇

2502.04376 2025-02-10 cs.CL cs.AI 62%

MEETING DELEGATE: Benchmarking LLMs on Attending Meetings on Our Behalf

Lingxiang Hu, Shurun Yuan, Xiaoting Qin, Jue Zhang, Qingwei Lin, Dongmei Zhang, Saravan Rajmohan, Qi Zhang

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21035 2025-02-10 cs.LG cs.CL 62%

Beyond Autoregression: Fast LLMs via Self-Distillation Through Time

Justin Deschenaux, Caglar Gulcehre

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02896 2025-02-06 cs.CL cs.AI 62%

A Benchmark for the Detection of Metalinguistic Disagreements between LLMs and Knowledge Graphs

Bradley P. Allen, Paul T. Groth

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 6 pages, 2 tables, to appear in Reham Alharbi, Jacopo de Berardinis, Paul Groth, Albert Meroño-Peñuela, Elena Simperl, Valentina Tamma (eds.), ISWC 2024 Special Session on Harmonising Generative AI and Semantic Web Technologies. CEUR-WS.org (forthcoming), for associated code and data see https://github.com/bradleypallen/trex-metalinguistic-disagreement

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12814 2025-02-06 cs.LG cs.CL cs.CR cs.CV 62%

Dissecting Adversarial Robustness of Multimodal LM Agents

Chen Henry Wu, Rishi Shah, Jing Yu Koh, Ruslan Salakhutdinov, Daniel Fried, Aditi Raghunathan

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.LG

Comments ICLR 2025. Also oral at NeurIPS 2024 Open-World Agents Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02249 2025-02-05 cs.CL cs.AI 62%

Conversation AI Dialog for Medicare powered by Finetuning and Retrieval Augmented Generation

Atharva Mangeshkumar Agrawal, Rutika Pandurang Shinde, Vasanth Kumar Bhukya, Ashmita Chakraborty, Sagar Bharat Shah, Tanmay Shukla, Sree Pradeep Kumar Relangi, Nilesh Mutyam

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments 12 pages

Journal ref ResMilitaris 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01676 2025-02-05 cs.CL cs.CY 62%

Benchmark on Peer Review Toxic Detection: A Challenging Task with a New Dataset

Man Luo, Bradley Peterson, Rafael Gan, Hari Ramalingame, Navya Gangrade, Ariadne Dimarogona, Imon Banerjee, Phillip Howard

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.CY

Comments Accepted to WiML workshop @Neurips 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00894 2025-02-04 cs.CL cs.AI 62%

MorphBPE: A Morpho-Aware Tokenizer Bridging Linguistic Complexity for Efficient LLM Training Across Morphologies

Ehsaneddin Asgari, Yassine El Kheir, Mohammad Ali Sadraei Javaheri

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17183 2025-02-04 cs.CL cs.AI 62%

LLM Evaluation Based on Aerospace Manufacturing Expertise: Automated Generation and Multi-Model Question Answering

Beiming Liu, Zhizhuo Cui, Siteng Hu, Xiaohua Li, Haifeng Lin, Zhengxin Zhang

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments conference paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.19271 2025-02-03 cs.AI cs.LG 62%

Concept-Based Explainable Artificial Intelligence: Metrics and Benchmarks

Halil Ibrahim Aysel, Xiaohao Cai, Adam Prugel-Bennett

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments 17 pages it total, 8 main pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18640 2025-02-03 cs.CL cs.CY cs.SI 62%

Divergent Emotional Patterns in Disinformation on Social Media? An Analysis of Tweets and TikToks about the DANA in Valencia

Iván Arcos, Paolo Rosso, Ramón Salaverría

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.CY

Journal ref Proceedings of the 17th International Conference on Agents and Artificial Intelligence (ICAART 2025), Porto, Portugal, February 23-25, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18006 2025-01-31 cs.LG cs.AI cs.CR 62%

Topological Signatures of Adversaries in Multimodal Alignments

Minh Vu, Geigh Zollicoffer, Huy Mai, Ben Nebgen, Boian Alexandrov, Manish Bhattarai

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12950 2025-01-31 cs.AI cs.CV cs.LG 62%

Beyond the Veil of Similarity: Quantifying Semantic Continuity in Explainable AI

Qi Huang, Emanuele Mezzi, Osman Mutlu, Miltiadis Kofinas, Vidya Prasad, Shadnan Azwad Khan, Elena Ranguelova, Niki van Stein

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 25 pages, accepted at the world conference of explainable AI, 2024, Malta

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.06059 2025-01-31 cs.CY cs.CV cs.LG 62%

Impact on Public Health Decision Making by Utilizing Big Data Without Domain Knowledge

Miao Zhang, Salman Rahman, Vishwali Mhasawade, Rumi Chunara

专题命中 安全评测 :alignment(abstract);分类 cs.CY、cs.LG

Journal ref Proceedings of the National Academy of Sciences 121.39 (2024): e2402387121

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17329 2025-01-30 cs.MA cs.AI cs.LG 62%

Anomaly Detection in Cooperative Vehicle Perception Systems under Imperfect Communication

Ashish Bastola, Hao Wang, Abolfazl Razi

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08980 2025-01-30 cs.HC cs.AI cs.LG 62%

Predicting Trust In Autonomous Vehicles: Modeling Young Adult Psychosocial Traits, Risk-Benefit Attitudes, And Driving Factors With Machine Learning

Robert Kaufman, Emi Lee, Manas Satish Bedmutha, David Kirsh, Nadir Weibel

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 24 pages (including references and appendix), Accepted to CHI Conference on Human Factors in Computing Systems (CHI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16348 2025-01-29 cs.LG cs.AI 62%

An Integrated Approach to AI-Generated Content in e-health

Tasnim Ahmed, Salimur Choudhury

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted for presentation at 2025 IEEE International Conference on Communications (IEEE ICC25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08027 2025-01-27 cs.CY cs.HC cs.LG 62%

iLLuMinaTE: An LLM-XAI Framework Leveraging Social Science Explanation Theories Towards Actionable Student Performance Feedback

Vinitra Swamy, Davide Romano, Bhargav Srinivasa Desikan, Oana-Maria Camburu, Tanja Käser

专题命中 安全评测 :alignment(abstract);分类 cs.CY、cs.LG

Comments Accepted at AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13479 2025-01-24 cs.LG cs.AI 62%

Adaptive Few-Shot Learning (AFSL): Tackling Data Scarcity with Stability, Robustness, and Versatility

Rishabh Agrawal

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16772 2025-01-23 cs.CV cs.CL cs.LG 62%

VisMin: Visual Minimal-Change Understanding

Rabiul Awal, Saba Ahmadi, Le Zhang, Aishwarya Agrawal

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted at NeurIPS 2024. Project URL at https://vismin.net/

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03595 2025-01-23 cs.CL cs.AI 62%

GREEN: Generative Radiology Report Evaluation and Error Notation

Sophie Ostmeier, Justin Xu, Zhihong Chen, Maya Varma, Louis Blankemeier, Christian Bluethgen, Arne Edward Michalson, Michael Moseley, Curtis Langlotz, Akshay S Chaudhari, Jean-Benoit Delbrouck

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Journal ref https://aclanthology.org/2024.findings-emnlp.21/

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09154 2025-01-17 cs.CL cs.AI 62%

Towards Multilingual LLM Evaluation for Baltic and Nordic languages: A study on Lithuanian History

Yevhen Kostiuk, Oxana Vitman, Łukasz Gagała, Artur Kiulian

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08292 2025-01-15 cs.CL cs.AI 62%

HALoGEN: Fantastic LLM Hallucinations and Where to Find Them

Abhilasha Ravichander, Shrusti Ghela, David Wadden, Yejin Choi

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19655 2025-01-14 cs.CL cs.AI 62%

Assessment and manipulation of latent constructs in pre-trained language models using psychometric scales

Maor Reuben, Ortal Slobodin, Aviad Elyshar, Idan-Chaim Cohen, Orna Braun-Lewensohn, Odeya Cohen, Rami Puzis

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06911 2025-01-14 cs.AI cs.CL 62%

Risk-Averse Finetuning of Large Language Models

Sapana Chaudhary, Ujwal Dinesha, Dileep Kalathil, Srinivas Shakkottai

专题命中 安全评测 :RLHF(abstract);分类 cs.CL、cs.AI

Comments Neurips 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06887 2025-01-14 cs.CV cs.AI cs.ET cs.LG 62%

MedGrad E-CLIP: Enhancing Trust and Transparency in AI-Driven Skin Lesion Diagnosis

Sadia Kamal, Tim Oates

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Accepted to 2025 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06432 2025-01-14 cs.LG cs.AI 62%

Deep Learning on Hester Davis Scores for Inpatient Fall Prediction

Hojjat Salehinejad, Ricky Rojas, Kingsley Iheasirim, Mohammed Yousufuddin, Bijan Borah

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted for presentation at IEEE SSCI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05566 2025-01-13 cs.CV cs.AI cs.CY 62%

Vision-Language Models for Autonomous Driving: CLIP-Based Dynamic Scene Understanding

Mohammed Elhenawy, Huthaifa I. Ashqar, Andry Rakotonirainy, Taqwa I. Alhadidi, Ahmed Jaber, Mohammad Abu Tami

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11484 2025-01-13 cs.AI cs.CL 62%

The Oscars of AI Theater: A Survey on Role-Playing with Language Models

Nuo Chen, Yan Wang, Yang Deng, Jia Li

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 28 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19432 2025-01-07 cs.LG cs.AI 62%

MicroFlow: An Efficient Rust-Based Inference Engine for TinyML

Matteo Carnelos, Francesco Pasti, Nicola Bellotto

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00522 2025-01-03 cs.CL cs.AI 62%

TinyHelen's First Curriculum: Training and Evaluating Tiny Language Models in a Simpler Language Environment

Ke Yang, Volodymyr Kindratenko, ChengXiang Zhai

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏