arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8057 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8057 篇

2412.07781 2024-12-12 cs.HC cs.LG 57%

Can LLMs faithfully generate their layperson-understandable 'self'?: A Case Study in High-Stakes Domains

Arion Das, Asutosh Mishra, Amitesh Patel, Soumilya De, V. Gurucharan, Kripabandhu Ghosh

机构 * Indian Institute of Information Technology, Ranchi(印度兰契信息技术学院) Indian Institute of Science Education and Research, Berhampur(印度贝汉布尔科学教育与研究学院) Indian Institute of Science Education and Research, Mohanpur(印度莫汉布尔科学教育与研究学院) Collaborative Dynamics, Texas, United States(美国得克萨斯州协作动力公司)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments 35 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07177 2024-12-11 cs.LG 57%

Effective Reward Specification in Deep Reinforcement Learning

Julien Roy

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07112 2024-12-11 cs.CV cs.CL 57%

Maya: An Instruction Finetuned Multilingual Multimodal Model

Nahid Alam, Karthik Reddy Kanjula, Surya Guthikonda, Timothy Chung, Bala Krishna S Vegesna, Abhipsha Das, Anthony Susevski, Ryan Sze-Yin Chan, S M Iftekhar Uddin, Shayekh Bin Islam, Roshan Santhosh, Snegha A, Drishti Sharma, Chen Liu, Isha Chaturvedi, Genta Indra Winata, Ashvanth. S, Snehanshu Mukherjee, Alham Fikri Aji

机构 * Cohere For AI Community(Cohere 人工智能社区) Indiana University Bloomington(印第安纳大学伯明顿分校) Imperial College London(伦敦帝国学院) Georgia Institute of Technology(佐治亚理工学院) The Alan Turing Institute(艾伦·图灵研究所) Bangladesh University of Engineering and Technology(孟加拉工程技术大学) University of Pennsylvania(宾夕法尼亚大学) IIT Bombay(印度理工学院孟买分校) TU Darmstadt(达姆施塔特工业大学) IIT Dhanbad(印度理工学院丹巴德分校) MBZUAI(穆罕默德·本·扎耶德人工智能大学) Cisco Meraki(思科 Meraki) Articul8 AI(Articul8 人工智能公司) Capital One(第一资本金融公司)

专题命中 其他安全 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06795 2024-12-11 cs.NE cs.AI 57%

SpikeFI: A Fault Injection Framework for Spiking Neural Networks

Theofilos Spyrou, Said Hamdioui, Haralampos-G. Stratigopoulos

机构 * Delft University of Technology(代尔夫特理工大学) Sorbonne Université(索邦大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03425 2024-12-11 cs.LG physics.chem-ph q-bio.BM 57%

Sculpting Molecules in Text-3D Space: A Flexible Substructure Aware Framework for Text-Oriented Molecular Optimization

Kaiwei Zhang, Yange Lin, Guangcheng Wu, Yuxiang Ren, Xuecang Zhang, Bo wang, Xiaoyu Zhang, Weitao Du

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06549 2024-12-10 cs.CV cs.LG 57%

Prediction of Occluded Pedestrians in Road Scenes using Human-like Reasoning: Insights from the OccluRoads Dataset

Melo Castillo Angie Nataly, Martin Serrano Sergio, Salinas Carlota, Sotelo Miguel Angel

机构 * Universidad de Alcalá(阿尔卡拉大学)

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.07951 2024-12-10 cs.CL 57%

Questioning the Survey Responses of Large Language Models

Ricardo Dominguez-Olmedo, Moritz Hardt, Celestine Mendler-Dünner

机构 * Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所) ELLIS Institute Tübingen(ELLIS图宾根研究所) Tübingen AI Center(图宾根人工智能中心)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05876 2024-12-10 cs.CV cs.AI 57%

MG-3D: Multi-Grained Knowledge-Enhanced 3D Medical Vision-Language Pre-training

Xuefeng Ni, Linshan Wu, Jiaxin Zhuang, Qiong Wang, Mingxiang Wu, Varut Vardhanabhuti, Lihai Zhang, Hanyu Gao, Hao Chen

机构 * The Hong Kong University of Science and Technology(香港科技大学) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院) Shenzhen People’s Hospital(深圳市人民医院) The University of Hong Kong(香港大学) Chinese PLA General Hospital(中国人民解放军总医院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 10 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05187 2024-12-09 cs.AI cs.CV cs.HC cs.RO 57%

SurgBox: Agent-Driven Operating Room Sandbox with Surgery Copilot

Jinlin Wu, Xusheng Liang, Xuexue Bai, Zhen Chen

机构 * CAIR, HKISI, CAS(中国科学院香港中文大学信息安全研究所智能计算研究中心) MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统实验室) Peking Union Medical College Hospital(北京协和医院)

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments This work is accepted by IEEE Big Data 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00573 2024-12-09 cs.AI 57%

Opus: A Large Work Model for Complex Workflow Generation

Théo Fagnoni, Bellinda Mesbah, Mahsun Altin, Phillip Kingston

机构 * AppliedAI(应用人工智能公司)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 25 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03882 2024-12-06 cs.CY 57%

A Multi-agent Simulation for the Mass School Shootings

Wei Dai, Yash Singh, Rui Zhang

专题命中 其他安全 :safety(abstract);分类 cs.CY

Comments 10 pages, 6 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02712 2024-12-05 cs.SI cs.CY 57%

Analyzing political stances on Twitter in the lead-up to the 2024 U.S. election

Hazem Ibrahim, Farhan Khan, Hend Alabdouli, Maryam Almatrooshi, Tran Nguyen, Talal Rahwan, Yasir Zaki

专题命中 其他安全 :alignment(abstract);分类 cs.CY

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02588 2024-12-04 cs.IR cs.AI 57%

Explainable CTR Prediction via LLM Reasoning

Xiaohan Yu, Li Zhang, Chong Chen

机构 * Huawei Cloud BU(华为云业务部) Institute of Finance Technology, UCL(伦敦大学学院金融科技学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments WSDM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02574 2024-12-04 cs.RO cs.AI cs.SE 57%

Generating Critical Scenarios for Testing Automated Driving Systems

Trung-Hieu Nguyen, Truong-Giang Vuong, Hong-Nam Duong, Son Nguyen, Hieu Dinh Vo, Toshiaki Aoki, Thu-Trang Nguyen

机构 * Faculty of Information Technology, University of Engineering and Technology, Vietnam National University(越南国家大学工程技术学院信息技术学院) School of Information Science, Japan Advanced Institute of Science and Technology(日本先进科学技术学院信息科学学院)

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02532 2024-12-03 cs.CL 57%

SpecExec: Massively Parallel Speculative Decoding for Interactive LLM Inference on Consumer Devices

Ruslan Svirschevski, Avner May, Zhuoming Chen, Beidi Chen, Zhihao Jia, Max Ryabinin

机构 * Carnegie Mellon University(卡内基梅隆大学) Meta AI Together AI Yandex HSE University(高等经济大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19589 2024-12-02 cs.CL 57%

Can Large Language Models Reason about the Region Connection Calculus?

Anthony G Cohn, Robert E Blackwell

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments 13 pages. arXiv admin note: text overlap with arXiv:2309.15577

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19514 2024-12-02 eess.IV cs.CV cs.LG 57%

Enhancing AI microscopy for foodborne bacterial classification via adversarial domain adaptation across optical and biological variability

Siddhartha Bhattacharya, Aarham Wasit, Mason Earles, Nitin Nitin, Luyao Ma, Jiyoon Yi

机构 * Michigan State University(密歇根州立大学) University of California, Davis(加利福尼亚大学戴维斯分校)

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16554 2024-11-26 cs.LG cs.CV 57%

Generating Out-Of-Distribution Scenarios Using Language Models

Erfan Aasi, Phat Nguyen, Shiva Sreeram, Guy Rosman, Sertac Karaman, Daniela Rus

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) UMass Amherst(马萨诸塞大学阿默斯特分校) TRI(丰田研究所) MIT LIDS(麻省理工学院决策与控制系统实验室)

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16754 2024-11-26 cs.CV cs.AI 57%

Towards Full-scene Domain Generalization in Multi-agent Collaborative Bird's Eye View Segmentation for Connected and Autonomous Driving

Senkang Hu, Zhengru Fang, Yiqin Deng, Xianhao Chen, Yuguang Fang, Sam Kwong

机构 * City University of Hong Kong(香港城市大学) The University of Hong Kong(香港大学) Lingnan University(岭南大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments Accepted by IEEE Transactions on Intelligent Transportation Systems (TITS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15202 2024-11-26 physics.ao-ph cs.LG stat.ML 57%

A Comparison of Machine Learning Algorithms for Predicting Sea Surface Temperature in the Great Barrier Reef Region

Dennis Quayesam, Jacob Akubire, Oliveira Darkwah

机构 * University of Cincinnati(辛辛那提大学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.07187 2024-11-26 cs.LG 57%

UPS: Efficiently Building Foundation Models for PDE Solving via Cross-Modal Adaptation

Junhong Shen, Tanya Marwah, Ameet Talwalkar

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments TMLR 2024; ICML 2024 AI for Science Workshop (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14907 2024-11-25 cs.SD cs.AI eess.AS 57%

DAIRHuM: A Platform for Directly Aligning AI Representations with Human Musical Judgments applied to Carnatic Music

Prashanth Thattai Ravikumar

机构 * Goldsmiths, University of London(伦敦大学金史密斯学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 4 Pages, ICASSP workshop submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13063 2024-11-25 physics.ao-ph cs.LG 57%

A Foundation Model for the Earth System

Cristian Bodnar, Wessel P. Bruinsma, Ana Lucic, Megan Stanley, Anna Vaughan, Johannes Brandstetter, Patrick Garvan, Maik Riechert, Jonathan A. Weyn, Haiyu Dong, Jayesh K. Gupta, Kit Thambiratnam, Alexander T. Archibald, Chun-Chieh Wu, Elizabeth Heider, Max Welling, Richard E. Turner, Paris Perdikaris

机构 * Microsoft Research, AI for Science(微软研究院科学人工智能部门) Silurian AI(硅鱼人工智能公司) University of Amsterdam(阿姆斯特丹大学) University of Cambridge(剑桥大学) JKU Linz(林茨约翰·开普勒大学) Microsoft Corporation(微软公司) National Taiwan University(台湾大学) Alan Turing Institute(艾伦·图灵研究所)

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14163 2024-11-22 cs.LO cs.CV cs.LG 57%

Creating a Formally Verified Neural Network for Autonomous Navigation: An Experience Report

Syed Ali Asadullah Bukhari, Thomas Flinkow, Medet Inkarbekov, Barak A. Pearlmutter, Rosemary Monahan

机构 * Maynooth University(梅努斯大学)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments In Proceedings FMAS2024, arXiv:2411.13215

Journal ref EPTCS 411, 2024, pp. 178-190

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13243 2024-11-21 cs.CV cs.AI 57%

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation

Ziyi Wang, Yanbo Wang, Xumin Yu, Jie Zhou, Jiwen Lu

机构 * Tsinghua University(清华大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments Accepted to NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10214 2024-11-20 cs.LG cs.NA math.NA 57%

Machine Learning Algorithms to Assess Site Closure Time Frames for Soil and Groundwater Contamination

Vu-Anh Le, Haruko Murakami Wainwright, Hansell Gonzalez-Raymat, Carol Eddy-Dilek

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments The paper will be withdrawn to fix some work issues with the sections on Bi-LSTM models

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03808 2024-11-20 cs.CY 57%

The Future of Office and Administrative Support Occupations in the Era of Artificial Intelligence: A Bibliometric Analysis

Priyadarshini R. Pennathur, Valerie Boksa, Arunkumar Pennathur, Andrew Kusiak, Beth Livingston

专题命中 其他安全 :safety(abstract);分类 cs.CY

Comments This work is being submitted to the IEEE for possible publication

Journal ref International Journal of Industrial Ergonomics, Vol 104, 2024, page 103665

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11560 2024-11-19 cs.DC cs.AI 57%

Topology-aware Preemptive Scheduling for Co-located LLM Workloads

Ping Zhang, Lei Su, Jinjie Yang, Xin Chen

机构 * Baichuan-Inc(百川公司)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 17 Pages, 11 Figures, 5 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22587 2024-11-19 cs.CL 57%

Toxicity of the Commons: Curating Open-Source Pre-Training Data

Catherine Arnett, Eliot Jones, Ivan P. Yamshchikov, Pierre-Carl Langlais

专题命中 其他安全 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09601 2024-11-15 cs.AI 57%

Accelerating Knowledge Graph and Ontology Engineering with Large Language Models

Cogan Shimizu, Pascal Hitzler

机构 * Wright State University(莱特州立大学) Kansas State University(堪萨斯州立大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏