arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-07-29 至 2025-07-29 共收录 242 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 67 篇

2507.19492 2025-07-29 cs.HC cs.AI cs.CV 77%

ChartGen: Scaling Chart Understanding Via Code-Guided Synthetic Chart Generation

Jovana Kondic, Pengyuan Li, Dhiraj Joshi, Zexue He, Shafiq Abedin, Jennifer Sun, Ben Wiesel, Eli Schwartz, Ahmed Nassar, Bo Wu, Assaf Arbelle, Aude Oliva, Dan Gutfreund, Leonid Karlinsky, Rogerio Feris

机构 * MIT(麻省理工学院) MIT-IBM Watson AI Labs(麻省理工-IBM沃森人工智能实验室) IBM Research(IBM研究院)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14515 2025-07-29 cs.LG cs.AI cs.SI 73%

Efficient Annotator Reliability Assessment and Sample Weighting for Knowledge-Based Misinformation Detection on Social Media

Owen Cook, Charlie Grimshaw, Ben Wu, Sophie Dillon, Jack Hicks, Luke Jones, Thomas Smith, Matyas Szert, Xingyi Song

机构 * School of Computer Science, The University of Sheffield(计算机科学学院,谢菲尔德大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

Comments 8 pages, 3 figures, 3 tables. Code available here: https://github.com/MiniEggz/ruc-misinfo; annotation framework available here: https://github.com/MiniEggz/EffiARA

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19525 2025-07-29 cs.LG cs.AI 73%

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs

Chenchen Zhao, Zhengyuan Shi, Xiangyu Wen, Chengjie Liu, Yi Liu, Yunhao Zhou, Yuxiang Zhao, Hefei Feng, Yinan Zhu, Gwok-Waa Wan, Xin Cheng, Weiyu Chen, Yongqi Fu, Chujie Chen, Chenhao Xue, Guangyu Sun, Ying Wang, Yibo Lin, Jun Yang, Ning Xu, Xi Wang, Qiang Xu

机构 * Department of Computer Science and Engineering, The Chinese University of Hong Kong(中国香港中文大学计算机科学与工程系) School of Electronic Science and Engineering, Nanjing University(南京大学电子科学与工程学院) School of Integrated Circuits, Peking University(北京大学集成电路学院) School of Intergrated Circuits, Southeast University(东南大学集成电路学院) School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) Department of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术系) National Center of Technology Innovation for EDA(EDA技术创新国家中心)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 1 figure, 5 tables. To appear in ICCAD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.00045 2025-07-29 cs.MM cs.AI cs.LG 73%

Detecting Multimedia Generated by Large AI Models: A Survey

Li Lin, Neeraj Gupta, Yue Zhang, Hainan Ren, Chun-Hao Liu, Feng Ding, Xin Wang, Xin Li, Luisa Verdoliva, Shu Hu

机构 * Department of Computer and Information Technology, Purdue University(普渡大学计算机与信息科技系) School of Software, Nanchang University(南昌大学软件学院) Amazon Prime Video(亚马逊Prime视频) Department of Epidemiology and Biostatistics, School of Public Health(公共卫生学院流行病学与生物统计学系) Department of Computer Science, College of Nanotechnology, Science, and Engineering(纳米技术、科学与工程学院计算机科学系) University at Albany, SUNY(阿尔巴尼大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20418 2025-07-29 cs.CV 71%

Can Foundation Models Predict Fitness for Duty?

Juan E. Tapia, Christoph Busch

机构 * da/sec-Biometrics and Internet Security Research Group, Darmstadt, Germany(达姆施塔特生物识别与互联网安全研究组,德国) Hochschule Darmstadt(达姆施塔特应用科学大学)

专题命中 评测与基准 :foundation model(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20994 2025-07-29 cs.CV cs.AI 70%

Security Tensors as a Cross-Modal Bridge: Extending Text-Aligned Safety to Vision in LVLM

Shen Li, Liuyi Yao, Wujia Niu, Lan Zhang, Yaliang Li

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

Comments Codes and data are available at https://github.com/listen0425/Security-Tensors

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20888 2025-07-29 cs.SE cs.CL 70%

Enhancing Project-Specific Code Completion by Inferring Internal API Information

Le Deng, Xiaoxue Ren, Chao Ni, Ming Liang, David Lo, Zhongxin Liu

机构 * State Key Laboratory of Blockchain and Data Security, Zhejiang University(区块链与数据安全国家重点实验室,浙江大学) Ant Group(蚂蚁集团) School of Computing and Information Systems, Singapore Management University(新加坡管理大学计算机与信息系)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20439 2025-07-29 cs.SE cs.AI 70%

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Maya Larbi, Amal Akli, Mike Papadakis, Rihab Bouyousfi, Maxime Cordy, Federica Sarro, Yves Le Traon

机构 * University of Luxembourg(卢森堡大学) University College London(伦敦大学学院)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19694 2025-07-29 math.OC cs.AI cs.GT cs.MA 70%

Ultracoarse Equilibria and Ordinal-Folding Dynamics in Operator-Algebraic Models of Infinite Multi-Agent Games

Faruk Alpay, Hamdi Alakkad, Bugra Kilictas, Taylan Alpay

机构 * Bahcesehir University(巴塞希尔大学) Turkish Aeronautical Association University(土耳其航空协会大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

Comments 15 pages, 2 figures; companion implementation available at https://github.com/farukalpay/ordinal-folding-index/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19500 2025-07-29 cs.HC cs.AI 70%

Gaze-Aware AI: Mathematical modeling of epistemic experience of the Marginalized for Human-Computer Interaction & AI Systems

Omkar Suresh Hatti

机构 * Independent Researcher(独立研究者)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01840 2025-07-29 cs.CL 70%

Minimal Pair-Based Evaluation of Code-Switching

Igor Sterner, Simone Teufel

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19782 2025-07-29 cs.HC 69%

KinemaFX: A Kinematic-Driven Interactive System for Particle Effect Exploration and Customization

Yifei Zhang, Lin-Ping Yuan, Yuheng Zhao, Jielin Feng, Siming Chen

专题命中 评测与基准 :large language model(abstract);language model(abstract);LLM(comments)

Comments Meta Review Overall Rating 3.5 Weakly Accept Contribution to HCI This paper presents KinemaFX, an LLM-powered interactive system leveraging semantic and kinematic inputs to help non-experts explore, customize, and compose particle effects

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21012 2025-07-29 cs.HC 67%

User-Centered Design with AI in the Loop: A Case Study of Rapid User Interface Prototyping with "Vibe Coding"

Tianyi Li, Tanay Maheshwari, Alex Voelker

专题命中 评测与基准 :large language model(abstract);language model(abstract)

Journal ref ACM Collective Intelligence 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14919 2025-07-29 cs.CV 67%

GenM$^3$: Generative Pretrained Multi-path Motion Model for Text Conditional Human Motion Generation

Junyu Shi, Lijiang Liu, Yong Sun, Zhiyuan Zhang, Jinni Zhou, Qiang Nie

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))

专题命中 评测与基准 :large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19975 2025-07-29 cs.RO cs.AI cs.LG 62%

A roadmap for AI in robotics

Aude Billard, Alin Albu-Schaeffer, Michael Beetz, Wolfram Burgard, Peter Corke, Matei Ciocarlie, Ravinder Dahiya, Danica Kragic, Ken Goldberg, Yukie Nagai, Davide Scaramuzza

专题命中 评测与基准 :language model(abstract);分类 cs.AI、cs.LG

Journal ref Nature Machine Intelligence (2025): 1-7

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14651 2025-07-29 cs.CL cs.AI 62%

Real-time Factuality Assessment from Adversarial Feedback

Sanxing Chen, Yukun Huang, Bhuwan Dhingra

专题命中 评测与基准 :LLM(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20614 2025-07-29 cs.CL 57%

Before the Outrage: Challenges and Advances in Predicting Online Antisocial Behavior

Anaïs Ollagnier

专题命中 评测与基准 :language model(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19396 2025-07-29 cs.CL 57%

Detection of Adverse Drug Events in Dutch clinical free text documents using Transformer Models: benchmark study

Rachel M. Murphy, Nishant Mishra, Nicolette F. de Keizer, Dave A. Dongelmans, Kitty J. Jager, Ameen Abu-Hanna, Joanna E. Klopotowska, Iacer Calixto

专题命中 评测与基准 :language model(abstract);分类 cs.CL

Comments 30 Pages, 5 Figures (Main Paper), 19 Pages, 2 Figures(Supplements). Rachel M. Murphy and Nishant Mishra are shared first authors. Joanna E. Klopotowska and Iacer Calixto are shared last authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13524 2025-07-29 cs.CR cs.AI 57%

Contextualized AI for Cyber Defense: An Automated Survey using LLMs

Christoforus Yoga Haryanto, Anne Maria Elvira, Trung Duc Nguyen, Minh Hieu Vu, Yoshiano Hartanto, Emily Lomempow, Arathi Arakala

机构 * School of Science RMIT University(科学学院 莡米大学)

专题命中 评测与基准 :LLM(abstract);分类 cs.AI

Comments 8 pages, 2 figures, 4 tables, accepted into 17th International Conference on Security of Information and Networks (SINCONF 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03328 2025-07-29 cs.CV cs.AI cs.NE 57%

Visual Enumeration Remains Challenging for Multimodal Generative AI

Alberto Testolin, Kuinan Hou, Marco Zorzi

机构 * Department of General Psychology and Department of Mathematics University of Padova(帕多瓦大学心理学系和数学系) Department of General Psychology University of Padova(帕多瓦大学心理学系) Department of General Psychology and Padova Neuroscience Center University of Padova(帕多瓦大学心理学系和帕多瓦神经科学中心) IRCSS San Camillo Hospital, Venice-Lido(威尼斯利多医院IRCSS桑卡莫医院)

专题命中 评测与基准 :foundation model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19657 2025-07-29 cs.NI cs.AI 57%

"X of Information'' Continuum: A Survey on AI-Driven Multi-dimensional Metrics for Next-Generation Networked Systems

Beining Wu, Jun Huang, Shui Yu

机构 * Department of Electrical Engineering and Computer Science, South Dakota State University(南达科他州立大学电气工程与计算机科学系) School of Computer Science, University of Technology Sydney(悉尼大学计算机科学学院)

专题命中 评测与基准 :LLM(abstract);分类 cs.AI

Comments 48 pages, 14 figures, submitted to IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07317 2025-07-29 cs.CV 50%

ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing Evaluation

Sherry X. Chen, Yi Wei, Luowei Zhou, Suren Kumar

机构 * Samsung AI Center(三星AI中心) Mountain View University of California, Santa Barbara(山景城加州大学圣巴巴拉分校)

专题命中 评测与基准 :language model(abstract)

Comments International Conference on Computer Vision (ICCV) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18796 2025-07-29 cs.HC cs.CY 50%

Artificial Intelligence Can Emulate Human Normative Judgments on Emotional Visual Scenes

Zaira Romeo, Alberto Testolin

专题命中 评测与基准 :language model(abstract)

Journal ref Royal Society Open Science, 12(7), 250128 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20397 2025-07-29 cs.CV 50%

VESPA: Towards un(Human)supervised Open-World Pointcloud Labeling for Autonomous Driving

Levente Tempfli, Esteban Rivera, Markus Lienkamp

机构 * Technical University of Munich(慕尼黑技术大学)

专题命中 评测与基准 :language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20286 2025-07-29 cs.CV cs.MM 50%

T$^\text{3}$SVFND: Towards an Evolving Fake News Detector for Emergencies with Test-time Training on Short Video Platforms

Liyuan Zhang, Zeyun Cheng, Yan Yang, Yong Liu, Jinke Ma

机构 * Heilongjiang University(黑龙江大学)

专题命中 评测与基准 :language model(abstract)

Comments 16 pages, 3 figures, published to DASFAA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20017 2025-07-29 cs.CV 50%

VAMPIRE: Uncovering Vessel Directional and Morphological Information from OCTA Images for Cardiovascular Disease Risk Factor Prediction

Lehan Wang, Hualiang Wang, Chubin Ou, Lushi Chen, Yunyi Liang, Xiaomeng Li

机构 * The Hong Kong University of Science and Technology, Hong Kong, China(香港科学与技术大学) Department of Radiology, Guangdong Provincial People's Hospital (Guangdong Academy of Medical Sciences), Southern Medical University, Guangzhou, China(广东省人民医院放射科(广东省医学科学院)南方医科大学) Health Management Center, Foshan First People's Hospital(佛山第一人民医院健康管理中心) Health Management Center, The Sixth Affiliated Hospital, School of Medicine, South China University of Technology(健康管理中心,华南理工大学医学院第六附属医院)

专题命中 评测与基准 :foundation model(abstract)

Comments Accepted in MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05390 2025-07-29 cs.CV eess.IV 50%

From General to Specialized: The Need for Foundational Models in Agriculture

Vishal Nedungadi, Xingguo Xiong, Aike Potze, Ron Van Bree, Tao Lin, Marc Rußwurm, Ioannis N. Athanasiadis

机构 * Wageningen University and Research(瓦赫宁根大学和研究学院) Zhejiang University(浙江大学)

专题命中 评测与基准 :foundation model(abstract)

Comments Accepted to the SEA Workshop (Sustainability with Earth Observation & AI) at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 效率与部署 49 篇

2408.08554 2025-07-29 cs.LG 92%

ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language Models

Chao Zeng, Songwei Liu, Yusheng Xie, Hong Liu, Xiaojian Wang, Miao Wei, Shu Yang, Fangmin Chen, Xing Mei

机构 * Project Leader(项目负责人)

专题命中 效率与部署 :LLM(title,abstract);large language model(title,abstract);language model(title,abstract);post-training(abstract)

Comments AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20786 2025-07-29 cs.CL 89%

Automating Thematic Review of Prevention of Future Deaths Reports: Replicating the ONS Child Suicide Study using Large Language Models

Sam Osian, Arpan Dutta, Sahil Bhandari, Iain E. Buchan, Dan W. Joyce

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL

Comments 8 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20749 2025-07-29 cs.CL cs.CV 89%

Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study

Yiran Huang, Lukas Thede, Massimiliano Mancini, Wenjia Xu, Zeynep Akata

机构 * Technical University of Munich, Germany(慕尼黑技术大学,德国) Helmholtz Munich, Munich Center for Machine Learning, Germany(海德堡慕尼黑,慕尼黑机器学习中心,德国) University of Tübingen, Tübingen AI Center, Germany(图宾根大学,图宾根人工智能中心,德国) University of Trento, Italy(特伦托大学,意大利) Beijing University of Posts and Telecommunications, China(北京邮电大学,中国)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);small language model(abstract);分类 cs.CL

Comments Accepted at GCPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏