arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-08-19 至 2025-08-19 共收录 229 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 58 篇

2508.13044 2025-08-19 cs.CL 70%

Büyük Dil Modelleri için TR-MMLU Benchmarkı: Performans Değerlendirmesi, Zorluklar ve İyileştirme Fırsatları

M. Ali Bayram, Ali Arda Fincan, Ahmet Semih Gümüş, Banu Diri, Savaş Yıldırım, Öner Aytaş

机构 * Yıldız Technical University(伊兹密尔技术大学) Yeditepe University(耶迪特佩大学) İstanbul Bilgi University(伊斯坦布尔比尔大学) Işık University(伊斯坎德大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

Comments 10 pages, in Turkish language, 5 figures. Presented at the 2025 33rd Signal Processing and Communications Applications Conference (SIU), 25--28 June 2025, Sile, Istanbul, Türkiye

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12754 2025-08-19 cs.AI 70%

Beyond Ethical Alignment: Evaluating LLMs as Artificial Moral Assistants

Alessio Galatolo, Luca Alberto Rappuoli, Katie Winkle, Meriem Beloucif

机构 * Uppsala University(乌普萨拉大学) University of St. Andrews(圣安德鲁大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

Comments Full version of the paper published in ECAI 2025 proceedings (IOS Press, CC BY-NC 4.0)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12355 2025-08-19 cs.CL 70%

Consensus or Conflict? Fine-Grained Evaluation of Conflicting Answers in Question-Answering

Eviatar Nachshoni, Arie Cattan, Shmuel Amar, Ori Shapira, Ido Dagan

机构 * Bar-Ilan University(巴伊兰大学) Google Research(谷歌研究) OriginAI

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

Comments no comments

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12291 2025-08-19 cs.AI 70%

RadarQA: Multi-modal Quality Analysis of Weather Radar Forecasts

Xuming He, Zhiyuan You, Junchao Gong, Couhua Liu, Xiaoyu Yue, Peiqin Zhuang, Wenlong Zhang, Lei Bai

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ZheJiang University(浙江大学) The Chinese University of Hong Kong(香港中文大学) Center for Earth System Modeling and Prediction of China Meteorological Administration(中国气象局地球系统模拟与预测中心)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09062 2025-08-19 cs.LG 70%

State-Space Modeling in Long Sequence Processing: A Survey on Recurrence in the Transformer Era

Matteo Tiezzi, Michele Casoni, Alessandro Betti, Marco Gori, Stefano Melacci

机构 * DIISM University of Siena(锡耶纳大学DIISM学院) IIT Ist. Italiano di Tecnologia(意大利技术研究所) IMT Scuola Alti Studi(IMT高级研究学院)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.LG

Comments Currently under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.05468 2025-08-19 cs.AI cond-mat.mtrl-sci cs.CV cs.IR 70%

Image and Data Mining in Reticular Chemistry Using GPT-4V

Zhiling Zheng, Zhiguo He, Omar Khattab, Nakul Rampal, Matei A. Zaharia, Christian Borgs, Jennifer T. Chayes, Omar M. Yaghi

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

Comments 36 pages, 24 figures

Journal ref Digital Discovery, 2024,3, 491-501

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12752 2025-08-19 cs.IR 67%

Deep Research: A Survey of Autonomous Research Agents

Wenlin Zhang, Xiaopeng Li, Yingyi Zhang, Pengyue Jia, Yichao Wang, Huifeng Guo, Yong Liu, Xiangyu Zhao

专题命中 评测与基准 :large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12626 2025-08-19 cs.SD 67%

Exploring the Feasibility of LLMs for Automated Music Emotion Annotation

Meng Yang, Jon McCormack, Maria Teresa Llano, Wanchao Su

专题命中 评测与基准 :large language model(abstract);language model(abstract)

Comments Accepted to be published at ISMIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12475 2025-08-19 cs.PL cs.FL 67%

Type-Driven Prompt Programming: From Typed Interfaces to a Calculus of Constraints

Abhijit Paul

专题命中 评测与基准 :large language model(abstract);language model(abstract)

Comments Accepted as Extended Abstract in TyDe Workshop 2025,co-located with ICFP

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19183 2025-08-19 cs.CV 67%

Segmenting Objectiveness and Task-awareness Unknown Region for Autonomous Driving

Mi Zheng, Guanglei Yang, Zitong Huang, Zhenhua Guo, Kevin Han, Wangmeng Zuo

机构 * Harbin Institute of Technology(哈尔滨工业大学) Tianyijiaotong Technology Ltd.(天翼交通科技有限公司) Independent Researcher(独立研究者)

专题命中 评测与基准 :large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.06670 2025-08-19 cs.LG cs.AI 62%

Unveiling the Unseen: A Comprehensive Survey on Explainable Anomaly Detection in Images and Videos

Yizhou Wang, Dongliang Guo, Sheng Li, Octavia Camps, Yun Fu

机构 * Department of Electrical and Computer Engineering, Northeastern University(电气与计算机工程系,东北大学) School of Data Science, University of Virginia(数据科学学院,弗吉尼亚大学) Department of Electrical and Computer Engineering and Khoury College of Computer Science, Northeastern University(电气与计算机工程系和Khoury计算机科学学院,东北大学)

专题命中 评测与基准 :foundation model(abstract);分类 cs.AI、cs.LG

Comments Under review at TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12282 2025-08-19 cs.CL cs.IR 57%

A Question Answering Dataset for Temporal-Sensitive Retrieval-Augmented Generation

Ziyang Chen, Erxue Min, Xiang Zhao, Yunxin Li, Xin Jia, Jinzhi Liao, Jichao Li, Shuaiqiang Wang, Baotian Hu, Dawei Yin

机构 * Laboratory for Big Data and Decision, National University of Defense Technology, Changsha, China(大数据与决策实验室,国防科技大学,长沙,中国) Baidu Inc., Beijing, China(百度公司,北京,中国) Department of Computer Science and Technology, Harbin Institute of Technology (Shenzhen), Shenzhen, China(计算机科学与技术系,哈尔滨工业大学(深圳),深圳,中国)

专题命中 评测与基准 :LLM(abstract);分类 cs.CL

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10872 2025-08-19 cs.CV cs.AI cs.ET 57%

V-RoAst: Visual Road Assessment. Can VLM be a Road Safety Assessor Using the iRAP Standard?

Natchapon Jongwiriyanurak, Zichao Zeng, June Moh Goo, Xinglei Wang, Ilya Ilyankou, Kerkritt Sriroongvikrai, Nicola Christie, Meihui Wang, Huanfa Chen, James Haworth

机构 * University College London(伦敦大学学院) Chulalongkorn University(朱拉隆梭大学)

专题命中 评测与基准 :language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11801 2025-08-19 cs.CV cs.CL 57%

VideoAVE: A Multi-Attribute Video-to-Text Attribute Value Extraction Dataset and Benchmark Models

Ming Cheng, Tong Wu, Jiazhen Hu, Jiaying Gong, Hoda Eldardiry

机构 * Virginia Tech(弗吉尼亚理工大学)

专题命中 评测与基准 :language model(abstract);分类 cs.CL

Comments 5 pages, 2 figures, 5 tables, accepted in CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04721 2025-08-19 cs.CL eess.AS 57%

Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities

Guan-Ting Lin, Jiachen Lian, Tingle Li, Qirui Wang, Gopala Anumanchipalli, Alexander H. Liu, Hung-yi Lee

机构 * Graduate Institute of Communication Engineering, National Taiwan University(国立台湾大学通信工程研究所) UC Berkeley(加州大学伯克利分校) University of Washington(华盛顿大学) MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)

专题命中 评测与基准 :language model(abstract);分类 cs.CL

Comments Accepted by ASRU 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12668 2025-08-19 cs.CV 50%

WP-CLIP: Leveraging CLIP to Predict Wölfflin's Principles in Visual Art

Abhijay Ghildyal, Li-Yun Wang, Feng Liu

机构 * Portland State University(波特兰州立大学)

专题命中 评测与基准 :language model(abstract)

Comments ICCV 2025 AI4VA workshop (oral), Code: https://github.com/abhijay9/wpclip

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12438 2025-08-19 cs.GR cs.CV 50%

Express4D: Expressive, Friendly, and Extensible 4D Facial Motion Generation Benchmark

Yaron Aloni, Rotem Shalev-Arkushin, Yonatan Shafir, Guy Tevet, Ohad Fried, Amit Haim Bermano

机构 * Tel Aviv University(特拉维夫大学) Reichman University(里奇曼大学)

专题命中 评测与基准 :LLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08038 2025-08-19 cs.CV 50%

TRIDE: A Text-assisted Radar-Image weather-aware fusion network for Depth Estimation

Huawei Sun, Zixu Wang, Hao Feng, Julius Ott, Lorenzo Servadei, Robert Wille

专题命中 评测与基准 :language model(abstract)

Comments Accepted by TMLR (2025.08)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04813 2025-08-19 cs.GR cs.CV 50%

WIR3D: Visually-Informed and Geometry-Aware 3D Shape Abstraction

Richard Liu, Daniel Fu, Noah Tan, Itai Lang, Rana Hanocka

机构 * University of Chicago(芝加哥大学)

专题命中 评测与基准 :foundation model(abstract)

Comments ICCV 2025 Oral Project page: https://threedle.github.io/wir3d/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16867 2025-08-19 cs.CV 50%

ETVA: Evaluation of Text-to-Video Alignment via Fine-grained Question Generation and Answering

Kaisi Guan, Zhengfeng Lai, Yuchong Sun, Peng Zhang, Wei Liu, Kieran Liu, Meng Cao, Ruihua Song

机构 * Renmin University of China(中国人民大学) Apple(苹果公司)

专题命中 评测与基准 :LLM(abstract)

Comments International Conference on Computer Vision 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 效率与部署 42 篇

2503.17811 2025-08-19 cs.CL cs.AI cs.DB 90%

Feather-SQL: A Lightweight NL2SQL Framework with Dual-Model Collaboration Paradigm for Small Language Models

Wenqi Pei, Hailing Xu, Hengyuan Zhao, Shizheng Hou, Han Chen, Zining Zhang, Pingyi Luo, Bingsheng He

专题命中 效率与部署 :language model(title,abstract);small language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI

Comments DL4C @ ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12920 2025-08-19 cs.AI cs.MA 89%

Do Large Language Model Agents Exhibit a Survival Instinct? An Empirical Study in a Sugarscape-Style Simulation

Atsushi Masumori, Takashi Ikegami

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11886 2025-08-19 cs.CV cs.AI cs.CL cs.LG eess.IV 89%

EVTP-IVS: Effective Visual Token Pruning For Unifying Instruction Visual Segmentation In Multi-Modal Large Language Models

Wenhui Zhu, Xiwen Chen, Zhipeng Wang, Shao Tang, Sayan Ghosh, Xuanzhao Dong, Rajat Koner, Yalin Wang

机构 * Arizona State University(亚利桑那州立大学) Clemson University(克莱姆森大学) LinkedIn Corporation(领英公司) Ludwig Maximilian University of Munich(慕尼黑路德维希-马克西米利安大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00819 2025-08-19 cs.CL 88%

Beyond Fixed: Training-Free Variable-Length Denoising for Diffusion Large Language Models

Jinsong Li, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Jiaqi Wang, Dahua Lin

机构 * The Chinese University of Hong Kong(香港中文大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

Comments Code is available at https://github.com/Li-Jinsong/DAEDAL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05934 2025-08-19 cs.CR cs.AI 88%

Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models

Ma Teng, Jia Xiaojun, Duan Ranjie, Li Xinfeng, Huang Yihao, Jia Xiaoshuang, Chu Zhixuan, Ren Wenqi

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13114 2025-08-19 cs.SE 85%

Understanding LLM-Centric Challenges for Deep Learning Frameworks: An Empirical Analysis

Yanzhou Mu, Rong Wang, Juan Zhai, Chunrong Fang, Xiang Chen, Jiacong Wu, An Guo, Jiawei Shen, Bingzhuo Li, Zhenyu Chen

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract)

Comments 46 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03865 2025-08-19 cs.CL cs.AI cs.LG 85%

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference

Seungjun Shin, Jaehoon Oh, Dokwan Oh

机构 * Samsung Advanced Institute of Technology, Korea(三星先进技术研究所,韩国)

专题命中 效率与部署 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ICML 2025 (final version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12590 2025-08-19 cs.LG cs.AI 84%

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding

Jihoon Park, Seungeun Oh, Seong-Lyun Kim

机构 * School of Electrical and Electronic Engineering, Yonsei University(延世大学电子与电气工程学院)

专题命中 效率与部署 :LLM(title,abstract);language model(abstract);分类 cs.AI、cs.LG

Comments 6 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12212 2025-08-19 cs.LG cs.AI q-bio.QM 84%

ProtTeX-CC: Activating In-Context Learning in Protein LLM via Two-Stage Instruction Compression

Chuanliu Fan, Zicheng Ma, Jun Gao, Nan Yu, Jun Zhang, Ziqiang Cao, Yi Qin Gao, Guohong Fu

专题命中 效率与部署 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13037 2025-08-19 cs.CL cs.AI 82%

Can Large Models Teach Student Models to Solve Mathematical Problems Like Human Beings? A Reasoning Distillation Method via Multi-LoRA Interaction

Xinhe Li, Jiajun Liu, Peng Wang

机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education(新一代人工智能技术及其交叉应用重点实验室(东南大学),教育部)

专题命中 效率与部署 :LLM(abstract);large language model(abstract);language model(abstract);small language model(abstract)

Comments Accepted by IJCAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏