arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-11-04 至 2025-11-04 共收录 368 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 92 篇

2503.08221 2025-11-04 cs.CV cs.AI cs.MM 70%

EgoBlind: Towards Egocentric Visual Assistance for the Blind

Junbin Xiao, Nanxin Huang, Hao Qiu, Zhulin Tao, Xun Yang, Richang Hong, Meng Wang, Angela Yao

机构 * National University of Singapore(新加坡国立大学) Communication University of China(中国传媒大学) University of Science and Technology of China(中国科学技术大学) Hefei University of Technoloy(合肥工业大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

Comments NeurIPS'25 (D&B Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01053 2025-11-04 cs.CL 70%

Building a Silver-Standard Dataset from NICE Guidelines for Clinical LLMs

Qing Ding, Eric Hua Qing Zhang, Felix Jozsa, Julia Ive

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

Comments Submitted to EFMI Medical Informatics Europe 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03529 2025-11-04 cs.CL 70%

Mafoko: Structuring and Building Open Multilingual Terminologies for South African NLP

Vukosi Marivate, Isheanesu Dzingirai, Fiskani Banda, Richard Lastrucci, Thapelo Sindane, Keabetswe Madumo, Kayode Olaleye, Abiodun Modupe, Unarine Netshifhefhe, Herkulaas Combrink, Mohlatlego Nakeng, Matome Ledwaba

机构 * Data Science for Social Impact, Dept. of Computer Science, University of Pretoria(南非普里尼亚大学计算机科学系数据科学与社会影响研究中心) AfriDSAI, University of Pretoria(南非普里尼亚大学非洲数据科学与人工智能研究所) Lelapa AI(莱拉帕人工智能) Economics and Management Sciences, University of the Free State(南非自由州大学经济与管理科学系) Interdisciplinary Centre for Digital Futures, University of the Free State(南非自由州大学跨学科数字未来研究中心)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

Comments Accepted for Sixth Workshop on Resources for African Indigenous Languages (RAIL) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16728 2025-11-04 cs.CL 70%

Natural Language Generation

Emiel van Miltenburg, Chenghua Lin

机构 * Tilburg University(蒂尔堡大学) University of Manchester(曼彻斯特大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

Comments 4 pages + references. Submitted for publication in the Encyclopedia of Language & Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06577 2025-11-04 cs.AI 70%

Survey Transfer Learning: Recycling Data with Silicon Responses

Ali Amini

机构 * Department of Government(政府系)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

Comments Revised and expanded version (v2). 32 pages, 11 figures. Under review at Political Analysis. Presented at SPSA 2025 (Political Methodology Panel)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01393 2025-11-04 cs.CR 67%

ConneX: Automatically Resolving Transaction Opacity of Cross-Chain Bridges for Security Analysis

Hanzhong Liang, Yue Duan, Xing Su, Xiao Li, Yating Liu, Yulong Tian, Fengyuan Xu, Sheng Zhong

专题命中 评测与基准 :large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00776 2025-11-04 cs.SE 67%

A Systematic Literature Review of Code Hallucinations in LLMs: Characterization, Mitigation Methods, Challenges, and Future Directions for Reliable AI

Cuiyun Gao, Guodong Fan, Chun Yong Chong, Shizhan Chen, Chao Liu, David Lo, Zibin Zheng, Qing Liao

专题命中 评测与基准 :large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12704 2025-11-04 cs.CV cs.MM 67%

SmartFreeEdit: Mask-Free Spatial-Aware Image Editing with Complex Instruction Understanding

Qianqian Sun, Jixiang Luo, Dell Zhang, Xuelong Li

机构 * Institute of Artificial Intelligence(TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信) The University of Hongkong(香港大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02866 2025-11-04 cs.CV 67%

OpenFACADES: An Open Framework for Architectural Caption and Attribute Data Enrichment via Street View Imagery

Xiucheng Liang, Jinheng Xie, Tianhong Zhao, Rudi Stouffs, Filip Biljecki

机构 * Department of Architecture, National University of Singapore(建筑系,新加坡国立大学) Department of Electrical and Computer Engineering, National University of Singapore(电气与计算机工程系,新加坡国立大学) School of Artificial Intelligence, Shenzhen Technology University(人工智能学院,深圳科技大学) Department of Real Estate, National University of Singapore(房地产系,新加坡国立大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract)

Journal ref ISPRS Journal of Photogrammetry and Remote Sensing 230: 918-942, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19006 2025-11-04 econ.TH 67%

Performance Rating Equilibrium

Mehmet Mars Seven

专题命中 评测与基准 :large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01359 2025-11-04 cs.CL cs.AI 62%

PrefixNLI: Detecting Factual Inconsistencies as Soon as They Arise

Sapir Harary, Eran Hirsch, Aviv Slobodkin, David Wan, Mohit Bansal, Ido Dagan

机构 * Bar-Ilan University(巴伊兰大学) UNC Chapel Hill(北卡罗来纳大学教堂山分校) cs.unc.edu(北卡罗来纳大学教堂山分校) cs.biu.ac.il(巴伊兰大学计算机科学系)

专题命中 评测与基准 :LLM(abstract);分类 cs.CL、cs.AI

Comments 9 pages + appendix. Code, datasets, and models are available at https://github.com/sapirharary/PrefixNLI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16791 2025-11-04 cs.LG cs.AI 62%

TabArena: A Living Benchmark for Machine Learning on Tabular Data

Nick Erickson, Lennart Purucker, Andrej Tschalzev, David Holzmüller, Prateek Mutalik Desai, David Salinas, Frank Hutter

机构 * Amazon Web Services(亚马逊网络服务) University of Freiburg(弗赖堡大学) University of Mannheim(曼海姆大学) INRIA Paris(巴黎国家信息与自动化研究所) Ecole Normale Supérieure(高等师范学校) PSL Research University(巴黎科学实验室大学) Prior Labs(先驱实验室) ELLIS Institute Tübingen(图宾根ELLIS研究所)

专题命中 评测与基准 :foundation model(abstract);分类 cs.AI、cs.LG

Comments Accepted (spotlight) at NeurIPS 2025 Datasets and Benchmarks Track. v4: fixed links in comments. v3: NeurIPS camera-ready version. v2: fixed author list. 51 pages. Code available at https://tabarena.ai/code and examples at https://tabarena.ai/code-examples and dataset curation at https://tabarena.ai/data-tabular-ml-iid-study and https://tabarena.ai/dataset-curation

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00416 2025-11-04 cs.CL cs.AI 62%

PADBen: A Comprehensive Benchmark for Evaluating AI Text Detectors Against Paraphrase Attacks

Yiwei Zha, Rui Min, Shanu Sushmita

机构 * Khoury College of Computer Science(科里学院计算机科学学院)

专题命中 评测与基准 :LLM(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00126 2025-11-04 cs.LG cs.AI 62%

Dynamic Model Selection for Trajectory Prediction via Pairwise Ranking and Meta-Features

Lu Bowen

机构 * Lu Bowen(路 Bowen)

专题命中 评测与基准 :LLM(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01775 2025-11-04 cs.CV cs.AI cs.MM 57%

How Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert Assessment

Zhen Chen, Qing Xu, Jinlin Wu, Biao Yang, Yuhao Zhai, Geng Guo, Jing Zhang, Yinlu Ding, Nassir Navab, Jiebo Luo

机构 * Yale University(耶鲁大学) University of Nottingham(诺丁汉大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Department of Neurosurgery, The First Hospital, Shanxi Medical University(山西医科大学第一医院神经外科部) Department of Gastrointestinal Surgery, The Second Qilu Hospital, Shandong University(山东大学第二齐鲁医院胃肠外科部) Technical University of Munich(慕尼黑技术大学) University of Rochester(罗切斯特大学)

专题命中 评测与基准 :foundation model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01380 2025-11-04 cs.CL 57%

Confounding Factors in Relating Model Performance to Morphology

Wessel Poelman, Thomas Bauwens, Miryam de Lhoneux

机构 * NLP, Department of Computer Science, KU Leuven(自然语言处理,计算机科学系,鲁文大学)

专题命中 评测与基准 :language model(abstract);分类 cs.CL

Comments EMNLP 2025: Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23112 2025-11-04 cs.CE cs.AI 57%

GroupSHAP-Guided Integration of Financial News Keywords and Technical Indicators for Stock Price Prediction

Minjoo Kim, Jinwoong Kim, Sangjin Park

专题命中 评测与基准 :language model(abstract);分类 cs.AI

Comments 6 pages

Journal ref ICAIF 2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00519 2025-11-04 cs.CL 57%

Exploring and Mitigating Gender Bias in Encoder-Based Transformer Models

Ariyan Hossain, Khondokar Mohammad Ahanaf Hannan, Rakinul Haque, Nowreen Tarannum Rafa, Humayra Musarrat, Shoaib Ahmed Dipu, Farig Yousuf Sadeque

机构 * Department of Computer Science and Engineering BRAC University(计算机科学与工程系 BRAC大学)

专题命中 评测与基准 :language model(abstract);分类 cs.CL

Comments 25 pages, 20 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22765 2025-11-04 cs.AI 57%

Jarvis: Towards Personalized AI Assistant via Personal KV-Cache Retrieval

Binxiao Xu, Junyu Feng, Shaolin Lu, Yulin Luo, Shilin Yan, Hao Liang, Ming Lu, Wentao Zhang

机构 * Peking University(北京大学) Xi’an Jiaotong University(西安交通大学) Alibaba Group(阿里巴巴集团) Intel Labs China(英特尔中国实验室)

专题命中 评测与基准 :language model(abstract);分类 cs.AI

Comments 19 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01206 2025-11-04 cs.CV cs.AI 57%

EndoGMDE: Generalizable Monocular Depth Estimation with Mixture of Low-Rank Experts for Diverse Endoscopic Scenes

Liangjing Shao, Chenkang Du, Benshuang Chen, Xueli Liu, Xinrong Chen

机构 * Department of Electronic Engineering, The Chinese University of Hong Kong(香港中文大学电子工程系) College of Biomedical Engineering, Fudan University(复旦大学生物医学工程学院) Shanghai Key laboratory of Medical Image Computing and Computer Assisted Intervention(上海医学影像计算与计算机辅助手术重点实验室) ENT Institute and Department of Otolaryngology, Eye & ENT Hospital of Fudan University(复旦大学眼耳鼻喉科医院耳鼻喉科及耳鼻喉科研究所)

专题命中 评测与基准 :foundation model(abstract);分类 cs.AI

Comments 12 pages, 12 figures, 7 tables. Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20900 2025-11-04 cs.SD cs.AI cs.MM 57%

Music Arena: Live Evaluation for Text-to-Music

Yonghyun Kim, Wayne Chi, Anastasios N. Angelopoulos, Wei-Lin Chiang, Koichi Saito, Shinji Watanabe, Yuki Mitsufuji, Chris Donahue

机构 * Carnegie Mellon University(卡内基梅隆大学) LMArena Sony AI(索尼人工智能) Georgia Tech(佐治亚理工学院)

专题命中 评测与基准 :LLM(abstract);分类 cs.AI

Comments NeurIPS 2025 Creative AI Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18161 2025-11-04 eess.AS cs.CL cs.SD 57%

Recent Trends in Distant Conversational Speech Recognition: A Review of CHiME-7 and 8 DASR Challenges

Samuele Cornell, Christoph Boeddeker, Taejin Park, He Huang, Desh Raj, Matthew Wiesner, Yoshiki Masuyama, Xuankai Chang, Zhong-Qiu Wang, Stefano Squartini, Paola Garcia, Shinji Watanabe

机构 * Carnegie Mellon University(卡内基梅隆大学) Paderborn University(帕德博恩大学) NVIDIA(NVIDIA公司) Meta Johns Hopkins University(约翰霍普金斯大学) Mitsubishi Electric Research Laboratories(三菱电机研究实验室) Southern University of Science and Technology(南方科技大学)

专题命中 评测与基准 :language model(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10449 2025-11-04 cs.CV cs.AI 57%

CoralVQA: A Large-Scale Visual Question Answering Dataset for Coral Reef Image Understanding

Hongyong Han, Wei Wang, Gaowei Zhang, Mingjie Li, Yi Wang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Technology Innovation Center for South China Sea Remote Sensing, Surveying and Mapping Collaborative Application, Ministry of Natural Resources(自然资源部南海遥感测绘协同应用技术创新中心) South China Sea Development Research Institute, Ministry of Natural Resources(自然资源部南海发展研究 institute) Inspur Computer Technology Co., Ltd(Inspur 计算机技术有限公司) Shandong Key Laboratory of Advanced Computing(山东先进计算重点实验室)

专题命中 评测与基准 :language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13776 2025-11-04 cs.AI cs.CY cs.HC 57%

Recommendations and Reporting Checklist for Rigorous & Transparent Human Baselines in Model Evaluations

Kevin L. Wei, Patricia Paskov, Sunishchal Dev, Michael J. Byun, Anka Reuel, Xavier Roberts-Gaal, Rachel Calcott, Evie Coxon, Chinmay Deshpande

机构 * Harvard University, Cambridge, MA, USA Stanford University, Stanford, CA, USA Center for Democracy \& Technology, Washington, D.C., USA Max Planck School of Cognition, Leipzig, Germany not those of RAND or its research sponsors, clients, or grantors.

专题命中 评测与基准 :foundation model(abstract);分类 cs.AI

Comments A version of this paper has been accepted to ICML 2025 as a position paper (spotlight), with the title: "Position: Human Baselines in Model Evaluations Need Rigor and Transparency (With Recommendations & Reporting Checklist)."

Journal ref Proceedings of the 42nd International Conference on Machine Learning, PMLR 267:82265-82325, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18829 2025-11-04 cs.AI cs.HC cs.OS 57%

LiteCUA: Computer as MCP Server for Computer-Use Agent on AIOS

Kai Mei, Xi Zhu, Hang Gao, Shuhang Lin, Yongfeng Zhang

机构 * Rutgers University(罗格斯大学)

专题命中 评测与基准 :language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17267 2025-11-04 cs.CL 57%

GreekBarBench: A Challenging Benchmark for Free-Text Legal Reasoning and Citations

Odysseas S. Chlapanis, Dimitrios Galanis, Nikolaos Aletras, Ion Androutsopoulos

机构 * Department of Informatics, Athens University of Economics and Business(信息学院,雅典经济与商业大学) Archimedes, Athena Research Center(阿提卡研究中心-阿基米德) Athena Research Center(阿提卡研究中心) University of Sheffield(谢菲尔德大学)

专题命中 评测与基准 :LLM(abstract);分类 cs.CL

Comments 19 pages, 17 figures, accepted in EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00162 2025-11-04 cs.LG math.OC 57%

Stochastic Subspace Descent Accelerated via Bi-fidelity Line Search

Nuojin Cheng, Alireza Doostan, Stephen Becker

机构 * Department of Applied Mathematics(应用数学系) University of Colorado Boulder(科罗拉多大学博尔德分校) Smead Aerospace Engineering Sciences Department(斯梅德航空航天工程科学系)

专题命中 评测与基准 :language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14904 2025-11-04 cs.LG 57%

Exploring Kolmogorov-Arnold Networks for Interpretable Time Series Classification

Irina Barašin, Blaž Bertalanič, Mihael Mohorčič, Carolina Fortuna

机构 * Department of Communication Systems, Jo z ef Stefan Institute Jamova ulica 39, 1000 Ljubljana, Slovenia

专题命中 评测与基准 :prompting(abstract);分类 cs.LG

Journal ref International Journal of Intelligent Systems, vol. 2025, no. 1, article 9553189, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01678 2025-11-04 cs.CV 50%

UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback

Ropeway Liu, Hangjie Yuan, Bo Dong, Jiazheng Xing, Jinwang Wang, Rui Zhao, Yan Xing, Weihua Chen, Fan Wang

机构 * Zhejiang University(浙江大学) DAMO Academy Alibaba Group(阿里云达摩院) Hupan Lab(虎扑实验室) National University of Singapore(新加坡国立大学)

专题命中 评测与基准 :language model(abstract)

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01625 2025-11-04 cs.DB 50%

UniDataBench: Evaluating Data Analytics Agents Across Structured and Unstructured Data

Han Weng, Zhou Liu, Yuanfeng Song, Xiaoming Yin, Xing Chen, Wentao Zhang

专题命中 评测与基准 :LLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏