arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-10-29 至 2025-10-29 共收录 186 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 47 篇

2510.24126 2025-10-29 cs.CL 79%

Reinforcement Learning for Long-Horizon Multi-Turn Search Agents

Vivek Kalyan, Martin Andrews

机构 * Red Cat Labs, Singapore(红猫实验室(新加坡)) Singapore(新加坡)

专题命中 评测与基准 :large language model(abstract,comments);language model(abstract,comments);LLM(abstract);分类 cs.CL

Comments 4 pages plus references and appendices. Accepted into the First Workshop on Multi-Turn Interactions in Large Language Models at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.17401 2025-10-29 cs.CL 79%

Evaluation of Geographical Distortions in Language Models

Rémy Decoupes, Roberto Interdonato, Mathieu Roche, Maguelonne Teisseire, Sarah Valentin

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL

Comments Accepted version. Published in Machine Learning (Springer) 114:263 (2025). Open access under a CC BY-NC-ND 4.0 license. DOI: 10.1007/s10994-025-06916-9

Journal ref Machine Learning (2025) 114:263

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24488 2025-10-29 cs.CL cs.AI 79%

A word association network methodology for evaluating implicit biases in LLMs compared to humans

Katherine Abramski, Giulio Rossetti, Massimo Stella

机构 * University of Pisa, Department of Computer Science(比萨大学计算机科学系) National Research Council of Italy, Institute of Information Science and Technologies(意大利国家研究委员会信息科学与技术研究所) University of Trento, Department of Psychology and Cognitive Science(特伦托大学心理学与认知科学系)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments 24 pages, 13 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24150 2025-10-29 cs.CL cs.AI 79%

Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean

Chanwoo Park, Suyoung Park, JiA Kang, Jongyeon Park, Sangho Kim, Hyunji M. Park, Sumin Bae, Mingyu Kang, Jaejin Lee

专题命中 评测与基准 :large language model(abstract);language model(abstract);prompting(abstract);分类 cs.CL、cs.AI

Comments submitted to ACL ARR Rolling Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23854 2025-10-29 cs.CL cs.AI 79%

Can LLMs Narrate Tabular Data? An Evaluation Framework for Natural Language Representations of Text-to-SQL System Outputs

Jyotika Singh, Weiyi Sun, Amit Agarwal, Viji Krishnamurthy, Yassine Benajiba, Sujith Ravi, Dan Roth

机构 * Oracle AI

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16947 2025-10-29 cs.LG cs.AI 79%

MixAT: Combining Continuous and Discrete Adversarial Training for LLMs

Csaba Dékány, Stefan Balauca, Robin Staab, Dimitar I. Dimitrov, Martin Vechev

机构 * INSAIT, Sofia University "St. Kliment Ohridski"(INSAIT,索菲亚大学"圣克莱孟·奥赫里迪斯基") ETH Zurich(苏黎世联邦理工学院) ELTE Eötvös Loránd University, Budapest, Hungary(ELTE 爱斯特霍特·洛兰德大学,布达佩斯,匈牙利)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

Comments Published at 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20510 2025-10-29 cs.CV 78%

CPathAgent: An Agent-based Foundation Model for Interpretable High-Resolution Pathology Image Analysis Mimicking Pathologists' Diagnostic Logic

Yuxuan Sun, Yixuan Si, Chenglu Zhu, Kai Zhang, Zhongyi Shui, Bowen Ding, Tao Lin, Lin Yang

机构 * College of Computer Science and Technology, Zhejiang University, China(浙江大学计算机科学与技术学院) Research Center for Industries of the Future and School of Engineering, Westlake University, China(未来产业研究中心和西湖大学工程学院) Department of Computer Science and Engineering, The Ohio State University, USA(俄亥俄州立大学计算机科学与工程系)

专题命中 评测与基准 :foundation model(title,abstract)

Comments 52 pages, 34 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24528 2025-10-29 cs.AI 77%

From Cross-Task Examples to In-Task Prompts: A Graph-Based Pseudo-Labeling Framework for In-context Learning

Zihan Chen, Song Wang, Xingbo Fu, Chengshuai Shi, Zhenyu Lei, Cong Shen, Jundong Li

机构 * Department of ECE, University of Virginia(电子工程系,弗吉尼亚大学)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21724 2025-10-29 cs.CV cs.AI cs.HC 77%

OmniResponse: Online Multimodal Conversational Response Generation in Dyadic Interactions

Cheng Luo, Jianghui Wang, Bing Li, Siyang Song, Bernard Ghanem

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

Comments 25 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24250 2025-10-29 cs.CL 77%

Evaluating LLMs on Generating Age-Appropriate Child-Like Conversations

Syed Zohaib Hassan, Pål Halvorsen, Miriam S. Johnson, Pierre Lison

机构 * Department of Holistic Systems, SimulaMet(整体系统系,SimulaMet) Department of Computer Science, Oslo Metropolitan University(计算机科学系,奥斯陆 Metropolitan 大学) Department of Behavioural Sciences, Oslo Metropolitan University(行为科学系,奥斯陆 Metropolitan 大学) Department of Psychology, Harvard University(心理学系,哈佛大学) Department of Informatics, University of Oslo(信息学系,奥斯陆大学) Norwegian Computing Center(挪威计算中心)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

Comments 11 pages excluding references and appendix. 3 figures and 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24073 2025-10-29 cs.CL 77%

Challenging Multilingual LLMs: A New Taxonomy and Benchmark for Unraveling Hallucination in Translation

Xinwei Wu, Heng Liu, Jiang Zhou, Xiaohu Zhao, Linlong Xu, Longyue Wang, Weihua Luo, Kaifu Zhang

机构 * Tianjin University(天津大学) Alibaba International Digital Commerce(阿里巴巴国际数字商业)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01271 2025-10-29 cs.LG cs.AI 74%

PULSE: Practical Evaluation Scenarios for Large Multimodal Model Unlearning

Tatsuki Kawakami, Kazuki Egashira, Atsuyuki Miyai, Go Irie, Kiyoharu Aizawa

机构 * The University of Tokyo(东京大学) Tokyo University of Science(东京科学大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG;LLM(comments)

Comments Accepted at NeurIPS 2025 Workshop: Evaluating the Evolving LLM Lifecycle

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24345 2025-10-29 cs.CL cs.AI 73%

LongWeave: A Long-Form Generation Benchmark Bridging Real-World Relevance and Verifiability

Zikai Xiao, Fei Huang, Jianhong Tu, Jianhui Wei, Wen Ma, Yuxuan Zhou, Jian Wu, Bowen Yu, Zuozhu Liu, Junyang Lin

机构 * Zhejiang University(浙江大学) Alibaba Group(阿里巴巴集团)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments EMNLP Findings 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16724 2025-10-29 cs.AI cs.CL 73%

A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications

Minhua Lin, Zongyu Wu, Zhichao Xu, Hui Liu, Xianfeng Tang, Qi He, Charu Aggarwal, Hui Liu, Xiang Zhang, Suhang Wang

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) The University of Utah(犹他大学) Amazon(亚马逊) Microsoft(微软) IBM T.J. Watson Research Center(IBM T.J.沃森研究中心) Michigan State University(密歇根州立大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments 38 pages, 4 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03490 2025-10-29 cs.CL cs.AI 73%

SEER: The Span-based Emotion Evidence Retrieval Benchmark

Aneesha Sampath, Oya Aran, Emily Mower Provost

机构 * Computer Science & Engineering University of Michigan(计算机科学与工程大学麦克马斯特大学) Data Science & AI Procter & Gamble(数据科学与人工智能普罗cter与甘贝尔)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04897 2025-10-29 cs.CV 71%

From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes

Tianxu Wang, Zhuofan Zhang, Ziyu Zhu, Yue Fan, Jing Xiong, Pengxiang Li, Xiaojian Ma, Qing Li

机构 * State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI) Tsinghua University(清华大学) Peking University(北京大学) Beijing Institute of Technology(北京理工大学)

专题命中 评测与基准 :large language model(abstract,comments);language model(abstract,comments)

Comments Update v3 of the NeurIPS 2025 Datasets and Benchmarks paper (v2), including additional evaluations of state-of-the-art multimodal large language models. Project page: https://anywhere-3d.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11842 2025-10-29 cs.CV cs.CL 70%

Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs

Xuannan Liu, Zekun Li, Zheqi He, Peipei Li, Shuhan Xia, Xing Cui, Huaibo Huang, Xi Yang, Ran He

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) University of California, Santa Barbara(加州大学圣芭芭拉分校) Center for Research on Intelligent Perception and Computing, NLPR, CASIA(中国科学院CASIA智能感知与计算中心)

专题命中 评测与基准 :LLM(abstract);language model(abstract);分类 cs.CL

Comments Accepted by NeurIPS 2025 Dataset and Benchmark Track, Project page: https://liuxuannan.github.io/Video-SafetyBench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24321 2025-10-29 cs.CV cs.AI 70%

Few-Shot Remote Sensing Image Scene Classification with CLIP and Prompt Learning

Ivica Dimitrovski, Vlatko Spasev, Ivan Kitanovski

机构 * Faculty of Computer Science and Engineering(计算机科学与工程学院) University Ss Cyril and Methodius(西里尔与美多西乌斯大学)

专题命中 评测与基准 :language model(abstract);prompting(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23995 2025-10-29 cs.CL 70%

M-Eval: A Heterogeneity-Based Framework for Multi-evidence Validation in Medical RAG Systems

Mengzhou Sun, Sendong Zhao, Jianyu Chen, Haochun Wang, Bin Qin

机构 * Faculty of Computing(计算学院) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23673 2025-10-29 cs.CR cs.AI 70%

MCPGuard : Automatically Detecting Vulnerabilities in MCP Servers

Bin Wang, Zexin Liu, Hao Yu, Ao Yang, Yenan Huang, Jing Guo, Huangsheng Cheng, Hui Li, Huiyu Wu

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23975 2025-10-29 q-bio.QM physics.bio-ph stat.ME stat.ML 67%

Machine learning approaches for interpretable antibody property prediction using structural data

Kevin Michalewicz, Mauricio Barahona, Barbara Bravi

专题命中 评测与基准 :large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23949 2025-10-29 cs.CL cs.AI 62%

Uncovering the Potential Risks in Unlearning: Danger of English-only Unlearning in Multilingual LLMs

Kyomin Hwang, Hyeonjin Kim, Seungyeon Kim, Sunghyun Wee, Nojun Kwak

专题命中 评测与基准 :LLM(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09060 2025-10-29 cs.LG cs.AI q-bio.GN 62%

Multimodal 3D Genome Pre-training

Minghao Yang, Pengteng Li, Yan Liang, Qianyi Cai, Zhihang Zheng, Shichen Zhang, Pengfei Zhang, Zhi-An Huang, Hui Xiong

机构 * Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou), China(人工智能研究所,香港科技大学(广州)) School of Artificial Intelligence, South China Normal University, China(人工智能学院,华南师范大学) Thrust of Bioscience and Biomedical Engineering, The Hong Kong University of Science and Technology (Guangzhou), China(生物科学与生物医学工程研究所,香港科技大学(广州)) Department of Computer Science, City University of Hong Kong (Dongguan), China(计算机科学系,香港城市大学(东莞)) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology Hong Kong SAR, China(计算机科学与工程系,香港科技大学香港特别行政区)

专题命中 评测与基准 :foundation model(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24402 2025-10-29 cs.IR cs.AI cs.CE 57%

Metadata-Driven Retrieval-Augmented Generation for Financial Question Answering

Michail Dadopoulos, Anestis Ladas, Stratos Moschidis, Ioannis Negkakis

机构 * Aristotle University of Thessaloniki(塞萨洛尼基亚里士多德大学) University of Macedonia(马其顿大学) University of Plymouth(普利茅斯大学)

专题命中 评测与基准 :LLM(abstract);分类 cs.AI

Comments Preprint version submitted to the International Journal of Accounting Information Systems; currently under major revision. 20 pages, 1 figure, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14788 2025-10-29 cs.IR cs.AI 57%

Cross-Scenario Unified Modeling of User Interests at Billion Scale

Manjie Xu, Cheng Chen, Xin Jia, Jingyi Zhou, Yongji Wu, Zejian Wang, Chi Zhang, Kai Zuo, Yibo Chen, Xu Tang, Yao Hu, Yixin Zhu

机构 * Peking University(北京大学) Fudan University(复旦大学) Xiaohongshu Inc.(小红书公司)

专题命中 评测与基准 :LLM(abstract);分类 cs.AI

Comments https://github.com/ariesssxu/RedSeqRec

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24194 2025-10-29 cs.RO cs.LG 57%

Blindfolded Experts Generalize Better: Insights from Robotic Manipulation and Videogames

Ev Zisselman, Mirco Mutti, Shelly Francis-Meretzki, Elisei Shafer, Aviv Tamar

机构 * Technion – Israel Institute of Technology(技术学院-以色列理工学院)

专题命中 评测与基准 :foundation model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24317 2025-10-29 cs.CR 50%

Cybersecurity AI Benchmark (CAIBench): A Meta-Benchmark for Evaluating Cybersecurity AI Agents

María Sanz-Gómez, Víctor Mayoral-Vilches, Francesco Balassone, Luis Javier Navarrete-Lozano, Cristóbal R. J. Veas Chavez, Maite del Mundo de Torres

专题命中 评测与基准 :LLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24124 2025-10-29 physics.geo-ph 50%

Global Chlorophyll-\textit{a} Retrieval algorithm from Sentinel 2 Using Residual Deep Learning and Novel Machine Learning Water Classification

Yotam Sherf, Bar Efrati, Gabriel Rozman, Moshe Harel

专题命中 评测与基准 :foundation model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23609 2025-10-29 math.MG cs.IT math.IT 50%

A short survey on almost orthogonal vectors in a few specific large dimensions

Rami Luisto

专题命中 评测与基准 :language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12841 2025-10-29 cs.CV 50%

AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning

Yiming Ren, Zhiqiang Lin, Yu Li, Gao Meng, Weiyun Wang, Junjie Wang, Zicheng Lin, Jifeng Dai, Yujiu Yang, Wenhai Wang, Ruihang Chu

机构 * Tsinghua University(清华大学) Shanghai AI Laboratory(上海人工智能实验室) Fudan University(复旦大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 评测与基准 :foundation model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏