arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-11-14 至 2025-11-14 共收录 171 信号源:cs.CL, cs.AI, cs.LG

1. 推理与问题求解 32 篇

2511.10648 2025-11-14 cs.CV 67%

Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling

Jiahao Wang, Weiye Xu, Aijun Yang, Wengang Zhou, Lewei Lu, Houqiang Li, Xiaohua Wang, Jinguo Zhu

机构 * Xi’an Jiaotong University(西安交通大学) University of Science and Technology of China(中国科学技术大学) SenseTime Research(商汤科技研究院)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract)

Comments Accepted to NeurIPS 2025 (The Thirty-Ninth Annual Conference on Neural Information Processing Systems)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16781 2025-11-14 cs.CV cs.AI cs.CL 62%

Xiaoice: Training-Free Video Understanding via Self-Supervised Spatio-Temporal Clustering of Semantic Features

Shihao Ji, Zihui Song

专题命中 推理与问题求解 :language model(abstract);分类 cs.CL、cs.AI

Comments This paper is being withdrawn because we have identified a significant error in the implementation of our self-supervised clustering approach. Specifically, our feature aggregation step inadvertently leaked temporal information across frames, which violates the core assumption of our training-free method. We sincerely apologize to the research community

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10611 2025-11-14 cs.NI cs.AI 57%

Towards an Agentic Workflow for Internet Measurement Research

Alagappan Ramanathan, Eunju Kang, Dongsu Han, Sangeetha Abdu Jyothi

机构 * University of California, Irvine(加州大学伊维奇分校)

专题命中 推理与问题求解 :LLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05508 2025-11-14 q-fin.GN cs.AI cs.CE 57%

Personalized Chain-of-Thought Summarization of Financial News for Investor Decision Support

Tianyi Zhang, Mu Chen

机构 * Computer Science Department University of Southern California Los Angeles, CA, USA(计算机科学系 美国南加州大学 洛杉矶 加州 美国)

专题命中 推理与问题求解 :language model(abstract);分类 cs.AI

Comments Proceedings of ICDM Workshops

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21538 2025-11-14 cs.CV cs.AI 57%

Caption This, Reason That: VLMs Caught in the Middle

Zihan Weng, Lucas Gomez, Taylor Whittington Webb, Pouya Bashivan

机构 * Integrated Program in Neuroscience (IPN) McGill University(神经科学联合计划 麦吉尔大学) Mila, University of Montreal(蒙特利尔大学Mila) Microsoft Research USA(微软研究院美国总部) Department of Physiology McGill University(生理学系 麦吉尔大学)

专题命中 推理与问题求解 :language model(abstract);分类 cs.AI

Comments Paper accepted by nips 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 评测与基准 44 篇

2510.01611 2025-11-14 cs.AI cs.CL 90%

PsychCounsel-Bench: Evaluating the Psychology Intelligence of Large Language Models

Min Zeng

机构 * Hong Kong University of Science and Technology(香港科技大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10418 2025-11-14 cs.DB 89%

CityVerse: A Unified Data Platform for Multi-Task Urban Computing with Large Language Models

Yaqiao Zhu, Hongkai Wen, Mark Birkin, Man Luo

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11034 2025-11-14 cs.SE 89%

The Impact of Large Language Models (LLMs) on Code Review Process

Antonio Collante, Samuel Abedu, SayedHassan Khatoonabadi, Ahmad Abdellatif, Ebube Alor, Emad Shihab

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18573 2025-11-14 cs.CL cs.AI 88%

FactReasoner: A Probabilistic Approach to Long-Form Factuality Assessment for Large Language Models

Radu Marinescu, Debarun Bhattacharjya, Junkyu Lee, Tigran Tchrakian, Javier Carnerero Cano, Yufang Hou, Elizabeth Daly, Alessandra Pascale

机构 * IBM Research(IBM研究院) IT:U - Interdisciplinary Transformation University Austria(interdisciplinary Transformation University Austria)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05294 2025-11-14 cs.RO cs.AI cs.LG 86%

Towards Embodied Agentic AI: Review and Classification of LLM- and VLM-Driven Robot Autonomy and Interaction

Sahar Salimpour, Lei Fu, Kajetan Rachwał, Pascal Bertrand, Kevin O'Sullivan, Robert Jakob, Farhad Keramat, Leonardo Militano, Giovanni Toffetti, Harry Edelman, Jorge Peña Queralta

机构 * Department of Computing, University of Turku(图尔库大学计算机系) Institute of Computer Science, Zurich University of Applied Sciences(应用科学大学计算机科学研究所) Centre for Artificial Ingelligence, Zurich University of Applied Sciences(应用科学大学人工智能中心) Agentic Systems Lab, Department of Management, Technology and Economics, ETH Zürich(苏黎世联邦理工学院管理、科技与经济系代理系统实验室) Faculty of Mathematics and Information Science, Warsaw University of Technology(华沙技术大学数学与信息科学学院)

专题命中 评测与基准 :LLM(title);large language model(abstract);language model(abstract);foundation model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08215 2025-11-14 cs.CV cs.LG 85%

Evaluating Gemini LLM in Food Image-Based Recipe and Nutrition Description with EfficientNet-B4 Visual Backbone

Rizal Khoirul Anam

机构 * Department of Computer Science and Technology(计算机科学与技术系) Nanjing University of Information Science and Technology(南京信息工程大学)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07127 2025-11-14 cs.LG 85%

REACT-LLM: A Benchmark for Evaluating LLM Integration with Causal Features in Clinical Prognostic Tasks

Linna Wang, Zhixuan You, Qihui Zhang, Jiunan Wen, Ji Shi, Yimin Chen, Yusen Wang, Fanqi Ding, Ziliang Feng, Li Lu

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09062 2025-11-14 cs.GT 85%

Pricing Online LLM Services with Data-Calibrated Stackelberg Routing Game

Zhendong Guo, Wenchao Bai, Jiahui Jin

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract)

Comments Extended version

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09964 2025-11-14 cs.SE cs.AI cs.PL 83%

EnvTrace: Simulation-Based Semantic Evaluation of LLM Code via Execution Trace Alignment -- Demonstrated at Synchrotron Beamlines

Noah van der Vleuten, Anthony Flores, Shray Mathur, Max Rakitin, Thomas Hopkins, Kevin G. Yager, Esther H. R. Tsai

专题命中 评测与基准 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02635 2025-11-14 cs.CL 83%

Test Set Quality in Multilingual LLM Evaluation

Chalamalasetti Kranti, Gabriel Bernier-Colborne, Yvan Gauthier, Sowmya Vajjala

专题命中 评测与基准 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL

Comments to appear in the proceedings of Eval4NLP workshop at AACL 2025. Camera ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19319 2025-11-14 cs.CL cs.AI 81%

FHIR-AgentBench: Benchmarking LLM Agents for Realistic Interoperable EHR Question Answering

Gyubok Lee, Elea Bach, Eric Yang, Tom Pollard, Alistair Johnson, Edward Choi, Yugang jia, Jong Ha Lee

机构 * Korea Advanced Institute of Science & Technology(韩国科学技术院) Verily Life Sciences(Verily 生物科技) Massachusetts Institute of Technology(麻省理工学院)

专题命中 评测与基准 :LLM(title,abstract);分类 cs.CL、cs.AI

Comments ML4H 2025 Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17337 2025-11-14 cs.CL cs.AI cs.IR 81%

Captions Speak Louder than Images: Generalizing Foundation Models for E-commerce from High-quality Multimodal Instruction Data

Xinyi Ling, Hanwen Du, Bo Peng, Zhihui Zhu, Xia Ning

机构 * Department of Computer Science and Engineering, The Ohio State University(计算机科学与工程系,俄亥俄州立大学) Translational Data Analytics Institute, The Ohio State University(转化数据分析研究所,俄亥俄州立大学) Department of Biomedical Informatics, The Ohio State University(生物医学信息学系,俄亥俄州立大学)

专题命中 评测与基准 :foundation model(title,abstract);分类 cs.CL、cs.AI

Comments IJCNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10615 2025-11-14 cs.CV cs.CL 79%

Towards Blind and Low-Vision Accessibility of Lightweight VLMs and Custom LLM-Evals

Shruti Singh Baghel, Yash Pratap Singh Rathore, Sushovan Jena, Anurag Pradhan, Amit Shukla, Arnav Bhavsar, Pawan Goyal

机构 * Indian Institute of Technology Mandi(印度理工学院曼迪分校) Vellore Institute of Technology(韦洛雷理工学院) Indian Institute of Technology Kharagpur(印度理工学院哈里科普分校)

专题命中 评测与基准 :LLM(title);language model(abstract);分类 cs.CL

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10177 2025-11-14 cs.CV cs.AI 79%

Utilizing a Geospatial Foundation Model for Coastline Delineation in Small Sandy Islands

Tishya Chhabra, Manisha Bajpai, Walter Zesk, Skylar Tibbits

专题命中 评测与基准 :foundation model(title,abstract);分类 cs.AI

Comments 8 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09969 2025-11-14 cs.CY cs.AI cs.HC 79%

Owlgorithm: Supporting Self-Regulated Learning in Competitive Programming through LLM-Driven Reflection

Juliana Nieto-Cardenas, Erin Joy Kramer, Peter Kurto, Ethan Dickey, Andres Bejarano

机构 * Purdue University(普渡大学)

专题命中 评测与基准 :LLM(title,abstract);分类 cs.AI

Comments 7 pages, 1 figure, to be published in SIGCSE '26

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10519 2025-11-14 cs.CL cs.AI 79%

Say It Differently: Linguistic Styles as Jailbreak Vectors

Srikant Panda, Avinash Rai

机构 * Independent Researcher(独立研究者)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10286 2025-11-14 q-bio.TO 78%

Diversity Over Scale: Whole-Slide Image Variety Enables H&E Foundation Model Training with Fewer Patches

Christoph Bosch, John K. L. Wong, Martin Paulikat, Myroslav Zapukhlyak, Bharti Arora, Manasi Aichmüller-Ratnaparkhe, Jens Baumann, Shivani Karn, Rutuja Kamble, Swapnil Karnik, Bhushan Khedkar, Serey Vathana Chhut, Witali Aswolinskiy, Christian Aichmüller

专题命中 评测与基准 :foundation model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09571 2025-11-14 q-bio.QM cs.AI 77%

General Intelligence-based Fragmentation (GIF): A framework for peak-labeled spectra simulation

Margaret R. Martin, Soha Hassoun

机构 * Department of Computer Science, Tufts University, Medford, MA 02155, USA(计算机科学系,塔夫茨大学,马萨诸塞州梅德福,02155,美国)

专题命中 评测与基准 :large language model(abstract);language model(abstract);prompting(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15554 2025-11-14 cs.CR cs.LG cs.SE 77%

Rethinking the Evaluation of Secure Code Generation

Shih-Chieh Dai, Jun Xu, Guanhong Tao

机构 * University of Utah(犹他大学)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.LG

Comments Accepted by ICSE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09724 2025-11-14 cs.CV cs.AI cs.RO 74%

PALMS+: Modular Image-Based Floor Plan Localization Leveraging Depth Foundation Model

Yunqian Cheng, Benjamin Princen, Roberto Manduchi

机构 * University of California, Santa Cruz(加州大学圣克ruz分校)

专题命中 评测与基准 :foundation model(title);分类 cs.AI

Comments Accepted to IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026, Application Track. Main paper: 8 pages, 5 figures. Supplementary material included

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10627 2025-11-14 cs.AI cs.CV cs.FL cs.LG 73%

Querying Labeled Time Series Data with Scenario Programs

Edward Kim, Devan Shanker, Varun Bharadwaj, Hongbeen Park, Jinkyu Kim, Hazem Torfah, Daniel J Fremont, Sanjit A Seshia

机构 * University of California, Berkeley(加州大学伯克利分校) Korea University(韩国大学) Chalmers University of Technology(查尔姆斯理工大学) University of Gothenburg(哥德堡大学) University of California, Santa Cruz(加州大学圣克ruz分校)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

Journal ref NASA Formal Methods Conference 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13595 2025-11-14 cs.CL cs.AI cs.IR 73%

MMTEB: Massive Multilingual Text Embedding Benchmark

Kenneth Enevoldsen, Isaac Chung, Imene Kerboua, Márton Kardos, Ashwin Mathur, David Stap, Jay Gala, Wissam Siblini, Dominik Krzemiński, Genta Indra Winata, Saba Sturua, Saiteja Utpala, Mathieu Ciancone, Marion Schaeffer, Gabriel Sequeira, Diganta Misra, Shreeya Dhakal, Jonathan Rystrøm, Roman Solomatin, Ömer Çağatan, Akash Kundu, Martin Bernstorff, Shitao Xiao, Akshita Sukhlecha, Bhavish Pahwa, Rafał Poświata, Kranthi Kiran GV, Shawon Ashraf, Daniel Auras, Björn Plüster, Jan Philipp Harries, Loïc Magne, Isabelle Mohr, Mariya Hendriksen, Dawei Zhu, Hippolyte Gisserot-Boukhlef, Tom Aarsen, Jan Kostkan, Konrad Wojtasik, Taemin Lee, Marek Šuppa, Crystina Zhang, Roberta Rocca, Mohammed Hamdy, Andrianos Michail, John Yang, Manuel Faysse, Aleksei Vatolin, Nandan Thakur, Manan Dey, Dipam Vasani, Pranjal Chitale, Simone Tedeschi, Nguyen Tai, Artem Snegirev, Michael Günther, Mengzhou Xia, Weijia Shi, Xing Han Lù, Jordan Clive, Gayatri Krishnakumar, Anna Maksimova, Silvan Wehrli, Maria Tikhonova, Henil Panchal, Aleksandr Abramov, Malte Ostendorff, Zheng Liu, Simon Clematide, Lester James Miranda, Alena Fenogenova, Guangyu Song, Ruqiya Bin Safi, Wen-Ding Li, Alessia Borghini, Federico Cassano, Hongjin Su, Jimmy Lin, Howard Yen, Lasse Hansen, Sara Hooker, Chenghao Xiao, Vaibhav Adlakha, Orion Weller, Siva Reddy, Niklas Muennighoff

机构 * Aarhus University(奥胡斯大学) Individual Contributor(个人贡献者) Esker(Esker公司) INSA Lyon(里昂INSA) University of Amsterdam(阿姆斯特丹大学) MBZUAI(穆罕默德·本·拉希德智能技术研究院) Jina AI(Jina AI公司) Microsoft Research(微软研究院) Wikit(Wikit公司)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments Accepted for ICLR: https://openreview.net/forum?id=zl3pfz4VCV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08045 2025-11-14 cs.CL cs.AI cs.CY 73%

Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs

Mohsinul Kabir, Ajwad Abrar, Sophia Ananiadou

机构 * Department of Computer Science, National Center for Text Mining, The University of Manchester(计算机科学系,文本挖掘国家中心,曼彻斯特大学) Department of Computer Science and Engineering, Islamic University of Technology(计算机科学与工程系,伊斯兰技术大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments Accepted at EMNLP 2025 (Main)

Journal ref Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15721 2025-11-14 cs.CL cs.AI 73%

EcomMMMU: Strategic Utilization of Visuals for Robust Multimodal E-commerce Models

Xinyi Ling, Hanwen Du, Zhihui Zhu, Xia Ning

机构 * Department of Computer Science and Engineering, The Ohio State University(俄亥俄州立大学计算机科学与工程系) Translational Data Analytics Institute, The Ohio State University(俄亥俄州立大学转化数据分析研究所) Department of Biomedical Informatics, The Ohio State University(俄亥俄州立大学生物医学信息学系)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments ICJNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10049 2025-11-14 cs.SE 71%

Continuous Benchmark Generation for Evaluating Enterprise-scale LLM Agents

Divyanshu Saxena, Rishikesh Maurya, Xiaoxuan Ou, Gagan Somashekar, Shachee Mishra Gupta, Arun Iyer, Yu Kang, Chetan Bansal, Aditya Akella, Saravan Rajmohan

专题命中 评测与基准 :LLM(title)

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏