arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-08-29 至 2025-08-29 共收录 117 信号源:cs.CL, cs.AI, cs.LG

1. 推理与问题求解 20 篇

2508.20912 2025-08-29 cs.DB cs.AI 85%

Research Challenges in Relational Database Management Systems for LLM Queries

Kerem Akillioglu, Anurag Chakraborty, Sairaj Voruganti, M. Tamer Özsu

机构 * University of Waterloo(多伦多大学)

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

Comments This paper will appear in the 6th International Workshop on Applied AI for Database Systems and Applications, AIDB Workshop at VLDB 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20384 2025-08-29 cs.AI 85%

Uncertainty Under the Curve: A Sequence-Level Entropy Area Metric for Reasoning LLM

Yongfu Zhu, Lin Sun, Guangxiang Zhao, Weihong Lin, Xiangzheng Zhang

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

Comments Under review for AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20279 2025-08-29 cs.CV cs.AI cs.CL 85%

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding

Zhuoran Yu, Yong Jae Lee

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract);instruction tuning(abstract)

Comments Accepted by COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20444 2025-08-29 cs.CR 85%

Ransomware 3.0: Self-Composing and LLM-Orchestrated

Md Raz, Meet Udeshi, P. V. Sai Charan, Prashanth Krishnamurthy, Farshad Khorrami, Ramesh Karri

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11302 2025-08-29 cs.CL 83%

Are formal and functional linguistic mechanisms dissociated in language models?

Michael Hanna, Yonatan Belinkov, Sandro Pezzelle

机构 * Institute for Logic, Language and Computation University of Amsterdam(逻辑、语言与计算研究所 阿姆斯特丹大学) Technion – Israel Institute of Technology(技术ion-以色列理工学院)

专题命中 推理与问题求解 :language model(title,abstract);large language model(abstract);分类 cs.CL

Comments To appear in Computational Linguistics. Pre-MIT Press publication version. 40 pages, 14 figures, 3 tables. Code available at https://github.com/hannamw/formal-functional-dissociation

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20151 2025-08-29 cs.AI 83%

IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinement

Yuanzhe Shen, Zisu Huang, Zhengkang Guo, Yide Liu, Guanxu Chen, Ruicheng Yin, Xiaoqing Zheng, Xuanjing Huang

专题命中 推理与问题求解 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI

Comments 17 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15576 2025-08-29 cs.CV cs.LG 79%

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models

Xin Huang, Ruibin Li, Tong Jia, Wei Zheng, Ya Wang

机构 * School of Artificial Intelligence and Software Engineering, Nanyang Normal University, Henan, China(人工智能与软件工程学院,南阳师范学院,河南) Institute for Artificial Intelligence, Peking University, Beijing, China(人工智能研究院,北京大学,北京) Collaborative Innovation Center of Intelligent Explosion-proof Equipment, Henan, China(智能防爆设备协同创新中心,河南)

专题命中 推理与问题求解 :language model(title,abstract);分类 cs.LG

Comments Accepted at the International Joint Conference on Artificial Intelligence (IJCAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20234 2025-08-29 cs.MA cs.AI cs.CY 77%

Validating Generative Agent-Based Models for Logistics and Supply Chain Management Research

Vincent E. Castillo

机构 * Fisher College of Business, The Ohio State University(俄亥俄州立大学费舍尔商学院)

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

Comments A version of this work is also available on SSRN (https://ssrn.com/abstract=5407742 or http://dx.doi.org/10.2139/ssrn.5407742). This preprint is distributed under the CC BY-NC-SA 4.0 License

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19724 2025-08-29 cs.CL cs.AI 73%

NLKI: A lightweight Natural Language Knowledge Integration Framework for Improving Small VLMs in Commonsense VQA Tasks

Aritra Dutta, Swapnanil Mukherjee, Deepanway Ghosal, Somak Aditya

机构 * IIT Kharagpur(印度理工学院Kharagpur分校) Ashoka University(阿什oka大学)

专题命中 推理与问题求解 :LLM(abstract);language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17387 2025-08-29 cs.LG cs.AI 73%

Graph-R1: Incentivizing the Zero-Shot Graph Learning Capability in LLMs via Explicit Reasoning

Yicong Wu, Guangyue Lu, Yuan Zuo, Huarong Zhang, Junjie Wu

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20395 2025-08-29 cs.CL cs.AI 73%

Measuring Reasoning Utility in LLMs via Conditional Entropy Reduction

Xu Guo

机构 * KTH Royal Institute of Technology(皇家理工学院)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20131 2025-08-29 cs.AI cs.LG 73%

ArgRAG: Explainable Retrieval Augmented Generation using Quantitative Bipolar Argumentation

Yuqicheng Zhu, Nico Potyka, Daniel Hernández, Yuan He, Zifeng Ding, Bo Xiong, Dongzhuoran Zhou, Evgeny Kharlamov, Steffen Staab

机构 * University of Stuttgart(斯图加特大学) Bosch Center for AI(博世人工智能中心) Cardiff University(卡迪夫大学) University of Oxford(牛津大学) University of Cambridge(剑桥大学) Stanford University(斯坦福大学) University of Oslo(奥斯陆大学) University of Southampton(南安普顿大学)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14862 2025-08-29 cs.CL 70%

Bitune: Leveraging Bidirectional Attention to Improve Decoder-Only LLMs

Dawid J. Kopiczko, Tijmen Blankevoort, Yuki M. Asano

机构 * Fundamental AI Lab University of Technology Nuremberg(基础人工智能实验室 汉诺威技术大学)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20941 2025-08-29 physics.ed-ph 67%

AI Reasoning Models for Problem Solving in Physics

Amir Bralin, N. Sanjay Rebello

专题命中 推理与问题求解 :large language model(abstract);language model(abstract)

Comments 6 pages, 1 table; Physics Education Research Conference (PERC) 2025 Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20408 2025-08-29 cs.IR 67%

Fact or Facsimile? Evaluating the Factual Robustness of Modern Retrievers

Haoyu Wu, Qingcheng Zeng, Kaize Ding

专题命中 推理与问题求解 :large language model(abstract);language model(abstract)

Comments Proceedings of the 34th ACM International Conference on Information and Knowledge Management

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20783 2025-08-29 cs.CV cs.AI 57%

Evaluating Compositional Generalisation in VLMs and Diffusion Models

Beth Pearson, Bilal Boulbarss, Michael Wray, Martha Lewis

机构 * University of Bristol(布里斯托大学) University of Amsterdam(阿姆斯特丹大学)

专题命中 推理与问题求解 :language model(abstract);分类 cs.AI

Comments 11 pages including references, 6 figures. Accepted at IWCS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20722 2025-08-29 cs.CL 57%

rStar2-Agent: Agentic Reasoning Technical Report

Ning Shang, Yifei Liu, Yi Zhu, Li Lyna Zhang, Weijiang Xu, Xinyu Guan, Buze Zhang, Bingcheng Dong, Xudong Zhou, Bowen Zhang, Ying Xin, Ziming Miao, Scarlett Li, Fan Yang, Mao Yang

机构 * Microsoft Research(微软研究院)

专题命中 推理与问题求解 :SFT(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20370 2025-08-29 cs.SE cs.AI 57%

Adaptive Root Cause Localization for Microservice Systems with Multi-Agent Recursion-of-Thought

Lingzhe Zhang, Tong Jia, Kangjin Wang, Weijie Hong, Chiming Duan, Minghua He, Ying Li

机构 * Peking University(北京大学) Alibaba Group(阿里巴巴集团)

专题命中 推理与问题求解 :LLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 评测与基准 27 篇

2508.21061 2025-08-29 cs.HC cs.AI cs.LG 91%

OnGoal: Tracking and Visualizing Conversational Goals in Multi-Turn Dialogue with Large Language Models

Adam Coscia, Shunan Guo, Eunyee Koh, Alex Endert

机构 * Georgia Institute of Technology(佐治亚理工学院) Adobe Research(Adobe研究)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);prompting(abstract)

Comments Accepted to UIST 2025. 18 pages, 9 figures, 2 tables. For a demo video, see https://youtu.be/uobhmxo6EIE

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18638 2025-08-29 cs.HC cs.AI 90%

Prompt Engineering and the Effectiveness of Large Language Models in Enhancing Human Productivity

Rizal Khoirul Anam

机构 * Department of Computer Science(计算机科学系)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);prompting(abstract)

Comments 38 pages, 15 tables, 5 figures. Submitted as a research paper draft for arXiv. Based on survey data collected in 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00182 2025-08-29 cs.CL cs.AI 90%

Zero-Shot Classification of Crisis Tweets Using Instruction-Finetuned Large Language Models

Emma McDaniel, Samuel Scheele, Jeff Liu

机构 * Computer Science Department Georgia State University(计算机科学系乔治亚州立大学) Disaster Relief Systems MIT Lincoln Laboratory(灾难救援系统麻省理工学院林肯实验室)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20385 2025-08-29 cs.CL 89%

CAPE: Context-Aware Personality Evaluation Framework for Large Language Models

Jivnesh Sandhan, Fei Cheng, Tushar Sandhan, Yugo Murawaki

机构 * Kyoto University(京都大学) IIT Kanpur(印度理工学院坎浦尔)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL

Comments Accepted at EMNLP25 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20583 2025-08-29 cs.CL cs.AI 86%

A Graph Talks, But Who's Listening? Rethinking Evaluations for Graph-Language Models

Soham Petkar, Hari Aakash K, Anirudh Vempati, Akshit Sinha, Ponnurangam Kumarauguru, Chirag Agarwal

机构 * plaksha.edu.in(普拉克斯哈大学) research.iiit.ac.in(IIIT研究机构)

专题命中 评测与基准 :language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15165 2025-08-29 cs.CL cs.AI cs.SE 86%

SoAy: A Solution-based LLM API-using Methodology for Academic Information Seeking

Yuanchun Wang, Jifan Yu, Zijun Yao, Jing Zhang, Yuyang Xie, Shangqing Tu, Yiyang Fu, Youhe Feng, Jinkai Zhang, Jingyao Zhang, Bowen Huang, Yuanyao Li, Huihui Yuan, Lei Hou, Juanzi Li, Jie Tang

机构 * Renmin University of China(中国人民大学) Key Laboratory of Data Engineering and Knowledge Engineering, MOE(数据工程与知识工程重点实验室) Tsinghua University(清华大学) Institute of Education(教育学院) Department of Computer Science and Technology(计算机科学与技术系) Engineering Research Center of Database and Business Intelligence, MOE(数据库与商务智能工程研究中心) School of Information(信息学院)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments KDD 2025; 22 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20818 2025-08-29 cs.LG cs.MA 85%

cMALC-D: Contextual Multi-Agent LLM-Guided Curriculum Learning with Diversity-Based Context Blending

Anirudh Satheesh, Keenan Powell, Hua Wei

机构 * Department of Computer Science, University of Maryland, College Park, Maryland, USA(美国马里兰大学计算机科学系) School of Computing and Augmented Intelligence, Arizona State University, Tempe, Arizona, USA(亚利桑那州立大学计算与增强智能学院)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.LG

Comments A shorter version has been accepted to the 2025 Conference on Information and Knowledge Management

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20420 2025-08-29 cs.CL 85%

CAMB: A comprehensive industrial LLM benchmark on civil aviation maintenance

Feng Zhang, Chengjie Pang, Yuehan Zhang, Chenyu Luo

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20750 2025-08-29 cs.CL 83%

Specializing General-purpose LLM Embeddings for Implicit Hate Speech Detection across Datasets

Vassiliy Cheremetiev, Quang Long Ho Ngo, Chau Ying Kot, Alina Elena Baia, Andrea Cavallaro

机构 * EPFL(苏黎世联邦理工学院) Idiap Research Institute(Idiap研究机构)

专题命中 评测与基准 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL

Comments Paper accepted at the DHOW Workshop at ACM Multimedia 2025. Code available at https://github.com/idiap/implicit-hsd

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10583 2025-08-29 cs.CV cs.CL 83%

Relative Drawing Identification Complexity is Invariant to Modality in Vision-Language Models

Diogo Freitas, Brigt Håvardstun, Cèsar Ferri, Darío Garigliotti, Jan Arne Telle, José Hernández-Orallo

机构 * Interactive Technologies Institute and NOVA LINCS Faculty of Exact Sciences and Engineering University of Madeira Portugal(互动技术研究所和NOVA LINCS精确科学与工程学院马德拉大学) Department of Informatics University of Bergen Norway(信息学院卑尔根大学挪威) Valencian Research Institute for Artificial Intelligence Universitat Politècnica de València Spain(瓦伦西亚人工智能研究机构瓦伦西亚理工大学西班牙) Leverhulme Centre for the Future of Intelligence and Valencian Research Institute for Artificial Intelligence Spain(未来智能中心和瓦伦西亚人工智能研究机构西班牙)

专题命中 评测与基准 :language model(title,abstract);large language model(abstract);分类 cs.CL

Comments 54 pages (42 pages of appendix). Accepted for publication at the ECAI 2025 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20227 2025-08-29 cs.CV cs.AI cs.CL cs.LG 82%

A Novel Framework for Automated Explain Vision Model Using Vision-Language Models

Phu-Vinh Nguyen, Tan-Hanh Pham, Chris Ngo, Truong Son Hy

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20976 2025-08-29 cs.SD cs.AI eess.AS 79%

WoW-Bench: Evaluating Fine-Grained Acoustic Perception in Audio-Language Models via Marine Mammal Vocalizations

Jaeyeon Kim, Heeseung Yun, Sang Hoon Woo, Chao-Han Huck Yang, Gunhee Kim

机构 * Carnegie Mellon University(卡内基梅隆大学) Seoul National University(首尔国立大学) NVIDIA(NVIDIA公司)

专题命中 评测与基准 :language model(title,abstract);分类 cs.AI

Comments Preprint. Project page: https://jaeyeonkim99.github.io/wow_bench/

详情

展开后加载摘要…

URL PDF HTML 收藏