arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-08-18 至 2025-08-18 共收录 30 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 30 篇

2506.06371 2025-08-18 cs.CL 89%

Relationship Detection on Tabular Data Using Statistical Analysis and Large Language Models

Panagiotis Koletsis, Christos Panagiotopoulos, Georgios Th. Papadopoulos, Vasilis Efthymiou

机构 * Harokopio University(哈罗科波斯大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);prompting(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08688 2025-08-18 cs.CL cs.AI 89%

The Fellowship of the LLMs: Multi-Model Workflows for Synthetic Preference Optimization Dataset Generation

Samee Arif, Sualeha Farid, Abdul Hameed Azeemi, Awais Athar, Agha Ali Raza

机构 * Lahore University of Management Sciences(拉合尔管理科学大学) University of Michigan - Ann Arbor(密歇根大学安娜堡分校) EMBL European Bioinformatics Institute(欧洲生物信息学研究所) Strategize Inc(Strategize公司)

专题命中 评测与基准 :preference optimization(title,abstract);LLM(abstract);large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10922 2025-08-18 cs.CV 88%

A Survey on Video Temporal Grounding with Multimodal Large Language Model

Jianlong Wu, Wei Liu, Ye Liu, Meng Liu, Liqiang Nie, Zhouchen Lin, Chang Wen Chen

机构 * School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) School of Computer Science and Technology, Shandong Jianzhu University(山东建筑大学计算机科学与技术学院) School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院) Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract)

Comments 20 pages,6 figures,survey

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.18013 2025-08-18 cs.CL cs.AI 86%

A Survey on Recent Advances in LLM-Based Multi-turn Dialogue Systems

Zihao Yi, Jiarui Ouyang, Zhe Xu, Yuwen Liu, Tianhao Liao, Haohao Luo, Ying Shen

机构 * Sun Yat-Sen University(中山大学)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments 35 pages, 10 figures, ACM Computing Surveys

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11061 2025-08-18 cs.CL 85%

BIPOLAR: Polarization-based granular framework for LLM bias evaluation

Martin Pavlíček, Tomáš Filip, Petr Sosík

机构 * Institute for Research and Applications of Fuzzy Modeling, University of Ostrava(模糊建模研究与应用研究所,奥斯特拉瓦大学) Institute of Computer Science, Faculty of Philosophy and Science, Silesian University in Opava(计算机科学研究所,哲学与科学学院,奥帕瓦西里西亚大学)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11736 2025-08-18 cs.CL 85%

Personalized LLM for Generating Customized Responses to the Same Query from Different Users

Hang Zeng, Chaoyue Niu, Fan Wu, Chengfei Lv, Guihai Chen

机构 * Shanghai Jiao Tong University(上海交通大学) Alibaba Group(阿里巴巴集团)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

Comments Accepted by CIKM'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08536 2025-08-18 cs.CV 82%

SenCLIP: Enhancing zero-shot land-use mapping for Sentinel-2 with ground-level prompting

Pallavi Jain, Dino Ienco, Roberto Interdonato, Tristan Berchoux, Diego Marcos

机构 * Mediterranean Agronomic Institute of Montpellier - CIHEAM-IAMM(地中海农业研究院-CIHEAM-IAMM) Inria(法国国家信息与自动化技术研究院) INRAE(法国国家农业研究咨询中心) Cirad(国际热带农业研究中心) UMR TETIS(TETIS联合研究单位) Univ. of Montpellier(蒙彼利埃大学)

专题命中 评测与基准 :prompting(title,abstract);language model(abstract)

Comments Accepted at WACV'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10927 2025-08-18 cs.CL cs.AI cs.CE cs.LG 80%

Modeling and Detecting Company Risks from News: A Case Study in Bloomberg News

Jiaxin Pei, Soumya Vadlamannati, Liang-Kang Huang, Daniel Preotiuc-Pietro, Xinyu Hua

机构 * University of Michigan(密歇根大学) Bloomberg(彭博)

专题命中 评测与基准 :large language model(abstract);language model(abstract);prompting(abstract);分类 cs.CL、cs.AI、cs.LG

Journal ref NAACL 2024: Human Language Technologies (Volume 6:Industry Track), pages 63 : 72

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11536 2025-08-18 cs.CL 79%

Language models align with brain regions that represent concepts across modalities

Maria Ryskina, Greta Tuckute, Alexander Fung, Ashley Malkin, Evelina Fedorenko

机构 * Vector Institute for AI(向量人工智能研究所) MIT(麻省理工学院)

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL

Comments Accepted to COLM 2025. Code and data can be found at https://github.com/ryskina/concepts-brain-llms

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11341 2025-08-18 cs.CV cs.CR cs.LG 79%

Semantically Guided Adversarial Testing of Vision Models Using Language Models

Katarzyna Filus, Jorge M. Cruz-Duarte

机构 * Institute of Theoretical and Applied Informatics, Polish Academy of Sciences(波兰科学院理论与应用信息学研究所) University of Lille, CNRS, Inria, Centrale Lille, UMR 9189 CRIStAL(里尔大学)

专题命中 评测与基准 :language model(title,abstract);分类 cs.LG

Comments 12 pages, 4 figures, 3 tables. Submitted for peer review

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11262 2025-08-18 cs.CV cs.AI 79%

Vision-Language Models display a strong gender bias

Aiswarya Konavoor, Raj Abhijit Dandekar, Rajat Dandekar, Sreedath Panat

机构 * Togo AI Labs(Togo人工智能实验室) Vizuara AI Labs(Vizuara人工智能实验室)

专题命中 评测与基准 :language model(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10972 2025-08-18 cs.CV cs.AI cs.HC 79%

Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision

Rosiana Natalie, Wenqian Xu, Ruei-Che Chang, Rada Mihalcea, Anhong Guo

专题命中 评测与基准 :language model(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10541 2025-08-18 cs.LG q-bio.QM 79%

Driving Accurate Allergen Prediction with Protein Language Models and Generalization-Focused Evaluation

Brian Shing-Hei Wong, Joshua Mincheol Kim, Sin-Hang Fung, Qing Xiong, Kelvin Fu-Kiu Ao, Junkang Wei, Ran Wang, Dan Michelle Wang, Jingying Zhou, Bo Feng, Alfred Sze-Lok Cheng, Kevin Y. Yip, Stephen Kwok-Wing Tsui, Qin Cao

专题命中 评测与基准 :language model(title,abstract);分类 cs.LG

Comments 59 pages, 5 main figures, 15 supplementary figures, 2 supplementary tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09346 2025-08-18 cs.CV cs.AI 79%

B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions

Hao Zhang, Wenqi Shao, Hong Liu, Yongqiang Ma, Ping Luo, Yu Qiao, Nanning Zheng, Kaipeng Zhang

机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence(人机混合增强智能国家重点实验室) National Engineering Research Center for Visual Information and Applications(视觉信息与应用国家工程研究中心) Institute of Artificial Intelligence and Robotics(人工智能与机器人研究院) Xi’an Jiaotong University(西安交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Osaka University(大阪大学)

专题命中 评测与基准 :language model(title,abstract);分类 cs.AI

Comments Accepted by IEEE Transactions on Information Forensics & Security

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11383 2025-08-18 cs.CL cs.AI 79%

When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs

Mikhail Seleznyov, Mikhail Chaichuk, Gleb Ershov, Alexander Panchenko, Elena Tutubalina, Oleg Somov

机构 * AIRI Skoltech(斯克里普切尔技术大学) Yandex MIPT(莫斯科理工学院) HSE University(俄罗斯高等经济大学) Sber AI ISP RAS Research Center for Trusted AI(俄罗斯科学院信息与系统问题研究所可信AI研究中心)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11310 2025-08-18 cs.CL cs.AI cs.IR 79%

SGSimEval: A Comprehensive Multifaceted and Similarity-Enhanced Benchmark for Automatic Survey Generation Systems

Beichen Guo, Zhiyuan Wen, Yu Yang, Peng Gao, Ruosong Yang, Jiaxing Shen

机构 * The Hong Kong Polytechnic University(香港理工大学) The Education University of Hong Kong(香港教育大学) China Mobile (Hong Kong) Innovation and Research Institute(中国移动(香港)创新与研究院) Lingnan University(岭南大学)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments Accepted to The 21st International Conference on Advanced Data Mining and Applications (ADMA2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11317 2025-08-18 cs.CV cs.MM 78%

Logic Unseen: Revealing the Logical Blindspots of Vision-Language Models

Yuchen Zhou, Jiayu Tang, Shuo Yang, Xiaoyan Xiao, Yuqin Dai, Wenhao Yang, Chao Gou, Xiaobo Xia, Tat-Seng Chua

专题命中 评测与基准 :language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10906 2025-08-18 cs.CL 77%

PersonaTwin: A Multi-Tier Prompt Conditioning Framework for Generating and Evaluating Personalized Digital Twins

Sihan Chen, John P. Lalor, Yi Yang, Ahmed Abbasi

机构 * Viterbi School of Engineering, University of Southern California(南加州大学维特比工程学院) Human-centered Analytics Lab, University of Notre Dame(诺特大学人本分析实验室) Department of IT, Analytics, and Operations, University of Notre Dame(诺特大学信息技术、分析与运营部门) Department of Information Systems, Business Statistics and Operations Management, HKUST(香港科技大学信息系统、商业统计与运营管理部门)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

Comments Presented at the Generation, Evaluation & Metrics (GEM) Workshop at ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05147 2025-08-18 cs.CR cs.LG 77%

Pr$εε$mpt: Sanitizing Sensitive Prompts for LLMs

Amrita Roy Chowdhury, David Glukhov, Divyam Anshumaan, Prasad Chalasani, Nicolas Papernot, Somesh Jha, Mihir Bellare

机构 * University of Michigan, Ann Arbor(密歇根大学安阿伯分校) University of Toronto(多伦多大学) Vector Institute(向量研究所) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Langroid Incorporated(Langroid公司) University of California, San Diego(加州大学圣地亚哥分校)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11414 2025-08-18 cs.CL 70%

Survey-to-Behavior: Downstream Alignment of Human Values in LLMs via Survey Questions

Shangrui Nie, Florian Mai, David Kaczér, Charles Welch, Zhixue Zhao, Lucie Flek

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

Comments 7 pages 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11163 2025-08-18 cs.CL 70%

MobQA: A Benchmark Dataset for Semantic Understanding of Human Mobility Data through Question Answering

Hikaru Asano, Hiroki Ouchi, Akira Kasuga, Ryo Yonetani

机构 * The University of Tokyo \ AIP Japan Tokyo Nara Institute of Science The University of Tokyo \ AIP

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

Comments 23 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10955 2025-08-18 cs.CV cs.CL cs.MM 70%

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey

Wenbin An, Jiahao Nie, Yaqiang Wu, Feng Tian, Shijian Lu, Qinghua Zheng

机构 * School of Automation Science and Engineering, Xi’an Jiaotong University(自动化科学与工程学院,西安交通大学) Interdisciplinary Graduate Programme, Nanyang Technological University(跨学科研究生项目,南洋理工大学) Lenovo Research, Lenovo(联想研究院,联想) School of Computer Science and Technology, Xi’an Jiaotong University(计算机科学与技术学院,西安交通大学) College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

Comments 21 pages, 361 references

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16636 2025-08-18 cs.CL cs.CV 70%

Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries

Yin Wu, Quanyu Long, Jing Li, Jianfei Yu, Wenya Wang

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

Comments 21 pages, 6 figures, 17 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11412 2025-08-18 cs.HC 67%

Towards Embodied Conversational Agents for Reducing Oral Exam Anxiety in Extended Reality

Jens Grubert, Yvonne Sedelmaier, Dieter Landes

专题命中 评测与基准 :large language model(abstract);language model(abstract)

Comments Accepted to the IEEE ISMAR-Adjunct Proceedings 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11192 2025-08-18 cs.CV 67%

Generating Dialogues from Egocentric Instructional Videos for Task Assistance: Dataset, Method and Benchmark

Lavisha Aggarwal, Vikas Bahirwani, Lin Li, Andrea Colaco

机构 * Google(谷歌)

专题命中 评测与基准 :large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11247 2025-08-18 cs.CL cs.AI 62%

Cross-Granularity Hypergraph Retrieval-Augmented Generation for Multi-hop Question Answering

Changjian Wang, Weihong Deng, Weili Guan, Quan Lu, Ning Jiang

专题命中 评测与基准 :LLM(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00184 2025-08-18 cs.LG cs.AI 62%

Text-to-Level Diffusion Models With Various Text Encoders for Super Mario Bros

Jacob Schrum, Olivia Kilday, Emilio Salas, Bess Hagan, Reid Williams

专题命中 评测与基准 :language model(abstract);分类 cs.AI、cs.LG

Comments Accepted to appear in The 21st AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment (November 10-14, 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09057 2025-08-18 cs.CL 57%

MVISU-Bench: Benchmarking Mobile Agents for Real-World Tasks by Multi-App, Vague, Interactive, Single-App and Unethical Instructions

Zeyu Huang, Juyuan Wang, Longfeng Chen, Boyi Xiao, Leng Cai, Yawen Zeng, Jin Xu

机构 * South China University of Technology(华南理工大学) Pazhou Lab(Pazhou 实验室)

专题命中 评测与基准 :language model(abstract);分类 cs.CL

Comments ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18932 2025-08-18 cs.MM cs.CL 57%

MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks

Lei Zhang, Xin Zhou, Chaoyue He, Di Wang, Yi Wu, Hong Xu, Wei Liu, Chunyan Miao

机构 * Nanyang Technological University(南洋理工大学) University College London(伦敦大学学院) Alibaba Group(阿里巴巴集团)

专题命中 评测与基准 :LLM(abstract);分类 cs.CL

Comments Accepted at ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10030 2025-08-18 cs.RO cs.AI 57%

EmbodiedAgent: A Scalable Hierarchical Approach to Overcome Practical Challenge in Multi-Robot Control

Hanwen Wan, Yifei Chen, Yixuan Deng, Zeyu Wei, Dongrui Li, Zexin Lin, Donghao Wu, Jiu Cheng, Xiaoqiang Ji

机构 * School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳)科学与工程学院) Shenzhen Institute of Artificial Intelligence and Robotics for Society, China(社会人工智能与机器人研究所) The School of Computer Science, The University of Sydney, Australia(悉尼大学计算机科学学院) Faculty of Computer and Mathematical Sciences, The Hong Kong Polytechnic University, China(香港理工大学计算机与数学科学学院)

专题命中 评测与基准 :LLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏