arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-11-13 至 2025-11-13 共收录 183 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 48 篇

2511.09185 2025-11-13 cs.CL 77%

Context is Enough: Empirical Validation of $\textit{Sequentiality}$ on Essays

Amal Sunny, Advay Gupta, Vishnu Sreekumar

机构 * IIIT-Hyderabad(IIIT-海得拉巴)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

Comments 9 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08866 2025-11-13 cs.CL 77%

BioVerge: A Comprehensive Benchmark and Study of Self-Evaluating Agents for Biomedical Hypothesis Generation

Fuyi Yang, Chenchen Ye, Mingyu Derek Ma, Yijia Xiao, Matthew Yang, Wei Wang

机构 * University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08217 2025-11-13 cs.AI 77%

MADD: Multi-Agent Drug Discovery Orchestra

Gleb V. Solovev, Alina B. Zhidkovskaya, Anastasia Orlova, Nina Gubina, Anastasia Vepreva, Rodion Golovinskii, Ilya Tonkii, Ivan Dubrovsky, Ivan Gurev, Dmitry Gilemkhanov, Denis Chistiakov, Timur A. Aliev, Ivan Poddiakov, Galina Zubkova, Ekaterina V. Skorb, Vladimir Vinogradov, Alexander Boukhanovsky, Nikolay Nikitin, Andrei Dmitrenko, Anna Kalyuzhnaya, Andrey Savchenko

机构 * ITMO University(ITMO大学) Sber AI Lab(Sber AI实验室) D ONE AG HSE University(高等经济大学)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

Comments EMNLP2025 accepted paper, Findings 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14389 2025-11-13 cs.CL cs.HC 77%

Leveraging Small LLMs for Argument Mining in Education: Argument Component Identification, Classification, and Assessment

Lucile Favero, Juan Antonio Pérez-Ortiz, Tanja Käser, Nuria Oliver

机构 * ELLIS Alicante, Spain(阿利坎特ELLIS研究所,西班牙) Universitat d’Alacant, Spain(阿利坎特大学,西班牙) École Polytechnique Fédérale de Lausanne, EPFL, Switzerland(洛桑联邦理工学院,瑞士)

专题命中 评测与基准 :large language model(abstract);language model(abstract);prompting(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21636 2025-11-13 cs.CR cs.AI 77%

The Feasibility of Topic-Based Watermarking on Academic Peer Reviews

Alexander Nemecek, Yuzhou Jiang, Erman Ayday

机构 * Case Western Reserve University(凯斯西储大学)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

Comments Accepted at AACL 25 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09337 2025-11-13 cs.HC cs.DB 75%

TempoQL: A Readable, Precise, and Portable Query System for Electronic Health Record Data

Ziyong Ma, Richard D. Boyce, Adam Perer, Venkatesh Sivaraman

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract)

Comments Accepted as a Proceedings paper at Machine Learning for Health (ML4H) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09443 2025-11-13 cs.CV cs.AI 74%

BronchOpt : Vision-Based Pose Optimization with Fine-Tuned Foundation Models for Accurate Bronchoscopy Navigation

Hongchao Shu, Roger D. Soberanis-Mukul, Jiru Xu, Hao Ding, Morgan Ringel, Mali Shen, Saif Iftekar Sayed, Hedyeh Rafii-Tari, Mathias Unberath

机构 * Johnson & Johnson MedTech(强生医疗科技)

专题命中 评测与基准 :foundation model(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09426 2025-11-13 cs.CL cs.LG 73%

BIG5-TPoT: Predicting BIG Five Personality Traits, Facets, and Items Through Targeted Preselection of Texts

Triet M. Le, Arjun Chandra, C. Anton Rytting, Valerie P. Karuzis, Vladimir Rife, William A. Simpson

机构 * The University of Maryland Applied Research Laboratory for Intelligence and Security (ARLIS)(马里兰大学应用研究实验室(情报与安全实验室)) The University of Maryland (UMD)(马里兰大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09179 2025-11-13 cs.CL cs.AI 73%

A Hybrid Search for Complex Table Question Answering in Securities Report

Daiki Shirafuji, Koji Tanaka, Tatsuhiko Saito

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments Accepted to IIAI AAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04962 2025-11-13 cs.CL cs.AI 73%

Too Good to be Bad: On the Failure of LLMs to Role-Play Villains

Zihao Yi, Qingxuan Jiang, Ruotian Ma, Xingyu Chen, Qu Yang, Mengru Wang, Fanghua Ye, Ying Shen, Zhaopeng Tu, Xiaolong Li, Linus

机构 * Tencent(腾讯)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00942 2025-11-13 cs.CL cs.AI eess.SP 73%

anyECG-chat: A Generalist ECG-MLLM for Flexible ECG Input and Multi-Task Understanding

Haitao Li, Ziyu Li, Yiheng Mao, Ziyi Liu, Zhoujian Sun, Zhengxing Huang

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17482 2025-11-13 cs.LG cs.AI cs.CV cs.HC 73%

What's Producible May Not Be Reachable: Measuring the Steerability of Generative Models

Keyon Vafa, Sarah Bentley, Jon Kleinberg, Sendhil Mullainathan

机构 * Harvard University(哈佛大学) MIT(麻省理工学院) Cornell University(康奈尔大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15526 2025-11-13 eess.IV cs.CV 71%

Multi-scale Cascaded Foundation Model for Whole-body Organs-at-risk Segmentation

Rui Hao, Dayu Tan, Qiankun Li, Chunhou Zheng, Weimin Zhong, Zhigang Zeng

机构 * School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(人工智能与自动化学院,华中科技大学) Institute of Artificial Intelligence, Huazhong University of Science and Technology(人工智能研究院,华中科技大学) Hubei Key Laboratory of Brain-Inspired Intelligent Systems, Huazhong University of Science and Technology(湖北省脑启发智能系统重点实验室,华中科技大学) Key Laboratory of Image Processing and Intelligent Control (Huazhong University of Science and Technology), Ministry of Education(图像处理与智能控制重点实验室(华中科技大学),教育部) Key Laboratory of Intelligent Computing and Signal Processing, Ministry of Education, Anhui University(智能计算与信号处理重点实验室(安徽大学),教育部) College of Computing and Data Science (CCDS), Nanyang Technological University(计算与数据科学学院(CCDS),南洋理工大学) East China University of Science and Technology(东华大学)

专题命中 评测与基准 :foundation model(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09448 2025-11-13 cs.MM cs.LG 70%

MCAD: Multimodal Context-Aware Audio Description Generation For Soccer

Lipisha Chaudhary, Trisha Mittal, Subhadra Gopalakrishnan, Ifeoma Nwogu, Jaclyn Pytlarz

机构 * University at Buffalo, SUNY(布法罗大学,SUNY) Dolby Laboratories Inc.(杜比实验室公司)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09407 2025-11-13 cs.CL 70%

CARE-Bench: A Benchmark of Diverse Client Simulations Guided by Expert Principles for Evaluating LLMs in Psychological Counseling

Bichen Wang, Yixin Sun, Junzhe Wang, Hao Yang, Xing Fu, Yanyan Zhao, Si Wei, Shijin Wang, Bing Qin

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15201 2025-11-13 cs.RO cs.AI 70%

Survey of Vision-Language-Action Models for Embodied Manipulation

Haoran Li, Yuhui Chen, Wenbo Cui, Weiheng Liu, Kai Liu, Mingcai Zhou, Zhengtao Zhang, Dongbin Zhao

专题命中 评测与基准 :foundation model(abstract);post-training(abstract);分类 cs.AI

Comments in Chinese language

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23990 2025-11-13 cs.AI 70%

Multi-RAG: A Multimodal Retrieval-Augmented Generation System for Adaptive Video Understanding

Mingyang Mao, Mariela M. Perez-Cabarcas, Utteja Kallakuri, Nicholas R. Waytowich, Xiaomin Lin, Tinoosh Mohsenin

机构 * Johns Hopkins Whiting School of Engineering(约翰霍普金斯大学惠廷工程学院) DEVCOM Army Research Laboratory(国防部陆军研究实验室)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13264 2025-11-13 cs.CL 70%

OpenGenAlign: A Preference Dataset and Benchmark for Trustworthy Reward Modeling in Open-Ended, Long-Context Generation

Hanning Zhang, Juntong Song, Juno Zhu, Yuanhao Wu, Tong Zhang, Cheng Niu

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) NewsBreak

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

Comments Preprint update

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08630 2025-11-13 cs.CY cs.AI cs.HC 70%

Hope, Aspirations, and the Impact of LLMs on Female Programming Learners in Afghanistan

Hamayoon Behmanush, Freshta Akhtari, Roghieh Nooripour, Ingmar Weber, Vikram Kamath Cannanure

机构 * Saarland Informatics Campus, Saarland University, Saarbrücken, Germany(萨尔兰州信息学校园,萨尔兰州大学,德国萨尔布吕肯) Computer Science Faculty, Parwan University, Charikar, Afghanistan(计算机科学学院,帕尔万大学,阿富汗查里卡) Department of Counseling, Qazvin Branch, Islamic Azad University, Qazvin, Iran(咨询部门,德黑兰分支,伊斯兰 Azad 大学,伊朗德黑兰)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09134 2025-11-13 cs.CR cs.SE 67%

One Signature, Multiple Payments: Demystifying and Detecting Signature Replay Vulnerabilities in Smart Contracts

Zexu Wang, Jiachi Chen, Zewei Lin, Wenqing Chen, Kaiwen Ning, Jianxing Yu, Yuming Feng, Yu Zhang, Weizhe Zhang, Zibin Zheng

专题命中 评测与基准 :large language model(abstract);language model(abstract)

Comments Accepted at ICSE2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07600 2025-11-13 cs.HC 67%

Look into your Heart -- Prototypes for a Speculative Design Exploration of Personal Heart Rate Visualization

Swaroop Panda

专题命中 评测与基准 :large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.12815 2025-11-13 cs.CR cs.AI cs.CL cs.LG 67%

Formalizing and Benchmarking Prompt Injection Attacks and Defenses

Yupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia, Neil Zhenqiang Gong

专题命中 评测与基准 :LLM(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Published in USENIX Security Symposium 2024; the model sizes for closed-source models are from blog posts. For slides, see https://people.duke.edu/~zg70/code/PromptInjection.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08927 2025-11-13 cs.AI cs.CL cs.CY 62%

The Double Contingency Problem: AI Recursion and the Limits of Interspecies Understanding

Graham L. Bishop

机构 * UC San Diego, Synthesis Program(UC圣地亚哥大学,合成计划)

专题命中 评测与基准 :foundation model(abstract);分类 cs.CL、cs.AI

Comments 5 pages, no figures, to be published in the NeurIPS 2025: AI for Non-Human Animal Communication Workshop Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09309 2025-11-13 cs.HC cs.AI 57%

TaskSense: Cognitive Chain Modeling and Difficulty Estimation for GUI Tasks

Yiwen Yin, Zhian Hu, Xiaoxi Xu, Chun Yu, Xintong Wu, Wenyu Fan, Yuanchun Shi

机构 * Tsinghua University(清华大学) University of Washington(华盛顿大学) Cornell University(康奈尔大学) University of Sydney(悉尼大学)

专题命中 评测与基准 :LLM(abstract);分类 cs.AI

Comments 22 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09042 2025-11-13 cs.LG 57%

GeoGNN: Quantifying and Mitigating Semantic Drift in Text-Attributed Graphs

Liangwei Yang, Jing Ma, Jianguo Zhang, Zhiwei Liu, Jielin Qiu, Shirley Kokane, Shiyu Wang, Haolin Chen, Rithesh Murthy, Ming Zhu, Huan Wang, Weiran Yao, Caiming Xiong, Shelby Heinecke

机构 * Salesforce AI Research(Salesforce AI研究院)

专题命中 评测与基准 :language model(abstract);分类 cs.LG

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08884 2025-11-13 cs.LG 57%

Spectral Predictability as a Fast Reliability Indicator for Time Series Forecasting Model Selection

Oliver Wang, Pengrui Quan, Kang Yang, Mani Srivastava

机构 * Electrical and Computer Engineering University of California, Los Angeles(电气与计算机工程大学加州大学洛杉矶分校)

专题命中 评测与基准 :foundation model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22354 2025-11-13 cs.CL 57%

LLMs Struggle to Reject False Presuppositions when Misinformation Stakes are High

Judith Sieker, Clara Lachenmaier, Sina Zarrieß

机构 * Computational Linguistics, Department of Linguistics, Bielefeld University, Germany(计算语言学系,语言学系,比勒菲尔德大学,德国)

专题命中 评测与基准 :LLM(abstract);分类 cs.CL

Comments 8 pages (including References). Published at CogSci 2025: https://escholarship.org/uc/item/4932r1hx

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11358 2025-11-13 cs.CR cs.AI 57%

DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks

Yupei Liu, Yuqi Jia, Jinyuan Jia, Dawn Song, Neil Zhenqiang Gong

专题命中 评测与基准 :LLM(abstract);分类 cs.AI

Comments Distinguished Paper Award in IEEE Symposium on Security and Privacy, 2025. For slides, see https://people.duke.edu/~zg70/code/PromptInjection.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09545 2025-11-13 cs.IR 50%

Practical RAG Evaluation: A Rarity-Aware Set-Based Metric and Cost-Latency-Quality Trade-offs

Etienne Dallaire

专题命中 评测与基准 :LLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02571 2025-11-13 cs.NI 50%

ASINT: Learning AS-to-Organization Mapping from Internet Metadata

Yongzhe Xu, Weitong Li, Eeshan Umrani, Taejoong Chung

专题命中 评测与基准 :LLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏