arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 32411 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 32411 篇

2409.12656 2024-09-20 cs.CL 86%

Efficient Performance Tracking: Leveraging Large Language Models for Automated Construction of Scientific Leaderboards

Furkan Şahinuç, Thy Thy Tran, Yulia Grishina, Yufang Hou, Bei Chen, Iryna Gurevych

专题命中 评测与基准 :large language model(title);language model(title);LLM(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.09158 2024-08-12 cs.AI 86%

Improving Large Language Models in Event Relation Logical Prediction

Meiqi Chen, Yubo Ma, Kaitao Song, Yixin Cao, Yan Zhang, Dongsheng Li

专题命中 评测与基准 :large language model(title);language model(title);LLM(abstract);分类 cs.AI

Comments ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.04270 2024-08-09 cs.CL 86%

Analysis of Argument Structure Constructions in the Large Language Model BERT

Pegah Ramezani, Achim Schilling, Patrick Krauss

专题命中 评测与基准 :language model(title,abstract);large language model(title);分类 cs.CL

Comments arXiv admin note: text overlap with arXiv:2408.03062

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02637 2024-08-06 cs.CR cs.LG 86%

Command-line Obfuscation Detection using Small Language Models

Vojtech Outrata, Michael Adam Polak, Martin Kopp

专题命中 评测与基准 :language model(title,abstract);small language model(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.02691 2024-08-06 cs.CL cs.CR 86%

InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents

Qiusi Zhan, Zhixiang Liang, Zifan Ying, Daniel Kang

专题命中 评测与基准 :large language model(title);language model(title);LLM(abstract);分类 cs.CL

Comments 36 pages, 6 figures, 13 tables (ACL 2024 Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15154 2024-07-23 cs.CL 86%

Fine-grained Gender Control in Machine Translation with Large Language Models

Minwoo Lee, Hyukhun Koh, Minsung Kim, Kyomin Jung

专题命中 评测与基准 :large language model(title);language model(title);prompting(abstract);分类 cs.CL

Comments NAACL 2024 Main track long paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.17760 2024-05-31 cs.CL 86%

Constructions Are So Difficult That Even Large Language Models Get Them Right for the Wrong Reasons

Shijia Zhou, Leonie Weissweiler, Taiqi He, Hinrich Schütze, David R. Mortensen, Lori Levin

专题命中 评测与基准 :large language model(title);language model(title);LLM(abstract);分类 cs.CL

Comments LREC-COLING 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12786 2024-05-31 cs.CL eess.AS 86%

Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations

Guan-Ting Lin, Cheng-Han Chiang, Hung-yi Lee

专题命中 评测与基准 :large language model(title);language model(title);LLM(abstract);分类 cs.CL

Comments Accepted by ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.07463 2024-04-04 cs.CL 86%

MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks

Sanchit Ahuja, Divyanshu Aggarwal, Varun Gumma, Ishaan Watts, Ashutosh Sathe, Millicent Ochieng, Rishav Hada, Prachi Jain, Maxamed Axmed, Kalika Bali, Sunayana Sitaram

专题命中 评测与基准 :large language model(title);language model(title);LLM(abstract);分类 cs.CL

Comments 40 pages, 35 figures and 34 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14691 2024-04-02 cs.CY cs.AI 86%

Large Language Models and User Trust: Consequence of Self-Referential Learning Loop and the Deskilling of Healthcare Professionals

Avishek Choudhury, Zaria Chaudhry

专题命中 评测与基准 :large language model(title);language model(title);LLM(abstract);分类 cs.AI

Comments 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.19548 2024-03-29 cs.CL 86%

WaterJudge: Quality-Detection Trade-off when Watermarking Large Language Models

Piotr Molenda, Adian Liusie, Mark J. F. Gales

专题命中 评测与基准 :large language model(title);language model(title);LLM(abstract);分类 cs.CL

Comments NAACL 2024 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16386 2024-03-26 cs.CV cs.AI 86%

Dia-LLaMA: Towards Large Language Model-driven CT Report Generation

Zhixuan Chen, Luyang Luo, Yequan Bie, Hao Chen

专题命中 评测与基准 :large language model(title);language model(title);LLM(abstract);分类 cs.AI

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09818 2024-03-14 cs.CL cs.PL 86%

SUQL: Conversational Search over Structured and Unstructured Data with Large Language Models

Shicheng Liu, Jialiang Xu, Wesley Tjangnaka, Sina J. Semnani, Chen Jie Yu, Monica S. Lam

专题命中 评测与基准 :large language model(title);language model(title);LLM(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03336 2024-03-07 cs.CL cs.SI 86%

Scope of Large Language Models for Mining Emerging Opinions in Online Health Discourse

Joseph Gatto, Madhusudan Basak, Yash Srivastava, Philip Bohlman, Sarah M. Preum

专题命中 评测与基准 :large language model(title);language model(title);LLM(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01711 2024-02-06 cs.CY cs.AI 86%

LLM on FHIR -- Demystifying Health Records

Paul Schmiedmayer, Adrit Rao, Philipp Zagar, Vishnu Ravi, Aydin Zahedivash, Arash Fereydooni, Oliver Aalami

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract,comments);language model(abstract,comments);分类 cs.AI

Comments Pre-print of the paper submitted to the Call for Papers for the Special Focus Issue on ChatGPT and Large Language Models (LLMs) in Biomedicine and Health at the Journal of the American Medical Informatics Association: https://academic.oup.com/jamia/pages/call-for-papers-for-special-focus-issue

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.17645 2024-02-01 cs.IR cs.AI 86%

ReSLLM: Large Language Models are Strong Resource Selectors for Federated Search

Shuai Wang, Shengyao Zhuang, Bevan Koopman, Guido Zuccon

专题命中 评测与基准 :large language model(title);language model(title);LLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.20499 2023-11-07 cs.CL 86%

Leveraging Word Guessing Games to Assess the Intelligence of Large Language Models

Tian Liang, Zhiwei He, Jen-tse Huang, Wenxuan Wang, Wenxiang Jiao, Rui Wang, Yujiu Yang, Zhaopeng Tu, Shuming Shi, Xing Wang

专题命中 评测与基准 :large language model(title);language model(title);LLM(abstract);分类 cs.CL

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.13233 2023-09-26 cs.CL 86%

User Simulation with Large Language Models for Evaluating Task-Oriented Dialogue

Sam Davidson, Salvatore Romeo, Raphael Shu, James Gung, Arshit Gupta, Saab Mansour, Yi Zhang

专题命中 评测与基准 :language model(title,abstract);large language model(title);分类 cs.CL

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.15078 2023-09-20 cs.CL 86%

Large Language Models are Diverse Role-Players for Summarization Evaluation

Ning Wu, Ming Gong, Linjun Shou, Shining Liang, Daxin Jiang

专题命中 评测与基准 :large language model(title);language model(title);prompting(abstract);分类 cs.CL

Comments NLPCC 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.01674 2023-09-06 cs.CV cs.AI cs.DL 86%

Prompt me a Dataset: An investigation of text-image prompting for historical image dataset creation using foundation models

Hassan El-Hajj, Matteo Valleriani

专题命中 评测与基准 :foundation model(title,abstract);prompting(title);分类 cs.AI

Comments 12 pages, 3 figures, Accepted in ICIAP2023, AI4DH workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.14852 2023-07-28 cs.CL 86%

ArcGPT: A Large Language Model Tailored for Real-world Archival Applications

Shitou Zhang, Jingrui Hou, Siyuan Peng, Zuchao Li, Qibiao Hu, Ping Wang

专题命中 评测与基准 :large language model(title);language model(title);LLM(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.10534 2023-06-07 cs.CL 86%

DISCO: Distilling Counterfactuals with Large Language Models

Zeming Chen, Qiyue Gao, Antoine Bosselut, Ashish Sabharwal, Kyle Richardson

专题命中 评测与基准 :language model(title,abstract);large language model(title);分类 cs.CL

Comments ACL 2023 camera ready, final title change

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.01318 2023-02-03 cs.CL 86%

Accelerating Large Language Model Decoding with Speculative Sampling

Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, John Jumper

专题命中 评测与基准 :language model(title,abstract);large language model(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.00607 2022-11-02 cs.CL 86%

A Systematic Investigation of Commonsense Knowledge in Large Language Models

Xiang Lorraine Li, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d'Autume, Phil Blunsom, Aida Nematzadeh

专题命中 评测与基准 :language model(title,abstract);large language model(title);分类 cs.CL

Comments Accepted to EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.11689 2022-10-24 cs.CL 86%

SLING: Sino Linguistic Evaluation of Large Language Models

Yixiao Song, Kalpesh Krishna, Rajesh Bhatt, Mohit Iyyer

专题命中 评测与基准 :language model(title,abstract);large language model(title);分类 cs.CL

Comments 29 pages, EMNLP 2022 camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.03374 2021-07-15 cs.LG 86%

Evaluating Large Language Models Trained on Code

Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, Wojciech Zaremba

专题命中 评测与基准 :language model(title,abstract);large language model(title);分类 cs.LG

Comments corrected typos, added references, added authors, added acknowledgements

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16344 2024-05-28 cs.RO 86%

Large Language Models Enable Automated Formative Feedback in Human-Robot Interaction Tasks

Emily Jensen, Sriram Sankaranarayanan, Bradley Hayes

专题命中 评测与基准 :large language model(title);language model(title);LLM(abstract,comments)

Comments Presented at Human-LLM Interaction Workshop at HRI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02950 2026-08-28 cs.CL cs.AI cs.HC 版本更新 86%

Distinct Profiles of Run-to-Run Score Reliability and Expert-Panel Alignment Across Four LLM Evaluators of Simulated Japanese-Language AI-to-AI Counseling

固定合成日语咨询对话中的结构化提示与自动评估

Keita Kiuchi, Yoshikazu Fujimoto, Hideyuki Gotō, Tomonori Hosokawa, Makoto Nishimura, Yōsuke Satō, Izumi Sezai, Tomohiro Inoue

机构 * Japan National Institute of Occupational Safety and Health(日本国立职业安全卫生研究所) Kaze To Taiyo(凯泽・太阳) Saga Occupational Health Association(Saga职业健康协会) Department of Pharmacy, Zikei Hospital/Zikei Institute of Psychiatry(药剂科,Zikei医院/Zikei精神医学研究所) Department of Medical Welfare, Suzuka University of Medical Science(医疗福祉科, Suzuka医科大学) Graduate School of Human Sciences, Ritsumeikan University(人类科学研究生院,立命馆大学) Faculty of Nursing, National Defense Medical College(护理学部,国家防卫医疗大学) Support Center for Students with Disabilities, Aoyama Gakuin University(残疾学生支持中心,上智大学)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 研究通过人工智能生成的18份固定日语咨询对话记录,对比GPT - minimal、GPT - SMDP和Claude - SMDP三种情况,咨询专家与新的语言模型分别评级,发现SMDP对话在多方面获更高专家评级,语言模型评级可重复但偏宽松。

Comments 55 pages, 2 figures, 31 tables; supplemental material included; data and code at https://doi.org/10.5281/zenodo.22106028; preregistration and amendments at https://doi.org/10.17605/OSF.IO/VU286

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01666 2026-08-27 cs.CL cs.AI 版本更新 86%

Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation

风格取胜,实质落败:对作为创意生成评判者的大语言模型(LLM)的诊断

Fengxian Ji, Yuke Li, Jingpu Yang, Juanfan Wu, Fan Zhang, Zhexuan Cui, Yu Xie, Min Peng, Qianqian Xie, Xiuying Chen, Zhuohan Xie

专题命中 评测与基准 :LLM(title,title_cn);分类 cs.CL、cs.AI

AI总结 本研究针对大语言模型作为创意评判者的风格偏差问题,提出SciStyleBench基准及配套的SciStyleExtractor模块,实验验证该模块可有效降低风格偏差、提升实质区分度,为科学创意评估提供了系统性解决方案。

Comments First three authors are co-first authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21766 2026-08-25 cs.CL cs.AI 新提交 86%

Evaluation Awareness in Language Models: Representation, Verbalization, and Control

语言模型中的评估感知:表示、语言化与控制

Farzaneh Heidari, Amin Memarian, Guillaume Rabusseau

专题命中 评测与基准 :language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 该研究系统探测了6个语言模型的评估感知,发现其可线性解码、与语言化部分对齐,且在Olmo模型中随微调放大,需评估时考虑模型内部表示与语言化的脱节。

Comments The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏