arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-02-16 至 2026-02-16 共收录 36 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 36 篇

2406.14045 2026-02-16 cs.LG cs.AI 90%

LTSM-Bundle: A Toolbox and Benchmark on Large Language Models for Time Series Forecasting

LTSM-Bundle: 一个用于时间序列预测的大型语言模型工具包和基准

Yu-Neng Chuang, Songchen Li, Jiayi Yuan, Guanchu Wang, Kwei-Herng Lai, Joshua Han, Zihang Xu, Songyuan Sui, Leisheng Yu, Sirui Ding, Chia-Yuan Chang, Alfredo Costilla Reyes, Daochen Zha, Xia Hu

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);prompting(abstract);分类 cs.AI、cs.LG

AI总结 LTSM-Bundle提出了一种用于时间序列预测的大型语言模型工具包和基准,通过模块化和多维度基准测试,结合最优设计选择,提升了零样本和少量样本性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13084 2026-02-16 cs.CL 88%

Exploring a New Competency Modeling Process with Large Language Models

探索基于大语言模型的新能力建模过程

Silin Du, Manqing Xin, Raymond Jia Wang

机构 * School of Economics and Management, Tsinghua University, Beijing, China(经济管理学院,清华大学,北京,中国) Bill-JC Technology, Wuhan, China(Bill-JC科技,武汉,中国)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 本研究提出基于大语言模型的新能力建模方法,通过结构化计算组件提升建模的透明性、数据驱动性和可验证性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12631 2026-02-16 cs.AI cs.HC cs.LG 86%

AI Agents for Inventory Control: Human-LLM-OR Complementarity

面向库存控制的AI代理:人类-大语言模型-运筹学互补性

Jackie Baek, Yaopeng Fu, Will Ma, Tianyi Peng

机构 * Stern School of Business, New York University(纽约大学斯特恩商学院) Columbia University(哥伦比亚大学) Graduate School of Business and Data Science Institute, Columbia University(哥伦比亚大学商学院与数据科学研究院)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文研究了人类、大语言模型和运筹学算法在库存控制中的互补性,通过实验发现人机协作能提升利润,且大量个体从中受益。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22968 2026-02-16 cs.CE cs.AI cs.CL 86%

Redefining Evaluation Standards: A Unified Framework for Evaluating the Korean Capabilities of Language Models

重新定义评估标准:一种统一评估韩语模型能力的框架

Hanwool Lee, Dasol Choi, Sooyong Kim, Ilgyun Jeong, Sangwon Baek, Guijin Son, Inseon Hwang, Naeun Lee, Seunghyeok Hong

机构 * AIM Intelligence Seoul National University(首尔国立大学) A.I.MATICS TigerCompany Catius National Assembly of Korea(韩国国会) Coupang Hankuk University of Foreign Studies(韩国外交大学)

专题命中 评测与基准 :language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出HRET框架,通过统一评估韩语大语言模型的能力,结合多种评估方法和数据集,提升模型在韩语任务中的表现和诊断能力。

Comments Accepted at LREC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13042 2026-02-16 cs.LG 85%

GPTZero: Robust Detection of LLM-Generated Texts

GPTZero:鲁棒的LLM生成文本检测

George Alexandru Adam, Alexander Cui, Edwin Thomas, Emily Napier, Nazar Shmatko, Jacob Schnell, Jacob Junqi Tian, Alekhya Dronavalli, Edward Tian, Dongwon Lee

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 GPTZero通过分层多任务架构实现鲁棒的LLM生成文本检测,提供准确可解释的检测方法和负责任的使用教育。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12779 2026-02-16 cs.HC 85%

iRULER: Intelligible Rubric-Based User-Defined LLM Evaluation for Revision

iRULER: 可理解的基于评分标准的用户自定义LLM评估用于修改

Jingwen Bai, Wei Soon Cheong, Philippe Muller, Brian Y Lim

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract)

AI总结 iRULER通过特定评分标准和可解释反馈,提升写作审查和修改的可理解性和有效性。

Comments To Appear at CHI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16882 2026-02-16 cs.AI cs.LG cs.SI 84%

SaVe-TAG: LLM-based Interpolation for Long-Tailed Text-Attributed Graphs

SaVe-TAG:基于大语言模型的长尾文本属性图插值

Leyao Wang, Yu Wang, Bo Ni, Yuying Zhao, Hanyu Wang, Yao Ma, Tyler Derr

机构 * Yale University(耶鲁大学) University of Oregon(俄勒冈大学) Vanderbilt University(范德比大学) Renmin University of China(中国人民大学) Rensselaer Polytechnic Institute(罗切斯特理工学院)

专题命中 评测与基准 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 SaVe-TAG通过大语言模型实现文本属性图的长尾问题解决,生成合成样本并结合图拓扑过滤提升学习效果。

Comments Accepted KDD 2026 Research Track Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12892 2026-02-16 cs.CV cs.AI cs.CL 82%

RADAR: Revealing Asymmetric Development of Abilities in MLLM Pre-training

RADAR: 揭示多模态大语言模型预训练中能力的非对称发展

Yunshuang Nie, Bingqian Lin, Minzhe Niu, Kun Xiang, Jianhua Han, Guowei Huang, Xingyue Quan, Hang Xu, Bokui Chen, Xiaodan Liang

机构 * Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) Peng Cheng Laboratory(鹏城实验室) Guangdong Key Laboratory of Big Data Analysis and Processing(广东大数据分析与处理重点实验室) Tsinghua Shenzhen International Graduate School(清华大学深圳国际 Graduate School) Tsinghua University(清华大学) Shanghai Jiao Tong University(上海交通大学) Yinwang Intelligent Technology Co., Ltd.(亿纬智能科技有限公司) Huawei’s 2012 Lab(华为2012实验室)

专题命中 评测与基准 :large language model(abstract);language model(abstract);pretraining(abstract);post-training(abstract)

AI总结 RADAR提出了一种高效的以能力为中心的评估框架,用于揭示多模态大语言模型预训练中感知和推理能力的非对称发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26510 2026-02-16 cs.LG stat.ML 81%

LLMs as In-Context Meta-Learners for Model and Hyperparameter Selection

大语言模型作为上下文元学习者用于模型和超参数选择

Youssef Attia El Hili, Albert Thomas, Malik Tiomoko, Abdelhakim Benechehab, Corentin Léger, Corinne Ancourt, Balázs Kégl

机构 * Huawei Noah’s Ark Lab(华为诺亚实验室) Centre de Recherche en Informatique, Mines Paris, PSL University(信息研究中心,巴黎 Mines,PSL 大学) Department of Data Science, EURECOM(数据科学系,EURECOM)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文提出利用大语言模型作为上下文元学习者,通过数据集元数据推荐模型和超参数,无需搜索即可实现高效优化。

Comments 27 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10661 2026-02-16 cs.CL 80%

Targeted Syntactic Evaluation of Language Models on Georgian Case Alignment

针对格鲁吉亚语的语法评估:变换语义分析

Daniel Gallagher, Gerhard Heyer

机构 * Institute for Applied Informatics (InfAI)(应用信息研究所)

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL

AI总结 本文针对格鲁吉亚语的语法评估,通过生成特定语法测试数据集,评估不同语言模型在处理分裂-反向语态对齐任务中的表现。

Comments To appear in Proceedings of The Second Workshop on Language Models for Low-Resource Languages (LoResLM), EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12806 2026-02-16 cs.CL cs.AI cs.CR cs.LG 80%

RAT-Bench: A Comprehensive Benchmark for Text Anonymization

RAT-Bench:文本匿名化的综合基准

Nataša Krčo, Zexi Yao, Matthieu Meeus, Yves-Alexandre de Montjoye

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 RAT-Bench通过评估文本匿名化工具的重新识别风险,揭示了LLM基工具在隐私与效用平衡上的优势,尽管计算成本较高且适用于多种语言。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01096 2026-02-16 cs.CV cs.CL 79%

Evaluating Vision Language Model Adaptations for Radiology Report Generation in Low-Resource Languages

评估用于低资源语言放射报告生成的视觉语言模型适应

Marco Salmè, Rosa Sicilia, Paolo Soda, Valerio Guarrasi

机构 * Department of Diagnostics and Intervention, Radiation Physics, Biomedical Engineering, Umeå University(诊断与介入系,辐射物理,生物医学工程,乌梅大学)

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL

AI总结 本研究评估了低资源语言中视觉语言模型在放射报告生成中的适应性,发现语言特定模型和领域特定训练能显著提升报告质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12659 2026-02-16 cs.CV cs.AI 79%

IndicFairFace: Balanced Indian Face Dataset for Auditing and Mitigating Geographical Bias in Vision-Language Models

IndicFairFace: 用于审计和缓解视觉-语言模型中地理偏见的平衡印度面部数据集

Aarish Shah Mohsin, Mohammed Tayyab Ilyas Khan, Mohammad Nadeem, Shahab Saquib Sohail, Erik Cambria, Jiechao Gao

专题命中 评测与基准 :language model(title,abstract);分类 cs.AI

AI总结 IndicFairFace是一个平衡的印度面部数据集,用于审计和缓解视觉-语言模型中的地理偏见,通过去偏方法减少偏见并保持检索准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08236 2026-02-16 cs.CL cs.CY 79%

Exploring Safety Alignment Evaluation of LLMs in Chinese Mental Health Dialogues via LLM-as-Judge

通过LLM-as-Judge探索LLM在中文心理健康对话中的安全对齐评估

Yunna Cai, Fan Wang, Haowei Wang, Kun Wang, Kailai Yang, Sophia Ananiadou, Moyan Li, Mingming Fan

机构 * HKUST(GZ)(香港科技大学(广州)) Wuhan Univ.(武汉大学) Univ. of Iowa(爱荷华大学) Univ. of Manchester(曼彻斯特大学)

专题命中 评测与基准 :LLM(title,abstract);分类 cs.CL

AI总结 本文提出PsyCrisis-Bench,通过LLM-as-Judge方法评估LLM在中文心理健康对话中的安全对齐性,提供高质量数据集和可解释的评估工具。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13148 2026-02-16 cs.SD 78%

Can Large Audio Language Models Understand Audio Well? Speech, Scene and Events Understanding Benchmark for LALMs

大音频语言模型能理解音频吗?LALMs的语音、场景和事件理解基准

Han Yin, Jung-Woo Choi

机构 * School of Electrical Engineering, KAIST, Daejeon, Republic of Korea(韩国成均馆大学电气工程学院)

专题命中 评测与基准 :language model(title,abstract)

AI总结 本文提出SSEU-Bench,首个考虑语音与非语音音频能量差异的音频理解基准,通过链式思维提升LALMs在联合理解任务中的性能。

Comments Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06971 2026-02-16 cs.CV cs.RO eess.IV 78%

Hallucinating 360°: Panoramic Street-View Generation via Local Scenes Diffusion and Probabilistic Prompting

生成360°全景图:通过局部场景扩散与概率提示的街道视图生成

Fei Teng, Kai Luo, Sheng Wu, Siyu Li, Pujun Guo, Jiale Wei, Jiaming Zhang, Kunyu Peng, Kailun Yang

机构 * School of Artificial Intelligence and Robotics and the National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University(人工智能与机器人学院和机器人视觉感知与控制技术国家工程研究中心,湖南大学) Institute for Anthropomatics and Robotics, Karlsruhe Institute of Technology(人机学与机器人研究所,卡尔斯鲁厄技术大学)

专题命中 评测与基准 :prompting(title,abstract)

AI总结 本文提出Percep360方法,通过局部场景扩散与概率提示技术,实现自动驾驶中的可控全景图像生成,提升真实场景下的鸟瞰图分割性能。

Comments Accepted to ICRA 2026. The source code will be publicly available at https://github.com/FeiT-FeiTeng/Percep360

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18401 2026-02-16 cs.CL 77%

Evaluating the Creativity of LLMs in Persian Literary Text Generation

评估大语言模型在波斯文学文本生成中的创造力

Armin Tourajmehr, Mohammad Reza Modarres, Yadollah Yaghoobzadeh

机构 * Tehran Institute for Advanced Studies, Khatam University, Iran(德黑兰高级研究学院,卡坦大学,伊朗) School of Electrical and Computer Engineering, College of Engineering, University of Tehran, Tehran, Iran(电气与计算机工程学院,工程学院,德黑兰大学,德黑兰,伊朗)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文评估了LLMs在生成波斯文学文本中的创造力,通过定制测试和人工验证,分析其在原创性、流畅性、灵活性和扩展性方面的表现,并探讨其在文学修辞手法应用上的能力。

Journal ref In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 14762-14774, Suzhou, China. Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19558 2026-02-16 cs.CY cs.LG 77%

PoliCon: Evaluating LLMs on Achieving Diverse Political Consensus Objectives

PoliCon:评估LLMs在实现多样化政治共识目标上的能力

Zhaowei Zhang, Xiaobo Wang, Minghua Yi, Mengmeng Wang, Fengshuo Bai, Zilong Zheng, Yipeng Kang, Yaodong Yang

机构 * Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院) USTC(中国科学技术大学) WHU(武汉大学) SJTU(上海交通大学) State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI) Zhongguancun Academy(中关村学院)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 PoliCon通过构建基于欧洲议会辩论记录的基准,评估LLMs在不同政治环境下生成共识决议的能力,揭示其在复杂任务中的不足及党派偏见。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07978 2026-02-16 cs.AI cs.CL cs.LG 75%

VoiceAgentBench: Are Voice Assistants ready for agentic tasks?

VoiceAgentBench: 聊天助手是否准备好处理代理任务?

Dhruv Jain, Harshit Shukla, Gautam Rajeev, Ashish Kulkarni, Chandra Khatri, Shubham Agarwal

机构 * OLA Electric(OLA电讯) Krutrim AI(Krutrim人工智能)

专题命中 评测与基准 :LLM(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 VoiceAgentBench评估语音模型在代理任务中的表现,发现ASR-LLM在英语任务中表现优于端到端SpeechLMs,但两者在多语言和安全评估中均存在局限。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11151 2026-02-16 cs.LG cs.CL cs.IR 73%

Diffusion-Pretrained Dense and Contextual Embeddings

基于扩散预训练的密集和上下文嵌入

Sedigheh Eslami, Maksim Gaiduk, Markus Krimmel, Louis Milliken, Bo Wang, Denis Bykov

专题命中 评测与基准 :language model(abstract);pretraining(abstract);分类 cs.CL、cs.LG

AI总结 本文提出pplx-embed模型,通过多阶段对比学习在扩散预训练语言模型基础上实现多语言嵌入,其中pplx-embed-v1在多个基准测试中表现优异,pplx-embed-context-v1在ConTEB基准中创纪录。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08543 2026-02-16 cs.CL cs.AI cs.IR 73%

GISA: A Benchmark for General Information-Seeking Assistant

GISA:通用信息检索助手的基准测试

Yutao Zhu, Xingshuo Zhang, Maosen Zhang, Jiajie Jin, Liancheng Zhang, Xiaoshuai Song, Kangzhi Zhao, Wencong Zeng, Ruiming Tang, Han Li, Ji-Rong Wen, Zhicheng Dou

机构 * Renmin University of China(中国人民大学) Kuaishou Technology(快手科技)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 GISA是一个针对通用信息检索助手的基准测试,包含373个人工设计的查询,旨在评估信息检索任务中深度推理与信息聚合的能力。

Comments Project repo: https://github.com/RUC-NLPIR/GISA

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13028 2026-02-16 cs.CV cs.CL 70%

Human-Aligned MLLM Judges for Fine-Grained Image Editing Evaluation: A Benchmark, Framework, and Analysis

面向细粒度图像编辑评估的人类对齐MLLM评判:一个基准、框架和分析

Runzhou Liu, Hailey Weingord, Sejal Mittal, Prakhar Dungarwal, Anusha Nandula, Bo Ni, Samyadeep Basu, Hongjie Chen, Nesreen K. Ahmed, Li Li, Jiayi Zhang, Koustava Goswami, Subhojyoti Mukherjee, Branislav Kveton, Puneet Mathur, Franck Dernoncourt, Yue Zhao, Yu Wang, Ryan A. Rossi, Zhengzhong Tu, Hongru Du

机构 * University of Virginia(弗吉尼亚大学) Columbia University(哥伦比亚大学) Vanderbilt University(范德比大学) Adobe Research(Adobe研究) Dolby Laboratories(杜比实验室) Cisco Research(思科研究) University of Southern California(南加州大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) University of Oregon(俄勒冈大学) Texas A&M University(德克萨斯大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出细粒度MLLM评判框架,通过分解十二个可解释因素,提升图像编辑评估的精度与实用性,为研究和改进图像编辑方法提供实用基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16543 2026-02-16 cs.AI 70%

Mathematics and Machine Creativity: A Survey on Bridging Mathematics with AI

数学与人工智能:连接数学与AI的综述

Shizhe Liang, Wei Zhang, Tianyang Zhong, Tianming Liu

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文综述了AI在数学研究中的应用,探讨了AI通过提供灵活框架和归纳推理能力,如何支持数学研究,并促进数学与AI的跨学科理解。

Comments This article is withdrawn due to internal authorship and supervisory considerations that require clarification before the work can proceed in its current form. After further review, I believe it is appropriate to pause and formally resolve these matters to ensure full compliance with institutional and collaborative research policies

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12641 2026-02-16 cs.NI cs.AI cs.HC cs.MM 70%

Artic: AI-oriented Real-time Communication for MLLM Video Assistant

Artic: 面向AI的实时通信用于多模态大语言模型视频助手

Jiangkai Wu, Zhiyuan Ren, Junquan Zhong, Liming Liu, Xinggong Zhang

机构 * Peking University(北京大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 Artic提出了一种面向AI的实时通信框架,通过自适应比特率、上下文感知流式传输和退化视频理解基准提升多模态大语言模型视频助手的准确性和降低延迟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12533 2026-02-16 cs.LG 70%

AMPS: Adaptive Modality Preference Steering via Functional Entropy

AMPS: 通过功能熵实现自适应模态偏好引导

Zihan Huang, Xintong Li, Rohan Surana, Tong Yu, Rui Wang, Julian McAuley, Jingbo Shang, Junda Wu

机构 * University of California, San Diego(加州大学圣地亚哥分校) Adobe Research(Adobe研究)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.LG

AI总结 AMPS通过实例感知的诊断度量和可学习模块,实现对多模态大语言模型模态偏好的自适应引导,有效调节偏好并降低生成错误率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12356 2026-02-16 cs.AI 70%

A Theoretical Framework for Adaptive Utility-Weighted Benchmarking

为适应性效用加权基准评估构建的理论框架

Philip Waggoner

机构 * Stanford University(斯坦福大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出了一种理论框架,通过多层次适应性网络连接评估指标、模型组件和利益相关者,以更全面地理解评估应代表什么,并通过人参与的更新规则实现动态演变的基准结构。

Comments 10 page, no figures, 40 equations

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22620 2026-02-16 cs.CL 70%

Layer-wise Swapping for Generalizable Multilingual Safety

分层交换以实现通用多语言安全

Hyunseo Shin, Wonseok Hwang

机构 * University of Seoul(首尔大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出分层交换方法,通过将英语安全专家的安全对齐转移到低资源语言专家,提升多语言安全性能,同时保持通用任务性能。

Comments EACL 2026 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09725 2026-02-16 cs.CL 70%

Assessing and Improving Punctuation Robustness in English-Marathi Machine Translation

评估和改进英语-马拉雅尔语机器翻译中的标点鲁棒性

Kaustubh Shivshankar Shejole, Sourabh Deoghare, Pushpak Bhattacharyya

机构 * Computation for Indian Language Technology (CFILT) Department of Computer Science and Engineering(印度语言技术计算(CFILT)部门计算机科学与工程系) Indian Institute of Technology Bombay(印度班加罗尔理工学院)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本研究通过英语到马拉雅尔语的翻译任务,评估并改进NMT系统在处理标点歧义时的鲁棒性,提出两种修复策略并发现任务特定方法优于大型语言模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12878 2026-02-16 cs.CY 67%

Understanding Cultural Alignment in Multilingual LLMs via Natural Debate Statements

通过自然辩论陈述理解多语言大语言模型中的文化契合

Vlad-Andrei Negru, Camelia Lemnaru, Mihai Surdeanu, Rodica Potolea

专题命中 评测与基准 :large language model(abstract);language model(abstract)

AI总结 本文通过分析自然辩论陈述,揭示了多语言大语言模型在不同文化背景下的价值观差异及其对用户社会文化背景适应能力的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24739 2026-02-16 cs.HC cs.IT math.IT stat.ME 67%

Human- vs. AI-generated tests: dimensionality and information accuracy in latent trait evaluation

人类生成与人工智能生成的测验:潜在特质评估中的维度性和信息准确性

Mario Angelelli, Morena Oliva, Serena Arima, Enrico Ciavolino

专题命中 评测与基准 :large language model(abstract);language model(abstract)

AI总结 本研究比较了AI生成与人类开发的BAQ问卷,发现AI生成的问卷在维度性和潜在特质信息分布上存在差异,强调了AI工具准确性验证的重要性。

Comments 28 pages, 12 figures. Minor corrections and comments added. The published version of this preprint is available in "Statistics" at the following DOI: 10.1080/02331888.2025.2610647

详情

展开后加载摘要…

URL PDF HTML 收藏