arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-11-21 至 2025-11-21 共收录 148 信号源:cs.CL, cs.AI, cs.LG

1. 预训练与数据 16 篇

2511.16135 2025-11-21 physics.optics cs.AI 90%

CoSP: Reconfigurable Multi-State Metamaterial Inverse Design via Contrastive Pretrained Large Language Model

CoSP: 通过对比预训练大语言模型实现可重构多状态超材料逆向设计

Shujie Yang, Xuzhe Zhao, Yuqi Zhang, Yansong Tang, Kaichen Dong

机构 * Center of Double Helix, Tsinghua Shenzhen International Graduate School, Tsinghua University(双螺旋中心,清华大学深圳国际研究生院,清华大学) Intelligent Passive Thermal Control Center, Research Institute of Tsinghua University in Shenzhen(智能被动热控制中心,清华大学深圳研究院)

专题命中 预训练与数据 :large language model(title,abstract);language model(title,abstract);LLM(abstract);pretraining(abstract)

AI总结 CoSP通过对比预训练大语言模型实现可重构多状态超材料的逆向设计,能高效生成具有目标光学性能的材料结构。

Comments 5 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14023 2025-11-21 cs.CL 90%

Synthetic Data Generation Using Large Language Models: Advances in Text and Code

利用大语言模型生成合成数据:文本与代码领域的进展

Mihai Nadas, Laura Diosan, Andreea Tomescu

机构 * Faculty of Mathematics and Computer Science, Babeş-Bolyai University(巴贝什-博耶亚大学数学与计算机科学系) KlusAI Research Lab(KlusAI研究实验室)

专题命中 预训练与数据 :large language model(title,abstract);language model(title,abstract);LLM(abstract);instruction tuning(abstract)

AI总结 本文探讨了利用大语言模型生成合成数据在文本和代码领域的新进展,分析了其在低资源任务和代码应用中的潜力及挑战。

Comments 24 pages, 6 tables, 1 figure, 64 references

Journal ref IEEE Access 13, 134615-134633 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09837 2025-11-21 cs.DC 88%

MoFa: A Unified Performance Modeling Framework for LLM Pretraining

MoFa: 一种统一的LLM预训练性能建模框架

Lu Zhao, Rong Shi, Shaoqing Zhang, Shangchao Su, Ziqing Yin, Zhiyan Cui, Hongfeng Sun, Baoguo He, Yueqiang Chen, Liang Dong, Xiyuan Li, Lingbin Wang, Lijun Ma, Qiang Huang, Ting Liu, Chong Wang, Can Wei

专题命中 预训练与数据 :LLM(title,abstract);pretraining(title,abstract)

AI总结 MoFa提出一种统一的LLM预训练性能建模框架,整合多维优化特征和容错机制,通过增强的成本模型和调优系统,提升预训练性能预测精度并提供系统设计指导。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04250 2025-11-21 stat.ME cs.AI cs.ET cs.IR stat.AP 85%

How many patients could we save with LLM priors?

使用 LLM 先验知识能挽救多少患者?

Shota Arai, David Selby, Andrew Vargo, Sebastian Vollmer

专题命中 预训练与数据 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 利用LLM生成的先验知识优化多中心临床试验中的不良事件建模,提高预测性能并减少患者数量需求。

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16577 2025-11-21 cs.CL cs.AI 84%

Integrating Symbolic Natural Language Understanding and Language Models for Word Sense Disambiguation

整合符号自然语言理解与语言模型用于词义消歧

Kexin Zhao, Ken Forbus

专题命中 预训练与数据 :language model(title,abstract);LLM(abstract);分类 cs.CL、cs.AI

AI总结 本文提出利用统计语言模型和符号NLU系统整合的方法,实现无需人工标注数据的词义消歧。

Comments 16 pages

Journal ref Proceedings of the Twelfth Annual Conference on Advances in Cognitive Systems ACS-2025 (333-348)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15714 2025-11-21 cs.AI 83%

Majority Rules: LLM Ensemble is a Winning Approach for Content Categorization

多数决策:LLM集成是内容分类的获胜方法

Ariel Kamen, Yakov Kamen

机构 * RingCentral Inc.(环中央公司) Relevad Corporation(Relevad公司)

专题命中 预训练与数据 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本研究提出eLLM集成方法,通过整合多个模型提升无结构文本分类的准确性和鲁棒性,实现接近人类专家水平的性能。

Comments 17 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.01144 2025-11-21 eess.IV cs.AI cs.CV cs.LG 81%

LEARNER: Contrastive Pretraining for Learning Fine-Grained Patient Progression from Coarse Inter-Patient Labels

LEARNER: 通过对比学习从粗粒度患者间标签中学习细粒度患者进展

Jana Armouti, Nikhil Madaan, Rohan Panda, Tom Fox, Laura Hutchins, Amita Krishnan, Ricardo Rodriguez, Bennett DeBoisblanc, Deva Ramanan, John Galeotti, Gautam Gare

机构 * Carnegie Mellon University(卡内基梅隆大学) LSUHSC Internal Medicine(LSUHSC内科) Cosmetic Surgery Facility LLC(美容外科诊所有限公司)

专题命中 预训练与数据 :pretraining(title,abstract);分类 cs.AI、cs.LG

AI总结 LEARNER通过对比学习利用患者间粗粒度标签,学习细粒度患者内变化,提升个性化医学中的治疗响应预测性能。

Comments Under review at ISBI 2026 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27237 2025-11-21 cs.CV 78%

Fusion of Multi-scale Heterogeneous Pathology Foundation Models for Whole Slide Image Analysis

多尺度异构病理基础模型融合用于全切片图像分析

Zhidong Yang, Xiuhui Shi, Wei Ba, Zhigang Song, Haijing Luan, Taiyuan Hu, Senlin Lin, Jiguang Wang, Shaohua Kevin Zhou, Rui Yan

机构 * School of Biomedical Engineering, Division of Life Sciences and Medicine, University of Science and Technology of China(中国科学技术大学生物医学工程学院) Division of Life Science, Department of Chemical and Biological Engineering, State Key Laboratory of Nervous System Disorders, The Hong Kong University of Science and Technology(香港科技大学生命科学系) SIAT-HKUST Joint Laboratory of Cell Evolution and Digital Health, HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute(深圳-香港联合创新研究院细胞进化与数字健康联合实验室) Department of Hepatobiliary Surgery, The First Affiliated Hospital of USTC, Division of Life Sciences and Medicine, University of Science and Technology of China(中国科学技术大学附属第一医院肝胆外科) Center for Medical Imaging, Robotics, Analytic Computing & Learning (MIRACLE), Suzhou Institute for Advanced Research, USTC, Suzhou, Jiangsu, China(中国科学技术大学苏州先进研究所) Computer Network Information Center, Chinese Academy of Sciences(中国科学院计算机网络信息中心) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Jiangsu Provincial Key Laboratory of Multimodal Digital Twin Technology, Suzhou, Jiangsu, China(江苏省多模态数字孪生技术重点实验室) Key Laboratory of Precision and Intelligent Chemistry, USTC, Hefei, Anhui, China(中国科学技术大学精准与智能化学重点实验室) Department of Pathology, Chinese PLA General Hospital, Beijing, China(中国人民解放军总医院病理科)

专题命中 预训练与数据 :foundation model(title,abstract)

AI总结 本文提出FuseCPath框架,通过多尺度异构病理基础模型融合提升全切片图像分析性能。

Comments 22 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15192 2025-11-21 cs.AI 77%

As If We've Met Before: LLMs Exhibit Certainty in Recognizing Seen Files

似曾相识:LLM在识别已见过的文件时表现出确定性

Haodong Li, Jingqi Zhang, Xiao Cheng, Peihua Mai, Haoyu Wang, Yan Pang

机构 * Huazhong University of Science and Technology(华中科技大学) National University of Singapore(国立新加坡大学) Macquarie University(麦考瑞大学)

专题命中 预训练与数据 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 COPYCHECK利用LLM的不确定性信号,通过双策略检测训练数据中的受版权内容,实现高准确率的版权检测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16028 2025-11-21 cs.CY 75%

Can Online GenAI Discussion Serve as Bellwether for Labor Market Shifts?

在线生成式AI讨论能否作为劳动力市场转变的风向标?

Shurui Cao, Wenyue Hua, William Yang Wang, Hong Shen, Fei Fang

专题命中 预训练与数据 :LLM(abstract);large language model(abstract);language model(abstract)

AI总结 本文研究在线生成式AI讨论是否能作为劳动力市场转变的早期预测指标,通过分析讨论强度与就业变化的关系,证明其对劳动力动态的预测有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10819 2025-11-21 cs.AI cs.LG 73%

PoE-World: Compositional World Modeling with Products of Programmatic Experts

PoE-World: 通过程序专家的乘积进行组合世界建模

Wasu Top Piriyakulkij, Yichao Liang, Hao Tang, Adrian Weller, Marta Kryven, Kevin Ellis

机构 * Cornell University(康奈尔大学) University of Cambridge(剑桥大学) The Alan Turing Institute(艾伦·图灵研究所) Dalhousie University(达尔豪斯大学)

专题命中 预训练与数据 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 PoE-World通过程序专家的乘积有效建模复杂非网格世界领域,利用LLMs学习世界模型并实现高效泛化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15967 2025-11-21 cs.CV 71%

InfoCLIP: Bridging Vision-Language Pretraining and Open-Vocabulary Semantic Segmentation via Information-Theoretic Alignment Transfer

InfoCLIP: 通过信息论对齐转移连接视觉语言预训练与开放词汇语义分割

Muyao Yuan, Yuanhong Zhang, Weizhan Zhang, Lan Ma, Yuan Gao, Jiangyong Ying, Yudeng Xin

专题命中 预训练与数据 :pretraining(title)

AI总结 InfoCLIP通过信息论对齐转移提升开放词汇语义分割的性能,有效解决预训练CLIP在微调过程中的过拟合问题。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18515 2025-11-21 cs.CR cs.AI cs.MA cs.SE 70%

Securing Smart Contract Languages with a Unified Agentic Framework for Vulnerability Repair in Solidity and Move

用统一的代理框架保障智能合约语言安全:在Solidity和Move中修复漏洞

Rabimba Karanjai, Lei Xu, Weidong Shi

机构 * University Of Houston(德克萨斯大学休斯敦分校) Kent State University(肯特州立大学)

专题命中 预训练与数据 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出Smartify,一种利用多代理框架和LLMs自动检测并修复Solidity和Move智能合约漏洞的方法,展现其在提升安全性和可靠性方面的显著成效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05274 2025-11-21 cs.CV 67%

From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos

从游戏到回放:面向时间精细粒度视频的组合视频检索

Animesh Gupta, Jay Parmar, Ishan Rajendrakumar Dave, Mubarak Shah

专题命中 预训练与数据 :LLM(abstract);prompting(abstract)

AI总结 TF-CoVR提出了一种针对时间精细粒度视频检索的框架,通过预训练视频编码器和对比学习提升检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16132 2025-11-21 cs.LG 61%

An Interpretability-Guided Framework for Responsible Synthetic Data Generation in Emotional Text

可解释性引导的负责任情感文本合成数据生成框架

Paula Joy B. Martinez, Jose Marie Antonio Miñoza, Sebastian C. Ibañez

专题命中 预训练与数据 :LLM(abstract);分类 cs.LG;foundation model(journal_ref)

AI总结 本文提出一种基于SHAP的可解释性引导框架,用于生成情感文本的合成数据,提升少数情绪类别的分类性能,同时揭示合成文本在词汇丰富性和表达复杂性上的局限性。

Journal ref Association for the Advancement of Artificial Intelligence (2026). Shaping Responsible Synthetic Data in the Era of Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16454 2025-11-21 cs.CV 50%

LLaVA$^3$: Representing 3D Scenes like a Cubist Painter to Boost 3D Scene Understanding of VLMs

LLaVA$^3$:像立体主义画家一样表示3D场景以提升VLM的3D场景理解

Doriand Petit, Steve Bourgeois, Vincent Gay-Bellile, Florian Chabot, Loïc Barthe

专题命中 预训练与数据 :language model(abstract)

AI总结 LLaVA$^3$通过立体主义方法提升VLM对3D场景的理解能力,利用多视角2D图像无需微调实现更优的3D场景理解。

Comments Accepted at AAAI'26

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 指令微调 9 篇

2409.20059 2025-11-21 cs.CL 87%

Is Preference Alignment Always the Best Option to Enhance LLM-Based Translation? An Empirical Analysis

基于偏好对齐是否总是提升基于大语言模型的翻译质量的最佳选项?一项实证分析

Hippolyte Gisserot-Boukhlef, Ricardo Rei, Emmanuel Malherbe, Céline Hudelot, Pierre Colombo, Nuno M. Guerreiro

专题命中 指令微调 :LLM(title);large language model(abstract);language model(abstract);SFT(abstract)

AI总结 本研究通过实证分析探讨了基于偏好的对齐在提升大语言模型翻译质量中的效果,发现CPO在高质量数据上表现优异,但可能在下游度量上存在不稳定性,同时基础模型生成翻译的表现与多个外部系统相当且更一致。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16375 2025-11-21 cs.LG cs.AI 84%

Are Foundation Models Useful for Bankruptcy Prediction?

基础模型在破产预测中是否有效?

Marcin Kostrzewa, Oleksii Furman, Roman Furman, Sebastian Tomczak, Maciej Zięba

机构 * wrocław University of Science and Technology(沃拉日大学科学与技术学院) Tooploox(Tooploox公司) Opera

专题命中 指令微调 :foundation model(title,abstract);LLM(abstract);分类 cs.AI、cs.LG

AI总结 本文研究了基础模型在破产预测中的有效性,发现传统机器学习模型在性能上优于基础模型,且基础模型在风险敏感场景中存在概率估计不可靠的问题。

Comments NeurIPS 2025 Workshop: Generative AI in Finance

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00236 2025-11-21 cs.RO 78%

Sim2Real Diffusion: Leveraging Foundation Vision Language Models for Adaptive Automated Driving

Sim2Real Diffusion:利用基础视觉语言模型实现自适应自动驾驶

Chinmay Vilas Samak, Tanmay Vilas Samak, Bing Li, Venkat Krovi

机构 * Department of Automotive Engineering, Clemson University International Center for Automotive Research (CU-ICAR)(汽车工程系,克莱姆森大学国际汽车研究中心(CU-ICAR))

专题命中 指令微调 :language model(title);foundation model(abstract)

AI总结 本文提出Sim2Real Diffusion框架,利用基础视觉语言模型实现自动驾驶的跨领域适应,通过条件潜在扩散提升sim2real转换性能。

Comments Accepted in IEEE Robotics and Automation Letters (RA-L)

Journal ref IEEE Robotics and Automation Letters, vol. 11, no. 1, pp. 177-184, Jan. 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16110 2025-11-21 cs.CR 78%

Multi-Faceted Attack: Exposing Cross-Model Vulnerabilities in Defense-Equipped Vision-Language Models

多面攻击:揭示配备防御机制的视觉语言模型中的跨模型漏洞

Yijun Yang, Lichao Wang, Jianping Zhang, Chi Harold Liu, Lanqing Hong, Qiang Xu

专题命中 指令微调 :language model(title,abstract)

AI总结 多面攻击揭示了配备防御机制的视觉语言模型中的跨模型安全漏洞,通过注意力转移攻击和轻量级转移增强算法,实现了58.5%的成功率。

Comments AAAI 2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16324 2025-11-21 cs.CL cs.AI 73%

SDA: Steering-Driven Distribution Alignment for Open LLMs without Fine-Tuning

SDA:无微调的开放语言模型分布对齐框架

Wei Xia, Zhi-Hong Deng

专题命中 指令微调 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 SDA 提出了一种无需微调的开源 LLMs 分布对齐框架,通过用户指令动态调整输出概率,提升模型与人类意图的一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20470 2025-11-21 cs.CV 67%

Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence

Conan:基于多尺度视觉证据的逐步学习以像侦探一样推理

Kun Ouyang, Yuanxin Liu, Linli Yao, Yishuo Cai, Hao Zhou, Jie Zhou, Fandong Meng, Xu Sun

机构 * State Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学) WeChat AI, Tencent Inc., China(微信AI,腾讯公司,中国)

专题命中 指令微调 :large language model(abstract);language model(abstract)

AI总结 Conan通过多阶段渐进冷启动策略和AIR RLVR框架,实现证据基础的多步视频推理,超越基线模型,达到最先进的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15923 2025-11-21 cs.CV 67%

RB-FT: Rationale-Bootstrapped Fine-Tuning for Video Classification

RB-FT:基于理由的视频分类微调

Meilong Xu, Di Fu, Jiaxing Zhang, Gong Yu, Jiayu Zheng, Xiaoling Hu, Dongdi Zhao, Feiyang Li, Chao Chen, Yong Cao

机构 * Stony Brook University(石溪大学) ByteDance Inc.(字节跳动公司) Harvard Medical School(哈佛医学院)

专题命中 指令微调 :language model(abstract);SFT(abstract)

AI总结 RB-FT通过自动生成的理由提升视频分类性能,无需额外标注,有效提升模型对领域特定视频内容的理解能力。

Comments 11 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10848 2025-11-21 cs.LG cs.AI 62%

STAMP: Spatial-Temporal Adapter with Multi-Head Pooling

STAMP:带有多头池化的空间-时间适配器

Brad Shook, Abby Turner, Jieshi Chen, Michał Wiliński, Mononito Goswami, Jonathan Elmer, Artur Dubrawski

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Pittsburgh School of Medicine(匹兹堡大学医学学院)

专题命中 指令微调 :foundation model(abstract);分类 cs.AI、cs.LG

AI总结 STAMP通过多头池化机制,利用通用TSFMs生成的单变量嵌入,隐式建模EEG数据的空间-时间特性,实现与现有EEGFMs相当的性能。

Comments Accepted as a Proceedings paper at Machine Learning for Health (ML4H) 2025, invited presentation at the Time Series for Health (TS4H) Workshop, NeurIPS 2025. v2: Updated author affiliation and corrected a duplicated word in the text. No other changes

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21456 2025-11-21 cs.CL 57%

Diagnosing the Performance Trade-off in Moral Alignment: A Case Study on Gender Stereotypes

在道德一致性中的性能权衡诊断:关于性别刻板印象的案例研究

Guangliang Liu, Bocheng Chen, Han Zi, Xitong Zhang, Kristen Marie Johnson

机构 * Michigan State University(密歇根州立大学) University of Mississippi(密苏里大学) Northeastern University(东北大学)

专题命中 指令微调 :language model(abstract);分类 cs.CL

AI总结 本文通过分析性别刻板印象缓解中的性能权衡,发现当前公平性目标存在局限,整体遗忘与下游任务性能密切相关,选择性遗忘无法有效降低整体遗忘。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 后训练与偏好优化 2 篇

2511.15767 2025-11-21 cs.LG cs.AI cs.PL 88%

TB or Not TB: Coverage-Driven Direct Preference Optimization for Verilog Stimulus Generation

TB 或 TB:基于覆盖率的直接偏好优化用于 Verilog 刺激生成

Bardia Nadimi, Khashayar Filom, Deming Chen, Hao Zheng

机构 * Core AI Department(核心人工智能部门) Cognichip Inc(认知芯片公司)

专题命中 后训练与偏好优化 :preference optimization(title,abstract);LLM(abstract);large language model(abstract);language model(abstract)

AI总结 TB or not TB通过覆盖率驱动的直接偏好优化,提升Verilog刺激生成的覆盖率和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15661 2025-11-21 cs.CV cs.AI cs.CL cs.LG 82%

VisPlay: Self-Evolving Vision-Language Models from Images

VisPlay: 从图像中自我进化视觉-语言模型

Yicheng He, Chengsong Huang, Zongxia Li, Jiaxin Huang, Yonghui Yang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Washington University in St. Louis(华盛顿大学圣路易斯分校) University of Maryland(马里兰大学) National University of Singapore(新加坡国立大学)

专题命中 后训练与偏好优化 :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 VisPlay通过自我进化强化学习框架,利用未标注图像数据提升视觉-语言模型的推理能力,实现多模态智能的可扩展发展。

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 长上下文与记忆 2 篇

2503.16356 2025-11-21 cs.CL cs.AI cs.CV cs.IR cs.LG 75%

CaKE: Circuit-aware Editing Enables Generalizable Knowledge Learners

CaKE:电路感知编辑实现通用知识学习

Yunzhi Yao, Jizhan Fang, Jia-Chen Gu, Ningyu Zhang, Shumin Deng, Huajun Chen, Nanyun Peng

机构 * Zhejiang University(浙江大学) National University of Singapore(新加坡国立大学) University of California, Los Angeles(美国加州大学洛杉矶分校)

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 CaKE通过电路感知编辑提升LLMs对更新知识的多跳推理能力,实现20%的准确率提升并降低内存消耗

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07067 2025-11-21 cs.LG cs.DC 70%

MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems

MoE-CAP:稀疏专家混合系统成本、准确率和性能的基准测试

Yinsicheng Jiang, Yao Fu, Yeqi Huang, Ping Nie, Zhan Lu, Leyang Xue, Congjie He, Man-Kit Sit, Jilong Xue, Li Dong, Ziming Miao, Dayou Du, Tairan Xu, Kai Zou, Edoardo Ponti, Luo Mai

机构 * University of Edinburgh(爱丁堡大学) Microsoft Research(微软研究院) Peking University(北京大学) NVIDIA Co-leading authors(NVIDIA)

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.LG

AI总结 MoE-CAP通过引入CAP雷达图和稀疏性能指标,系统性地评估稀疏混合专家系统在成本、准确率和性能之间的权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 推理与问题求解 23 篇

2511.14813 2025-11-21 cs.LG 89%

DEVAL: A Framework for Evaluating and Improving the Derivation Capability of Large Language Models

DEVAL:一个用于评估和提升大语言模型推导能力的框架

Yifan Li, Qin Li, Min Zhang, Min Zhang

机构 * East China Normal University(华东师范大学)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);prompting(abstract);分类 cs.LG

AI总结 DEVAL框架通过评估和提升大语言模型的推导能力,提出了一种新的提示工程方法DP,显著提高了模型在问题解决中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏