arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-01-12 至 2026-01-12 共收录 158 信号源:cs.CL, cs.AI, cs.LG

1. 领域大模型 17 篇

2601.05821 2026-01-12 cs.CL 81%

LLMs as Science Journalists: Supporting Early-stage Researchers in Communicating Their Science to the Public

大语言模型作为科学记者:帮助初级研究人员向公众沟通其科学发现

Milad Alshomary, Grace Li, Anubhav Jangra, Yufang Hou, Kathleen McKeown, Smaranda Muresan

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本研究提出训练大语言模型作为科学记者的框架,以帮助初级研究人员更有效地向公众传达科学发现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05567 2026-01-12 cs.AI cs.CL 79%

WildSci: Advancing Scientific Reasoning from In-the-Wild Literature

WildSci: 从真实文献中推进科学推理

Tengxiao Liu, Deepak Nathani, Zekun Li, Kevin Yang, William Yang Wang

机构 * University of California, Santa Barbara(加州大学圣巴巴拉分校)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 WildSci通过自动合成领域特定科学问题数据集,提升科学推理的可扩展性和可持续性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00742 2026-01-12 physics.comp-ph 75%

Materials Informatics: Emergence To Autonomous Discovery In The Age Of AI

材料信息学:在人工智能时代的涌现与自主发现

Turab Lookman, YuJie Liu, Zhibin Gao

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract)

AI总结 本文探讨了材料信息学在人工智能时代的发展,从基础理论到AI驱动的自主发现,强调其作为不断演进的生态系统,通过反向设计和自主实验室推动材料科学的变革。

Comments 44 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17802 2026-01-12 cs.LG 74%

LEKA:LLM-Enhanced Knowledge Augmentation

LEKA:大语言模型增强的知识增强

Xinhao Zhang, Jinghan Zhang, Fengran Mo, Dongjie Wang, Yanjie Fu, Kunpeng Liu

机构 * Portland State University(波特兰州立大学) University of Montreal(蒙特利尔大学) University of Kansas(堪萨斯大学) Arizona State University(亚利桑那州立大学)

专题命中 领域大模型 :LLM(title);分类 cs.LG

AI总结 LEKA通过主动检索合适知识源,提升跨领域知识转移效率,减少计算成本并优化迁移学习效果。

Comments Accepted by IJCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18695 2026-01-12 cs.AI cs.CE cs.CL 73%

KALE-LM-Chem: Vision and Practice Toward an AI Brain for Chemistry

KALE-LM-Chem:迈向化学人工智能脑的愿景与实践

Weichen Dai, Yezeng Chen, Zijie Dai, Yubo Liu, Zhijie Huang, Yixuan Pan, Baiyang Song, Chengli Zhong, Xinhe Li, Zeyu Wang, Zhuoying Feng, Yi Zhou

机构 * University of Science and Technology of China(科学技术大学) ShanghaiTech University(上海科技大学) USTC Knowledge Computing Lab(USTC知识计算实验室) State Key Laboratory of Communication Content Cognition People's Daily Online(通信内容认知国家重点实验室)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出KALE-LM-Chem模型,旨在通过整合领域知识和逻辑,推动化学领域的智能AI发展,提升科学发现效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05609 2026-01-12 cs.CL 70%

Data Augmented Pipeline for Legal Information Extraction and Reasoning

法律信息抽取与推理的数据增强管道

Nguyen Minh Phuong, Ha-Thanh Nguyen, May Myo Zin, Ken Satoh

机构 * Center for Juris-Informatics, ROIS-DS(司法信息中心,ROIS-DS) Japan Advanced Institute of Science and Technology(日本先进科学研究所) Research and Development Center for LLMs, NII(大语言模型研发中心,NII)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出了一种基于大型语言模型的数据增强管道,用于提升法律领域信息抽取与推理的效率和鲁棒性,同时具备跨领域应用的通用性。

Comments Accepted in the Demonstration Track at ICAIL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05502 2026-01-12 cs.SE cs.AI 70%

Evaluating the Use of LLMs for Automated DOM-Level Resolution of Web Performance Issues

评估LLMs在自动化网页性能问题DOM层面解决中的应用

Gideon Peters, SayedHassan Khatoonabadi, Emad Shihab

机构 * Concordia University(Concordia大学)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本研究评估了九种先进LLMs在自动化网页性能问题DOM层面解决中的效果,发现其在SEO和可访问性问题上表现优异,但在性能关键的DOM修改中效果不一。

Comments Accepted to the The ACM International Conference on Mining Software Repositories (MSR) (MSR 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05399 2026-01-12 cs.CV cs.AI cs.IR 57%

Multi-task Cross-modal Learning for Chest X-ray Image Retrieval

多任务跨模态学习用于胸部X射线图像检索

Zhaohui Liang, Sivaramakrishnan Rajaraman, Niccolo Marini, Zhiyun Xue, Sameer Antani

专题命中 领域大模型 :foundation model(abstract);分类 cs.AI

AI总结 本文提出多任务学习框架,通过改进BiomedCLIP模型,提升胸部X射线图像与文本的跨模态检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05264 2026-01-12 cs.IR cs.AI 57%

Engineering the RAG Stack: A Comprehensive Review of the Architecture and Trust Frameworks for Retrieval-Augmented Generation Systems

构建RAG堆栈:对检索增强生成系统架构和信任框架的全面综述

Dean Wampler, Dave Nielson, Alireza Seddighi

机构 * The AI Alliance(AI联盟) IBM Research(IBM研究院)

专题命中 领域大模型 :LLM(abstract);分类 cs.AI

AI总结 本文综述了RAG系统架构和信任框架,提出统一分类学和评估框架,为构建安全且领域适应性强的RAG系统提供指导。

Comments 86 pages, 2 figures, 37 tables. A comprehensive review of Retrieval-Augmented Generation (RAG) architectures and trust frameworks (2018-2025), encompassing a unified taxonomy, evaluation benchmarks, and trust-safety modeling

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05887 2026-01-12 cs.CR 50%

Cybersecurity AI: A Game-Theoretic AI for Guiding Attack and Defense

网络空间AI:一种基于博弈论的AI用于引导攻击与防御

Víctor Mayoral-Vilches, María Sanz-Gómez, Francesco Balassone, Stefan Rass, Lidia Salas-Espejo, Benjamin Jablonski, Luis Javier Navarrete-Lozano, Maite del Mundo de Torres, Cristóbal R. J. Veas Chavez

专题命中 领域大模型 :LLM(abstract)

AI总结 本文提出G-CTR,一种基于博弈论的AI指导层,通过生成摘要引导攻击与防御行为,提升网络安全测试效率与成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05255 2026-01-12 cs.IR cs.HC 50%

CourtNav: Voice-Guided, Anchor-Accurate Navigation of Long Legal Documents in Courtrooms

CourtNav: 语音引导、锚点精确的法庭长法律文档导航

Sai Khadloya, Kush Juvekar, Arghya Bhattacharya, Utkarsh Saxena

专题命中 领域大模型 :LLM(abstract)

AI总结 CourtNav通过语音引导和锚点导航技术,显著提升法庭长法律文档的检索效率,将相关性查找时间从分钟级降至秒级。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 知识编辑与模型理解 12 篇

2406.03505 2026-01-12 cs.LG cs.AI 84%

Dynamic and Adaptive Feature Generation with LLM

动态和自适应特征生成与大语言模型

Xinhao Zhang, Jinghan Zhang, Banafsheh Rekabdar, Yuanchun Zhou, Pengfei Wang, Kunpeng Liu

机构 * Portland State University(波特兰州立大学) Computer Network Information Center, Chinese Academy of Sciences(中国科学院计算机网络信息中心) University of Chinese Academy of Sciences, Chinese Academy of Sciences(中国科学院大学)

专题命中 知识编辑与模型理解 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出利用大语言模型和特征生成提示,实现动态和自适应的特征生成方法,以提高特征生成的可解释性、适用性和灵活性。

Comments Accepted by IJCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05298 2026-01-12 cs.AI 77%

Mathematical Knowledge Graph-Driven Framework for Equation-Based Predictive and Reliable Additive Manufacturing

基于数学知识图谱的方程驱动框架用于基于方程的预测和可靠的增材制造

Yeongbin Cha, Namjung Kim

专题命中 知识编辑与模型理解 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本研究提出基于数学知识图谱的框架,结合大语言模型与增材制造知识图谱,实现可靠的知识提取和外推建模,提升预测的准确性和物理一致性。

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05772 2026-01-12 cs.SE cs.CR 75%

StriderSPD: Structure-Guided Joint Representation Learning for Binary Security Patch Detection

StriderSPD:基于结构的二进制安全补丁检测联合表示学习

Qingyuan Li, Chenchen Yu, Chuanyi Li, Xin-Cheng Wen, Cheryl Lee, Cuiyun Gao, Bin Luo

专题命中 知识编辑与模型理解 :LLM(abstract);large language model(abstract);language model(abstract)

AI总结 StriderSPD通过整合图分支与大型语言模型,利用结构信息提升二进制安全补丁检测的准确性与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05254 2026-01-12 cs.CV cs.AI cs.LG 73%

Explaining Low Perception Model Competency with High-Competency Counterfactuals

用高能力反事实解释低感知模型能力

Sara Pohland, Claire Tomlin

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出五种生成高能力反事实图像的方法,用于解释模型预测不确定性的原因,并通过实验验证了其在生成语言解释中的有效性。

Journal ref Explainable Artificial Intelligence. xAI 2025. Communications in Computer and Information Science, vol 2580

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06240 2026-01-12 cs.CL 70%

Bridging External and Parametric Knowledge: Mitigating Hallucination of LLMs with Shared-Private Semantic Synergy in Dual-Stream Knowledge

弥合外部与参数化知识:通过共享-私有语义协同缓解LLM的幻觉

Yi Sui, Chaozhuo Li, Chen Zhang, Dawei song, Qiuchi Li

机构 * Beijing Institute of Technology, China(北京理工大学) Beijing University of Posts and Telecommunications, China(北京邮电大学) Meituan, China(美团) The Open University, UK(开放大学)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出DSSP-RAG框架,通过共享-私有语义协同机制,缓解LLM在整合外部知识时的幻觉问题,并提升生成性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05495 2026-01-12 cs.CV cs.CL 70%

MMViR: A Multi-Modal and Multi-Granularity Representation for Long-range Video Understanding

MMViR:一种多模态和多粒度表示用于长视频理解

Zizhong Li, Haopeng Zhang, Jiawei Zhang

机构 * IFM Lab, University of California, Davis(加州大学戴维斯分校信息融合实验室) ALOHA Lab, University of Hawaii at Mānoa(夏威夷大学马诺阿分校ALOHA实验室)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 MMViR通过多模态和多粒度表示提升长视频理解,实现更高效的检索和更优的性能表现。

Comments 13 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14367 2026-01-12 cs.CV 67%

Hallucination Score: Towards Mitigating Hallucinations in Generative Image Super-Resolution

幻觉分数:朝着减轻生成图像超分辨率中的幻觉

Weiming Ren, Raghav Goyal, Zhiming Hu, Tristan Ty Aumentado-Armstrong, Iqbal Mohomed, Alex Levinshtein

机构 * University of Waterloo(多伦多大学) AI Center – Toronto, Samsung Electronics(三星电子人工智能中心)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract)

AI总结 本文提出幻觉分数以衡量和减轻生成图像超分辨率中的幻觉问题,通过多模态大语言模型生成 HS 并用于模型微调。

Comments 31 pages, 21 figures, and 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16996 2026-01-12 econ.TH cs.GT 67%

Artificial Intelligence Clones

人工智能克隆

Annie Liang

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract)

AI总结 本文研究了人工智能克隆在匹配搜索中的影响,发现有限的面对面接触比无限的人工智能表示搜索更有效,尤其是在高维人格空间中。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25475 2026-01-12 cs.AI cs.LG 62%

TDHook: A Lightweight Framework for Interpretability

TDHook:一个轻量级的可解释性框架

Yoann Poupart

机构 * LIP6(LIP6实验室) Sorbonne University(索邦大学)

专题命中 知识编辑与模型理解 :language model(abstract);分类 cs.AI、cs.LG

AI总结 TDHook是一个轻量级的可解释性框架,旨在通过处理复杂组合模型和提供灵活的API,提升现代可解释性流水线的易用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05647 2026-01-12 cs.LG cs.AI 62%

Transformer Is Inherently a Causal Learner

Transformer本质上是一种因果学习者

Xinyue Wang, Stephen Wang, Biwei Huang

机构 * Halıcıoğlu Data Science Institute(哈利奇奥格鲁数据科学研究所) University of California San Diego(加州大学圣地亚哥分校) ABEL Intelligence, Inc.(ABEL智能公司)

专题命中 知识编辑与模型理解 :foundation model(abstract);分类 cs.AI、cs.LG

AI总结 本文证明Transformer通过自回归训练自然编码因果结构,无需显式因果目标即可恢复因果图,且在复杂场景中表现优于传统方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05939 2026-01-12 cs.CV 50%

Context-Aware Decoding for Faithful Vision-Language Generation

面向上下文的解码方法用于忠实的视觉-语言生成

Mehrdad Fazli, Bowen Wei, Ziwei Zhu

机构 * Department of Computer Science, George Mason University(计算机科学系,乔治·马歇尔大学)

专题命中 知识编辑与模型理解 :language model(abstract)

AI总结 本文提出了一种无需训练的上下文嵌入注入方法,通过利用上下文嵌入信号在解码过程中保持视觉一致性,有效减少视觉-语言模型中的幻觉问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05721 2026-01-12 cs.SE 50%

From Issues to Insights: RAG-based Explanation Generation from Software Engineering Artifacts

从问题到洞察:基于RAG的软件工程制品解释生成

Daniel Pöttgen, Mersedeh Sadeghi, Max Unterbusch, Andreas Vogelsang

专题命中 知识编辑与模型理解 :language model(abstract)

AI总结 本文首次应用RAG方法从问题跟踪数据生成解释,实现90%的人工解释一致性,展示了结构化问题数据在提升软件系统可解释性方面的潜力。

Comments Accepted at NLBSE 2026, Rio de Janeiro, Brazil

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他LLM 15 篇

2511.02108 2026-01-12 cs.SE cs.AI 89%

Metamorphic Testing of Large Language Models for Natural Language Processing

大型语言模型在自然语言处理中的变形测试

Steven Cho, Stefano Ruberto, Valerio Terragni

机构 * University of Auckland(奥克兰大学)

专题命中 其他LLM :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 本文提出了一种针对大型语言模型的变形测试方法,通过收集和实验191个变形关系,评估了变形测试在自然语言处理任务中的有效性与局限性。

Journal ref Proc. 2025 IEEE Int. Conf. on Software Maintenance and Evolution (ICSME), pp. 174-186, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05542 2026-01-12 cs.SE cs.AI cs.LG 89%

Understanding LLM-Driven Test Oracle Generation

理解由大语言模型驱动的测试 oracle 生成

Adam Bodicoat, Gunel Jahangirova, Valerio Terragni

机构 * University of Auckland(奥克兰大学) King's College London(伦敦国王学院)

专题命中 其他LLM :LLM(title,abstract);large language model(abstract);language model(abstract);foundation model(abstract)

AI总结 本文研究了大语言模型在生成测试 oracle 以暴露软件故障中的有效性,探讨了提示策略和上下文输入对 oracle 质量的影响。

Comments Accepted for presentation at the 2nd ACM/IEEE International Conference on AI-powered Software (AIware 2025)

Journal ref Proc. 2nd ACM/IEEE International Conference on AI-powered Software (AIware 2025), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05903 2026-01-12 cs.CL 85%

HAPS: Hierarchical LLM Routing with Joint Architecture and Parameter Search

HAPS: 基于联合架构和参数搜索的分层LLM路由

Zihang Tian, Rui Li, Jingsen Zhang, Xiaohe Bo, Wei Huo, Xu Chen

机构 * Renmin University of China(中国人民大学) Wireless Technology Lab, Huawei Technologies Co., Ltd.(华为技术有限公司无线技术实验室)

专题命中 其他LLM :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 HAPS通过联合搜索模型架构和参数,提升大型语言模型在多样化任务中的路由性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05835 2026-01-12 cs.CL 85%

Left, Right, or Center? Evaluating LLM Framing in News Classification and Generation

左、右或中间?评估LLM在新闻分类和生成中的框架

Molly Kennedy, Ali Parker, Yihong Liu, Hinrich Schütze

专题命中 其他LLM :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 研究评估了LLM在新闻分类和生成中的框架倾向,发现存在系统性的中间框架倾向,Grok 4表现最意识形态,Claude Sonnet 4.5和Llama 3.1在商业和开放权重模型中表现最佳。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16035 2026-01-12 cs.CL cs.AI 84%

Liars' Bench: Evaluating Lie Detectors for Language Models

说谎的检验台:评估语言模型的说谎检测器

Kieron Kretschmar, Walter Laurito, Sharan Maiya, Samuel Marks

机构 * Cadenza Labs(Cadenza实验室) FZI(弗劳恩霍夫研究所) University of Cambridge(剑桥大学)

专题命中 其他LLM :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出LIARS' BENCH,通过多个数据集和模型生成72,863个谎言和诚实响应,评估三种谎言检测技术,揭示现有方法在识别特定类型谎言上的局限性。

Comments *Kieron Kretschmar and Walter Laurito contributed equally to this work. 10 pages, 2 figures; plus appendix. Code at https://github.com/Cadenza-Labs/liars-bench and datasets at https://huggingface.co/datasets/Cadenza-Labs/liars-bench Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21434 2026-01-12 hep-ph cs.AI cs.LG hep-ex physics.data-an 81%

Foundation models for high-energy physics

高能物理中的基础模型

Anna Hallin

机构 * Institute for Experimental Physics, University of Hamburg(汉堡大学实验物理研究所)

专题命中 其他LLM :foundation model(title,abstract);分类 cs.AI、cs.LG

AI总结 本文综述了高能物理中基础模型的应用现状,探讨了其在粒子物理数据中的潜在应用与研究进展。

Comments Submitted to SciPost Physics Proceedings (EuCAIFCon 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05579 2026-01-12 cs.DB cs.AI cs.CL cs.SE 79%

RISE: Rule-Driven SQL Dialect Translation via Query Reduction

RISE: 通过查询简化实现规则驱动的SQL方言翻译

Xudong Xie, Yuwei Zhang, Wensheng Dou, Yu Gao, Ziyu Cui, Jiansen Song, Rui Yang, Jun Wei

机构 * Institute of Software at CAS, China(中国科学院软件研究所)

专题命中 其他LLM :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 RISE通过查询简化技术,利用LLM实现高效准确的SQL方言翻译,显著提升翻译准确率。

Comments Accepted by ICSE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏