arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-02-04 至 2026-02-04 共收录 18 信号源:cs.CL, cs.AI, cs.LG

1. 领域大模型 18 篇

2602.00041 2026-02-04 cs.CY cs.AI cs.HC 88%

Student Perceptions of Large Language Models Use in Self-Reflection and Design Critique in Architecture Studio

学生对在建筑工作室中使用大型语言模型进行自我反思和设计批评的看法

Juan David Salazar Rodriguez, Sam Conrad Joyce, Nachamma Sockalingam, Khoo Eng Tat, Julfendi

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 本研究探讨了大型语言模型在建筑工作室自我反思与设计批评中的应用,发现学生将其视为协作工具,帮助构建批判性思维并提升设计迭代效率。

Comments Keywords: Architectural Education, Design Studio Pedagogy, Large Lan-guage Models, Generative AI in Education, Design Critique

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17873 2026-02-04 cs.CV cs.AI 88%

SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model

SurgVidLM:迈向多粒度外科视频理解的大型语言模型

Guankun Wang, Junyi Wang, Wenjin Mo, Long Bai, Kun Yuan, Ming Hu, Jinlin Wu, Junjun He, Yiming Huang, Nicolas Padoy, Zhen Lei, Hongbin Liu, Nassir Navab, Hongliang Ren

机构 * The Chinese University of Hong Kong(香港中文大学) Sun Yat-sen University(中山大学) University of Strasbourg(斯特拉斯堡大学) Technical University of Munich(慕尼黑技术大学) Monash University(墨尔本大学) Centre for Artificial Intelligence and Robotics, HKISI-CAS(人工智能与机器人中心,HKISI-CAS) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 SurgVidLM通过多粒度分析提升外科视频理解能力,结合全局与局部机制实现更精确的手术流程解析。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02752 2026-02-04 cs.SE 87%

Beyond the Prompt: Assessing Domain Knowledge Strategies for High-Dimensional LLM Optimization in Software Engineering

超越提示:评估领域知识策略在软件工程高维LLM优化中的应用

Srinath Srinivasan, Tim Menzies

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本研究比较了人类与人工智能策略在生成领域知识方面的效果,通过四种不同架构评估如何利用结构化知识整合提升LLM在高维优化中的表现。

Comments Accepted at MSR 2026 (Registered Reports Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19120 2026-02-04 cs.IR cs.AI cs.LG 86%

RobustExplain: Evaluating Robustness of LLM-Based Explanation Agents for Recommendation

RobustExplain: 评估基于LLM的推荐解释代理的鲁棒性

Guilin Zhang, Kai Zhao, Jeffrey Friedman, Xu Chu

机构 * Workday

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 RobustExplain提出首个评估框架,用于衡量LLM生成推荐解释的鲁棒性,揭示当前模型鲁棒性较低,较大模型稳定性更高。

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03439 2026-02-04 cs.AI cs.IR 85%

Ontology-to-tools compilation for executable semantic constraint enforcement in LLM agents

本体到工具的编译用于LLM代理中的可执行语义约束强制

Xiaochi Zhou, Patrick Bulter, Changxuan Yang, Simon D. Rihm, Thitikarn Angkanaporn, Jethro Akroyd, Sebastian Mosbach, Markus Kraft

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本体到工具的编译用于LLM代理中的可执行语义约束强制,通过生成知识图谱实例来强制执行语义约束,减少手动工程。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02509 2026-02-04 cs.CY cs.AI 85%

CodeGuard: Improving LLM Guardrails in CS Education

CodeGuard: 提高计算机科学教育中的LLM防护机制

Nishat Raihan, Noah Erdachew, Jayoti Devi, Joanna C. S. Santos, Marcos Zampieri

机构 * George Mason University(乔治·马歇尔大学) University of Oklahoma(俄克拉荷马大学) University of Notre Dame(圣约翰斯大学)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 CodeGuard通过引入新的分类法、数据集和PromptShield模型,有效提升教育AI系统对不安全提示的检测能力,减少有害代码生成,同时保持教育任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03400 2026-02-04 cs.SE cs.AI 81%

Precision in Practice: Knowledge Guided Code Summarizing Grounded in Industrial Expectations

实践中的精确性:基于工业期望的知识引导代码摘要

Jintai Li, Songqiang Chen, Shuo Jin, Xiaoyuan Xie

机构 * School of Computer Science, Wuhan University, China(武汉大学计算机科学学院) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology, China(香港科学与技术大学计算机科学与工程系) Department of Computer Science(计算机科学系)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 ExpSum 通过整合元数据抽象和领域知识检索,生成符合工业期望的代码摘要,显著提升摘要质量与实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22975 2026-02-04 cs.AI 81%

Golden Goose: A Simple Trick to Synthesize Unlimited RLVR Tasks from Unverifiable Internet Text

金 Goose:从不可验证的互联网文本中合成无限 RLVR 任务的简单技巧

Ximing Lu, David Acuna, Jaehun Jung, Jian Hu, Di Zhang, Shizhe Diao, Yunheng Zou, Shaokun Zhang, Brandon Cui, Mingjie Liu, Hyunwoo Kim, Prithviraj Ammanabrolu, Jan Kautz, Yi Dong, Yejin Choi

机构 * NVIDIA(NVIDIA公司) University of Washington(华盛顿大学) University of California San Diego(圣地亚哥大学)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);post-training(abstract)

AI总结 Golden Goose 通过从不可验证的互联网文本中合成无限 RLVR 任务,提升大型语言模型在复杂推理和网络安全领域的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02824 2026-02-04 cs.CL 79%

CATNIP: LLM Unlearning via Calibrated and Tokenized Negative Preference Alignment

CATNIP:通过校准和令牌化的负偏好对齐实现LLM去学习

Zhengbang Yang, Yisheng Zhong, Junyuan Hong, Zhuangdi Zhu

机构 * George Mason University(乔治·马歇尔大学) University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 领域大模型 :LLM(title,abstract);分类 cs.CL

AI总结 CATNIP通过校准和令牌化的负偏好对齐方法,实现有效LLM去学习,无需保留数据或对比对,提升知识遗忘与保留的平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03569 2026-02-04 cs.AI cs.LG 79%

EHRWorld: A Patient-Centric Medical World Model for Long-Horizon Clinical Trajectories

EHRWorld: 一个以患者为中心的医疗世界模型用于长周期临床轨迹

Linjie Mu, Zhongzhen Huang, Yannian Gu, Shengqian Qin, Shaoting Zhang, Xiaofan Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 EHRWorld通过因果序列范式训练,有效解决长周期临床模拟中的误差累积问题,提升医疗世界模型的稳定性与可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03320 2026-02-04 cs.CV cs.AI 77%

MedSAM-Agent: Empowering Interactive Medical Image Segmentation with Multi-turn Agentic Reinforcement Learning

MedSAM-Agent: 通过多轮代理强化学习赋能交互式医学图像分割

Shengyuan Liu, Liuxin Bao, Qi Yang, Wanting Geng, Boyun Zheng, Chenxin Li, Wenting Chen, Houwen Peng, Yixuan Yuan

机构 * Chinese University of Hong Kong, Hong Kong SAR, China(香港中文大学) Hunyuan Group, Tencent(腾讯洪音集团) Institute of Automation, the Chinese Academy of Sciences, Beijing, China(中国科学院自动化研究所) Dalian University of Technology, Dalian, China(大连理工大学) Stanford University, Stanford, USA(斯坦福大学)

专题命中 领域大模型 :large language model(abstract);language model(abstract);prompting(abstract);分类 cs.AI

AI总结 MedSAM-Agent通过多轮代理强化学习提升医学图像分割的交互效率与准确性

Comments 23 Pages, 4 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02919 2026-02-04 cs.AI cs.LG 73%

DeltaEvolve: Accelerating Scientific Discovery through Momentum-Driven Evolution

DeltaEvolve:通过动量驱动的进化加速科学发现

Jiachen Jiang, Tianyu Ding, Zhihui Zhu

专题命中 领域大模型 :LLM(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 DeltaEvolve通过动量驱动的语义delta进化框架,在减少令牌消耗的同时提升科学发现效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02731 2026-02-04 cs.CL cs.AI 73%

Predicting first-episode homelessness among US Veterans using longitudinal EHR data: time-varying models and social risk factors

利用纵向电子健康记录数据预测美国退伍军人首次住房问题:时间变化模型和社会风险因素

Rohan Pandey, Haijuan Yan, Hong Yu, Jack Tsai

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本研究利用纵向电子健康记录数据,结合时间变化模型和社会风险因素,提高了预测退伍军人首次住房问题的准确性,展示了数据驱动的预防策略潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02584 2026-02-04 cs.SE cs.AI cs.CR 70%

Constitutional Spec-Driven Development: Enforcing Security by Construction in AI-Assisted Code Generation

宪法式规范驱动开发:通过构建确保AI辅助代码生成中的安全

Srinivas Rao Marri

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本研究提出宪法规范驱动开发方法,通过在规范层嵌入不可协商的安全原则,有效减少AI生成代码中的安全缺陷,提升开发过程中的安全性。

Comments 15 pages, 2 figures, 5 tables, 11 code listings, 14 references. Includes reference implementation and compliance traceability matrix

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15113 2026-02-04 cs.CL 70%

Inferring Scientific Cross-Document Coreference and Hierarchy with Definition-Augmented Relational Reasoning

基于定义增强的 relational 推理进行科学跨文档指代与层次推断

Lior Forer, Tom Hope

机构 * The Hebrew University of Jerusalem(希伯来大学)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出基于定义增强的relational推理方法,用于科学文本中的跨文档指代和层次推断,通过生成上下文依赖的概念定义和关系定义,提升模型对复杂科学概念的推理能力。

Comments Accepted to TACL. Pre-MIT Press publication version

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03069 2026-02-04 cs.DB 67%

Skill-Based Autonomous Agents for Material Creep Database Construction

基于技能的自主代理用于材料蠕变数据库构建

Yue Wu, Tianhao Su, Shunbo Hu, Deng Pan

专题命中 领域大模型 :large language model(abstract);language model(abstract)

AI总结 本文提出基于技能的自主代理框架,用于从科学文献中自动提取高保真的材料蠕变数据,实现物理自洽的数据库构建。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01757 2026-02-04 cs.CL cs.LG 62%

Zero2Text: Zero-Training Cross-Domain Inversion Attacks on Textual Embeddings

Zero2Text: 无训练跨域反向攻击文本嵌入

Doohyun Kim, Donghwa Kang, Kyungjae Lee, Hyeongboo Baek, Brent Byunghoon Kang

机构 * School of Computing, Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院计算学院) Graduate School of Information Security, Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院信息安全研究生院) Department of Artificial Intelligence, University of Seoul, Seoul, Republic of Korea(首尔大学人工智能系) Department of Computer Science, University of Seoul, Seoul, Republic of Korea(首尔大学计算机科学系)

专题命中 领域大模型 :LLM(abstract);分类 cs.CL、cs.LG

AI总结 Zero2Text通过递归在线对齐机制,在无训练情况下实现跨域文本嵌入反向攻击,有效对抗传统防御手段,实验显示其在恢复句子方面表现优异。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19121 2026-02-04 cs.IR cs.AI cs.LG cs.MA 62%

LLMs as Orchestrators: Constraint-Compliant Multi-Agent Optimization for Recommendation Systems

LLMs作为协调者:用于推荐系统的约束合规多智能体优化

Guilin Zhang, Kai Zhao, Jeffrey Friedman, Xu Chu

机构 * Workday

专题命中 领域大模型 :LLM(abstract);分类 cs.AI、cs.LG

AI总结 本文提出DualAgent-Rec框架,利用LLM协调多智能体优化,实现推荐系统中100%的约束满足和提升帕累托超体积。

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏