arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-02-11 至 2026-02-11 共收录 39 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 39 篇

2602.09540 2026-02-11 cs.SE 88%

SWE-Bench Mobile: Can Large Language Model Agents Develop Industry-Level Mobile Applications?

SWE-Bench Mobile: 大语言模型代理能否开发行业级移动应用?

Muxin Tian, Zhe Wang, Blair Yang, Zhenwei Tang, Kunlun Zhu, Honghua Dong, Hanchen Li, Xinni Xie, Guangjing Wang, Jiaxuan You

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract)

AI总结 SWE-Bench Mobile评估大语言模型代理在开发行业级移动应用中的能力,发现商业代理表现优于开源替代品,且简单提示策略更有效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02076 2026-02-11 cs.AI cs.MA 87%

Leveraging LLM Agents and Digital Twins for Fault Handling in Process Plants

利用大语言模型代理和数字孪生进行过程厂故障处理

Milapji Singh Gill, Javal Vyas, Artan Markaj, Felix Gehlhoff, Mehmet Mercangöz

机构 * Institute of Automation Technology Helmut Schmidt University Hamburg(海因里希·施密特大学汉堡自动化技术研究所) Autonomous Industrial Systems Lab, Imperial College London(伦敦帝国理工学院自主工业系统实验室)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文提出利用大语言模型代理和数字孪生技术,实现过程厂故障的自动处理与纠正,通过模拟验证有效缓解管道堵塞问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10017 2026-02-11 cs.CL 85%

SCORE: Specificity, Context Utilization, Robustness, and Relevance for Reference-Free LLM Evaluation

SCORE:特定性、上下文利用、鲁棒性与相关性用于无参考LLM评估

Homaira Huda Shomee, Rochana Chaturvedi, Yangxinyu Xie, Tanwi Mallick

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) Argonne National Laboratory(阿贡国家实验室) University of Pennsylvania(宾夕法尼亚大学)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出了一种无参考评估框架,用于评估LLM在高风险领域任务中的特定性、鲁棒性、相关性和上下文利用,通过精心编纂的数据集和人工评估验证了多指标评估的必要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09624 2026-02-11 cs.CL 85%

MILE-RefHumEval: A Reference-Free, Multi-Independent LLM Framework for Human-Aligned Evaluation

MILE-RefHumEval: 一种无需参考的、多独立LLM评估框架用于人类对齐评估

Nalin Srun, Parisa Rastin, Guénaël Cabanes, Lydia Boudjeloud Assala

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 MILE-RefHumEval是一种无需参考的多独立LLM评估框架,通过人类对齐方案和任务特定提示,实现高效、稳健的LLM评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08146 2026-02-11 cs.SE 85%

Test vs Mutant: Adversarial LLM Agents for Robust Unit Test Generation

测试与突变:对抗性LLM代理用于鲁棒性单元测试生成

Pengyu Chang, Yixiong Fang, Silin Chen, Yuling Shi, Beijun Shen, Xiaodong Gu

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract)

AI总结 AdverTest通过对抗性代理提升单元测试生成的鲁棒性,提高故障检测率和覆盖率

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08712 2026-02-11 cs.CL cs.AI cs.DC 84%

A Survey on Parallel Text Generation: From Parallel Decoding to Diffusion Language Models

关于并行文本生成的综述:从并行解码到扩散语言模型

Lingzhe Zhang, Liancheng Fang, Chiming Duan, Minghua He, Leyi Pan, Pei Xiao, Shiyu Huang, Yunpeng Zhai, Xuming Hu, Philip S. Yu, Aiwei Liu

机构 * Peking University(北京大学) University of Illinois Chicago(伊利诺伊大学香槟分校) Tsinghua University(清华大学) XPENG Alibaba Group(阿里巴巴集团) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 评测与基准 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI

AI总结 本文综述了并行文本生成技术,分析了基于AR和非AR的方法,评估了其在速度、质量和效率上的权衡,并指出了未来研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09383 2026-02-11 cs.CL cs.AI cs.SE 81%

BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation

BiasScope: 向LLM-as-a-Judge评估中的偏见检测自动化迈进

Peng Lai, Zhihao Ou, Yong Wang, Longyue Wang, Jian Yang, Yun Chen, Guanhua Chen

机构 * Southern University of Science and Technology(南方科技大学) Alibaba Group(阿里巴巴集团) Beihang University(北航) Shanghai University of Finance and Economics(上海财经大学)

专题命中 评测与基准 :LLM(title,abstract);分类 cs.CL、cs.AI

AI总结 BiasScope通过自动化发现LLM-as-a-Judge评估中的潜在偏见,提升了评估的稳健性,并提出了更具有挑战性的JudgeBench-Pro基准。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09914 2026-02-11 cs.CL cs.IR 79%

AmharicIR+Instr: A Two-Dataset Resource for Neural Retrieval and Instruction Tuning

AmharicIR+Instr: 一种双数据集资源用于神经检索与指令微调

Tilahun Yeshambel, Moncef Garouani, Josiane Mothe

机构 * Computer Science Department, Addis Ababa University(亚的斯亚贝巴大学计算机科学系) Univ. Toulouse Capitole, IRIT, UMR5505 CNRS(图卢兹卡普利大学、IRIT、UMR5505 CNRS) UT2J, Univ. de Toulouse, IRIT, UMR5505 CNRS(UT2J、图卢兹大学、IRIT、UMR5505 CNRS)

专题命中 评测与基准 :instruction tuning(title);LLM(abstract);分类 cs.CL

AI总结 本研究发布了一个包含两个数据集的阿姆哈拉语资源,用于神经检索和指令微调,支持对比训练和生成模型研究。

Comments 7 pages, Submitted to resource track

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09252 2026-02-11 cs.CV cs.AI cs.MA 79%

VLM-Guided Iterative Refinement for Surgical Image Segmentation with Foundation Models

基于基础模型的手术图像分割迭代细化方法

Ange Lou, Yamin Li, Qi Chang, Nan Xi, Luyuan Xie, Zichao Li, Tianyu Luan

机构 * School of Automation Science and Engineering, South China University of Technology, Guangzhou, China(自动化科学与工程学院,华南理工大学,广州,中国) Vanderbilt University, Nashville, TN, USA(范德比尔特大学,纳什维尔,田纳西州,美国) The Pennsylvania State University, University Park, PA , USA(宾夕法尼亚州立大学,大学园,宾夕法尼亚州,美国) Peking University, Beijing , China(北京大学,北京,中国) Virginia Commonwealth University, Richmond, VA, USA(弗吉尼亚共同wealth大学,里士满,弗吉尼亚州,美国) University of California San Diego, San Diego, CA, USA(加州大学圣地亚哥分校,圣地亚哥,加利福尼亚州,美国)

专题命中 评测与基准 :foundation model(title);language model(abstract);分类 cs.AI

AI总结 本文提出IR-SIS,一种基于自然语言描述的手术图像分割迭代细化系统,结合视觉-语言模型和自适应工作流,实现高精度分割与临床交互。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08493 2026-02-11 cs.CR cs.AI 79%

LLM-based Vulnerable Code Augmentation: Generate or Refactor?

基于LLM的易受攻击代码增强:生成或重构?

Dyna Soumhane Ouchebara, Stéphane Dupont

机构 * University of Mons - Computer science department(蒙斯大学-计算机科学系)

专题命中 评测与基准 :LLM(title,abstract);分类 cs.AI

AI总结 本文提出基于LLM的漏洞代码增强方法,通过生成或重构技术提升漏洞分类器性能,实验表明混合策略效果最佳。

Comments 15 pages, Accepted by ESAAN 2026, version with added appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22516 2026-02-11 cs.LG cs.CV 79%

Ice-FMBench: A Foundation Model Benchmark for Sea Ice Type Segmentation

Ice-FMBench: 一个用于海冰类型分割的基础模型基准

Samira Alkaee Taleghan, Morteza Karimzadeh, Andrew P. Barrett, Walter N. Meier, Farnoush Banaei-Kashani

机构 * University of Colorado Denver(科罗拉多大学丹佛分校) University of Colorado Boulder(科罗拉多大学波德分校) National Snow and Ice Data Center (NSIDC), CIRES, University of Colorado Boulder(国家冰雪数据研究中心(NSIDC)、CIRES、科罗拉多大学波德分校)

专题命中 评测与基准 :foundation model(title,abstract);分类 cs.LG

AI总结 Ice-FMBench为海冰类型分割任务提供了一个基准框架,评估基础模型的性能,并通过多教师知识蒸馏方法提升模型的时空可转移性。

Journal ref ACM ACM SIGSPATIAL PoIDS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09444 2026-02-11 cs.CL cs.AI 79%

Conceptual Cultural Index: A Metric for Cultural Specificity via Relative Generality

概念文化指数:一种通过相对普遍性衡量文化特异性的指标

Takumi Ohashi, Hitoshi Iyatomi

机构 * Hosei University(恒生大学)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出概念文化指数(CCI)用于衡量句子层面的文化特异性,通过比较不同文化间的普遍性估计,有效评估文化特异性并提升二元可分性性能。

Comments 9 pages, 2 figures, 8 tables. Accepted at the First Workshop on Multilingual Multicultural Evaluation (MME) @ EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04337 2026-02-11 cs.CL cs.AI cs.HC cs.IR 79%

Modelling and Classifying the Components of a Literature Review

文献综述组件的建模与分类

Francisco Bolaños, Angelo Salatino, Francesco Osborne, Enrico Motta

机构 * Knowledge Media Institute, The Open University(开放大学知识媒体研究所) The Open University(开放大学) Department of Business and Law, University of Milano Bicocca(米兰Bicocca大学商学院)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出了一种新的注释方案和评估方法,用于对文献综述中的句子进行修辞角色分类,并评估了多种大型语言模型的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09214 2026-02-11 cs.CV 78%

VLM-UQBench: A Benchmark for Modality-Specific and Cross-Modality Uncertainties in Vision Language Models

VLM-UQBench:一种用于视觉语言模型中模态特定和跨模态不确定性的基准

Chenyu Wang, Tianle Chen, H. M. Sabbir Ahmad, Kayhan Batmanghelich, Wenchao Li

机构 * Boston University(波士顿大学)

专题命中 评测与基准 :language model(title,abstract)

AI总结 VLM-UQBench通过评估不同UQ方法在模态特定和跨模态不确定性上的表现,揭示了现有方法在细粒度不确定性检测上的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08268 2026-02-11 cs.AI 77%

Puda: Private User Dataset Agent for User-Sovereign and Privacy-Preserving Personalized AI

Puda:用户主权与隐私保护的个性化AI用户数据代理

Akinori Maeda, Yuto Sekiya, Sota Sugimura, Tomoya Asai, Yu Tsuda, Kohei Ikeda, Hiroshi Fujii, Kohei Watanabe

机构 * Research Institute of Advanced Technology, SoftBank Corp.(软银公司先进科技研究所) Turnt Up Technologies, Inc.(Turnt Up技术公司) TechArts Co., Ltd.(TechArts公司) Acutus Software, Inc.(Acutus软件公司)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 Puda通过用户主权架构实现多粒度数据管理,平衡隐私保护与个性化AI的使用需求。

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10913 2026-02-11 cs.CV cs.CL 77%

Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP

了解“不”:一种数据驱动的方法用于增强CLIP中的否定意识

Junsung Park, Jungbeom Lee, Jongyoon Song, Sangwon Yu, Dahuin Jung, Sungroh Yoon

机构 * Department of Electrical and Computer Engineering, Seoul National University(首尔国立大学电子与计算机工程系) Amazon(亚马逊) Samsung Research(三星研究院) School of Computer Science and Engineering, Soongsil University(顺天大学计算机科学与工程学院) IPAI, AIIS, ASRI, INMC, and ISRC, Seoul National University(首尔国立大学IPAI、AIIS、ASRI、INMC和ISRC)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出NegationCLIP,通过生成包含否定的数据增强CLIP的否定意识,同时提出NegRefCOCOg基准用于评估多模态模型的否定理解能力。

Comments Accepted to ICCV 2025

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025, pp. 2825-2835

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19379 2026-02-11 cs.LG cs.AI cs.MM 76%

OmniMER: Auxiliary-Enhanced LLM Adaptation for Indonesian Multimodal Emotion Recognition

OmniMER: 增辅增强的LLM适应用于印度尼西亚多模态情感识别

Xueming Yan, Boyan Xu, Yaochu Jin, Lixian Xiao, Wenlong Ye, Runyang Cai, Zeqi Zheng, Jingfa Liu, Aimin Yang, Yongduan Song

机构 * School of Information Science and Technology, Guangdong University of Foreign Studies(广东外语外贸大学信息科学与技术学院) School of Computer Science, Guangdong University of Technology(广东工业大学计算机学院) Faculty of Asian Languages and Cultures, Guangdong University of Foreign Studies(广东外语外贸大学亚洲语言文化学院) School of Engineering, Westlake University(西湖大学工程学院) School of Computer Science and Intelligence Education, Lingnan Normal University(岭南师范学院计算机科学与智能教育学院) School of Automation, Chongqing University(重庆大学自动化学院)

专题命中 评测与基准 :LLM(title);分类 cs.AI、cs.LG

AI总结 OmniMER通过三种辅助任务提升印度尼西亚多模态情感识别性能,实现情感分类和识别的显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09930 2026-02-11 cs.SE 75%

JMigBench: A Benchmark for Evaluating LLMs on Source Code Migration (Java 8 to Java 11)

JMigBench: 一个用于评估大语言模型在源代码迁移(Java 8到Java 11)任务的基准

Nishil Amin, Zhiwei Fei, Xiang Li, Justyna Petke, He Ye

专题命中 评测与基准 :large language model(abstract);language model(abstract);prompting(abstract)

AI总结 JMigBench基准评估了大语言模型在Java 8到Java 11源代码迁移任务中的表现,发现其在简单API替换上有效,但对复杂迁移任务仍存在不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18075 2026-02-11 cs.SE cs.ET 75%

A Context-Driven Approach for Co-Auditing Smart Contracts with The Support of GPT-4 code interpreter

基于GPT-4代码解释器的上下文驱动智能合约共审计方法

Mohamed Salah Bouafif, Chen Zheng, Ilham Ahmed Qasse, Ed Zulkoski, Mohammad Hamdaqa, Foutse Khomh

专题命中 评测与基准 :large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文提出一种基于GPT-4代码解释器的上下文驱动智能合约共审计方法,通过优化提示设计提升漏洞检测效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09079 2026-02-11 cs.LG 74%

Patient foundation model for risk stratification in low-risk overweight patients

患者风险分层的基础模型用于低风险超重患者

Zachary N. Flamholz, Dillon Tracy, Ripple Khera, Jordan Wolinsky, Nicholas Lee, Nathaniel Tann, Xiao Yin Zhu, Harry Phillips, Jeffrey Sherman

机构 * Zephyr AI, Inc.(Zephyr AI公司)

专题命中 评测与基准 :foundation model(title);分类 cs.LG

AI总结 PatientTPP通过整合临床知识和时间序列数据,为低风险超重患者的风险分层提供可解释的通用模型,有效提升心血管相关医疗成本的预测能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09346 2026-02-11 cs.CL 70%

Digital Linguistic Bias in Spanish: Evidence from Lexical Variation in LLMs

西班牙的数字语言偏见:来自LLMs词形变异的证据

Yoshifumi Kawasaki

机构 * University of Tokyo(东京大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本研究通过评估LLMs对西班牙语地域词形变异的识别能力,揭示了模型在不同方言表现上的系统性差异,并指出数据量之外的因素影响方言表征。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24654 2026-02-11 cs.CL 70%

Evolving Interactive Diagnostic Agents in a Virtual Clinical Environment

在虚拟临床环境中进化交互诊断代理

Pengcheng Qiu, Chaoyi Wu, Junwei Liu, Qiaoyu Zheng, Yusheng Liao, Haowen Wang, Yun Yue, Qianrui Fan, Shuai Zhen, Jian Wang, Jinjie Gu, Yanfeng Wang, Ya Zhang, Weidi Xie

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本研究提出DiagAgent,通过强化学习在虚拟临床环境中训练,实现多轮交互式诊断,显著提升诊断准确性和检查推荐效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16596 2026-02-11 cs.CV cs.AI 70%

SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense

SHIELD:通过偏差和脆弱性防御抑制LVLM编码器中的幻觉

Yiyang Huang, Liang Shi, Yitian Zhang, Yi Xu, Yun Fu

机构 * Department of Electrical and Computer Engineering, Northeastern University(电气与计算机工程系,东北大学) Khoury College of Computer Science, Northeastern University(计算机科学学院,东北大学)

专题命中 评测与基准 :LLM(abstract);language model(abstract);分类 cs.AI

AI总结 SHIELD通过减少统计偏差、对抗固有偏差和解决脆弱性,有效抑制LVLM编码器中的对象幻觉。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15935 2026-02-11 cs.DB cs.CL cs.CR 70%

MAPS: A Multilingual Benchmark for Agent Performance and Security

MAPS: 一种多语言评估基准,用于代理性能和安全

Omer Hofman, Jonathan Brokman, Oren Rachmil, Shamik Bose, Vikas Pahuja, Toshiya Shimizu, Trisha Starostina, Kelly Marchisio, Seraphina Goldfarb-Tarrant, Roman Vainshtein

机构 * Fujitsu Research of Europe(富士通欧洲研究部) Fujitsu Limited(富士通有限公司) Cohere

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 MAPS是一个多语言评估基准,用于评估代理AI在不同语言和任务中的性能与安全性。

Comments Accepted to EACL 2026 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14708 2026-02-11 cs.AI cs.DC cs.ET cs.HC 70%

Creation of AI-driven Smart Spaces for Enhanced Indoor Environments -- A Survey

构建AI驱动的智能空间以提升室内环境——综述

Aygün Varol, Naser Hossein Motlagh, Mirka Leino, Sasu Tarkoma, Johanna Virkki

机构 * Department of Computing Sciences(计算科学系) Tampere University(塔尔皮奥大学) Department of Computer Science(计算机科学系) University of Helsinki(赫尔辛基大学) Faculty of Technology(技术学院) Satakunta University of Applied Sciences(萨塔昆塔应用科学大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文综述了AI驱动智能空间的基础组件和关键技术,探讨了传统和新兴AI方法在提升室内环境中的应用与挑战,为未来研究提供了方向。

Comments 39 pages, 3 figures, 1 table, journal

Journal ref Internet of Things, Volume 36, 2026, 101876

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09772 2026-02-11 cs.RO 67%

Design and Evaluation of an Assisted Programming Interface for Behavior Trees in Robotics

行为树在机器人中的辅助编程接口设计与评估

Jonathan Styrud, Matteo Iovino, Rebecca Stower, Mart Kartašev, Mikael Norrlöf, Mårten Björkman, Christian Smith

专题命中 评测与基准 :large language model(abstract);language model(abstract)

AI总结 本文提出BETR-GUI,结合AI助手与拖放编辑器,通过整合多种技术提升机器人行为树编程效率,实验证明人类用户优于纯AI助手。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09210 2026-02-11 eess.SP cs.SD eess.AS 67%

AI-Driven Cardiorespiratory Signal Processing: Separation, Clustering, and Anomaly Detection

由AI驱动的心脏呼吸信号处理:分离、聚类和异常检测

Yasaman Torabi

专题命中 评测与基准 :large language model(abstract);language model(abstract)

AI总结 本文提出利用AI技术对心脏呼吸信号进行分离、聚类和异常检测,并探讨了新一代传感器在智能医疗诊断中的应用。

Comments PhD thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00185 2026-02-11 cs.AI cs.LG 62%

Chunking Strategies for Multimodal AI Systems

多模态AI系统中的分块策略

Shashanka B R, Mohith Charan R, Seema Banu F

专题命中 评测与基准 :foundation model(abstract);分类 cs.AI、cs.LG

AI总结 本文综述了多模态系统中分块策略的分类和技术分析,探讨了不同模态的数据处理方法及挑战。

Comments 50 pages, 5 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04072 2026-02-11 cs.AI cs.LG 62%

Among Us: A Sandbox for Measuring and Detecting Agentic Deception

Among Us: 一种用于衡量和检测代理欺骗的沙盒

Satvik Golechha, Adrià Garriga-Alonso

机构 * MATS FAR AI

专题命中 评测与基准 :LLM(abstract);分类 cs.AI、cs.LG

AI总结 Among Us通过开放的沙盒游戏评估LLM的欺骗能力,发现强化学习训练的模型更擅长生成欺骗而非检测,且基于激活的逻辑回归和SAEs在检测中表现优异。

Comments 21 pages, preprint

Journal ref NeurIPS 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04772 2026-02-11 cs.CV cs.AI cs.LG 62%

Federated Learning for Surgical Vision in Appendicitis Classification: Results of the FedSurg EndoVis 2024 Challenge

在阑尾炎分类中用于手术视觉的联邦学习:FedSurg EndoVis 2024挑战的结果

Max Kirchner, Hanna Hoffmann, Alexander C. Jenke, Oliver L. Saldanha, Kevin Pfeiffer, Weam Kanjo, Julia Alekseenko, Claas de Boer, Santhi Raj Kolamuri, Lorenzo Mazza, Nicolas Padoy, Sophia Bano, Annika Reinke, Lena Maier-Hein, Danail Stoyanov, Jakob N. Kather, Fiona R. Kolbinger, Sebastian Bodenstedt, Stefanie Speidel

机构 * Department of Translational Surgical Oncology, National Center for Tumor Diseases (NCT), NCT/UCC Dresden, a partnership between DKFZ, Faculty of Medicine and University Hospital Carl Gustav Carus, TUD Dresden University of Technology(转化外科肿瘤学部,肿瘤疾病国家中心(NCT),NCT/UCC德累斯顿,DKFZ、医学院和卡尔·戈斯瓦尔德·卡尔医院之间的合作) Centre for Tactile Internet with Human-in-the-Loop (CeTI), TUD Dresden University of Technology(人机协同触觉互联网中心(CeTI),德累斯顿技术大学) Department of Medicine I, Faculty of Medicine and University Hospital Carl Gustav Carus, TUD Dresden University of Technology(医学部I,医学院和卡尔·戈斯瓦尔德·卡尔医院,德累斯顿技术大学) Medical Oncology, National Center for Tumor Diseases (NCT), University Hospital Heidelberg(肿瘤医学部,肿瘤疾病国家中心(NCT),海德堡大学医院) Weldon School of Biomedical Engineering, Purdue University(生物医学工程学院,普渡大学) Else Kroener Fresenius Center for Digital Health, Faculty of Medicine and University Hospital Carl Gustav Carus, TUD Dresden University of Technology(Else Kroener Fresenius数字健康中心,医学院和卡尔·戈斯瓦尔德·卡尔医院,德累斯顿技术大学) IHU Strasbourg, Institute of Image-Guided Surgery(斯特拉斯堡IHU,影像引导手术研究所) University of Strasbourg, CNRS, INSERM, ICube, UMR7357(斯特拉斯堡大学,CNRS,INSERM,ICube,UMR7357) Dr. NTR University of Health Sciences(Dr. NTR健康科学大学) Department of Visceral, Thoracic and Vascular Surgery, Faculty of Medicine and University Hospital Carl Gustav Carus, TUD Dresden University of Technology(visceral、胸腔和血管外科部,医学院和卡尔·戈斯瓦尔德·卡尔医院,德累斯顿技术大学) Digital Technologies, Medtronic(数字技术,美敦力) UCL Hawkes Institute and Department of Computer Science, University College London(UCL Hawkes研究所和计算机科学系,伦敦大学学院) German Cancer Research Center (DKFZ) Heidelberg, Division of Intelligent Medical Systems(德国癌症研究中心(DKFZ)海德堡,智能医学系统部) German Cancer Research Center (DKFZ) Heidelberg, Helmholtz Imaging(德国癌症研究中心(DKFZ)海德堡,海德堡成像) National Center for Tumor Diseases (NCT) Heidelberg, a partnership between DKFZ and Heidelberg University Hospital(肿瘤疾病国家中心(NCT)海德堡,DKFZ和海德堡大学医院之间的合作) Faculty of Mathematics and Computer Science, Heidelberg University(数学和计算机科学学院,海德堡大学) Helmholtz Information and Data Science School for Health(海德堡信息与数据科学健康学院) Heidelberg University Hospital, Surgical Clinic, Surgical AI Research Group(海德堡大学医院,外科诊所,外科AI研究组) Krankenhaus St. Joseph-Stift Dresden GmbH(德累斯顿圣约瑟夫修道院医院)

专题命中 评测与基准 :foundation model(abstract);分类 cs.AI、cs.LG

AI总结 FedSurg挑战评估了联邦学习在手术视频分类中的性能,发现局部微调提升效果有限,但ViViT模型表现最佳,突显了架构选择和预处理的重要性。

Comments A challenge report pre-print (31 pages), including 7 tables and 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏