arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-02-11 至 2026-02-11 共收录 200 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 39 篇

2512.08493 2026-02-11 cs.CR cs.AI 79%

LLM-based Vulnerable Code Augmentation: Generate or Refactor?

基于LLM的易受攻击代码增强:生成或重构?

Dyna Soumhane Ouchebara, Stéphane Dupont

机构 * University of Mons - Computer science department(蒙斯大学-计算机科学系)

专题命中 评测与基准 :LLM(title,abstract);分类 cs.AI

AI总结 本文提出基于LLM的漏洞代码增强方法,通过生成或重构技术提升漏洞分类器性能,实验表明混合策略效果最佳。

Comments 15 pages, Accepted by ESAAN 2026, version with added appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22516 2026-02-11 cs.LG cs.CV 79%

Ice-FMBench: A Foundation Model Benchmark for Sea Ice Type Segmentation

Ice-FMBench: 一个用于海冰类型分割的基础模型基准

Samira Alkaee Taleghan, Morteza Karimzadeh, Andrew P. Barrett, Walter N. Meier, Farnoush Banaei-Kashani

机构 * University of Colorado Denver(科罗拉多大学丹佛分校) University of Colorado Boulder(科罗拉多大学波德分校) National Snow and Ice Data Center (NSIDC), CIRES, University of Colorado Boulder(国家冰雪数据研究中心(NSIDC)、CIRES、科罗拉多大学波德分校)

专题命中 评测与基准 :foundation model(title,abstract);分类 cs.LG

AI总结 Ice-FMBench为海冰类型分割任务提供了一个基准框架,评估基础模型的性能,并通过多教师知识蒸馏方法提升模型的时空可转移性。

Journal ref ACM ACM SIGSPATIAL PoIDS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09444 2026-02-11 cs.CL cs.AI 79%

Conceptual Cultural Index: A Metric for Cultural Specificity via Relative Generality

概念文化指数:一种通过相对普遍性衡量文化特异性的指标

Takumi Ohashi, Hitoshi Iyatomi

机构 * Hosei University(恒生大学)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出概念文化指数(CCI)用于衡量句子层面的文化特异性,通过比较不同文化间的普遍性估计,有效评估文化特异性并提升二元可分性性能。

Comments 9 pages, 2 figures, 8 tables. Accepted at the First Workshop on Multilingual Multicultural Evaluation (MME) @ EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04337 2026-02-11 cs.CL cs.AI cs.HC cs.IR 79%

Modelling and Classifying the Components of a Literature Review

文献综述组件的建模与分类

Francisco Bolaños, Angelo Salatino, Francesco Osborne, Enrico Motta

机构 * Knowledge Media Institute, The Open University(开放大学知识媒体研究所) The Open University(开放大学) Department of Business and Law, University of Milano Bicocca(米兰Bicocca大学商学院)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出了一种新的注释方案和评估方法,用于对文献综述中的句子进行修辞角色分类,并评估了多种大型语言模型的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09214 2026-02-11 cs.CV 78%

VLM-UQBench: A Benchmark for Modality-Specific and Cross-Modality Uncertainties in Vision Language Models

VLM-UQBench:一种用于视觉语言模型中模态特定和跨模态不确定性的基准

Chenyu Wang, Tianle Chen, H. M. Sabbir Ahmad, Kayhan Batmanghelich, Wenchao Li

机构 * Boston University(波士顿大学)

专题命中 评测与基准 :language model(title,abstract)

AI总结 VLM-UQBench通过评估不同UQ方法在模态特定和跨模态不确定性上的表现,揭示了现有方法在细粒度不确定性检测上的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08268 2026-02-11 cs.AI 77%

Puda: Private User Dataset Agent for User-Sovereign and Privacy-Preserving Personalized AI

Puda:用户主权与隐私保护的个性化AI用户数据代理

Akinori Maeda, Yuto Sekiya, Sota Sugimura, Tomoya Asai, Yu Tsuda, Kohei Ikeda, Hiroshi Fujii, Kohei Watanabe

机构 * Research Institute of Advanced Technology, SoftBank Corp.(软银公司先进科技研究所) Turnt Up Technologies, Inc.(Turnt Up技术公司) TechArts Co., Ltd.(TechArts公司) Acutus Software, Inc.(Acutus软件公司)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 Puda通过用户主权架构实现多粒度数据管理,平衡隐私保护与个性化AI的使用需求。

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10913 2026-02-11 cs.CV cs.CL 77%

Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP

了解“不”:一种数据驱动的方法用于增强CLIP中的否定意识

Junsung Park, Jungbeom Lee, Jongyoon Song, Sangwon Yu, Dahuin Jung, Sungroh Yoon

机构 * Department of Electrical and Computer Engineering, Seoul National University(首尔国立大学电子与计算机工程系) Amazon(亚马逊) Samsung Research(三星研究院) School of Computer Science and Engineering, Soongsil University(顺天大学计算机科学与工程学院) IPAI, AIIS, ASRI, INMC, and ISRC, Seoul National University(首尔国立大学IPAI、AIIS、ASRI、INMC和ISRC)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出NegationCLIP,通过生成包含否定的数据增强CLIP的否定意识,同时提出NegRefCOCOg基准用于评估多模态模型的否定理解能力。

Comments Accepted to ICCV 2025

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025, pp. 2825-2835

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19379 2026-02-11 cs.LG cs.AI cs.MM 76%

OmniMER: Auxiliary-Enhanced LLM Adaptation for Indonesian Multimodal Emotion Recognition

OmniMER: 增辅增强的LLM适应用于印度尼西亚多模态情感识别

Xueming Yan, Boyan Xu, Yaochu Jin, Lixian Xiao, Wenlong Ye, Runyang Cai, Zeqi Zheng, Jingfa Liu, Aimin Yang, Yongduan Song

机构 * School of Information Science and Technology, Guangdong University of Foreign Studies(广东外语外贸大学信息科学与技术学院) School of Computer Science, Guangdong University of Technology(广东工业大学计算机学院) Faculty of Asian Languages and Cultures, Guangdong University of Foreign Studies(广东外语外贸大学亚洲语言文化学院) School of Engineering, Westlake University(西湖大学工程学院) School of Computer Science and Intelligence Education, Lingnan Normal University(岭南师范学院计算机科学与智能教育学院) School of Automation, Chongqing University(重庆大学自动化学院)

专题命中 评测与基准 :LLM(title);分类 cs.AI、cs.LG

AI总结 OmniMER通过三种辅助任务提升印度尼西亚多模态情感识别性能,实现情感分类和识别的显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09930 2026-02-11 cs.SE 75%

JMigBench: A Benchmark for Evaluating LLMs on Source Code Migration (Java 8 to Java 11)

JMigBench: 一个用于评估大语言模型在源代码迁移(Java 8到Java 11)任务的基准

Nishil Amin, Zhiwei Fei, Xiang Li, Justyna Petke, He Ye

专题命中 评测与基准 :large language model(abstract);language model(abstract);prompting(abstract)

AI总结 JMigBench基准评估了大语言模型在Java 8到Java 11源代码迁移任务中的表现,发现其在简单API替换上有效,但对复杂迁移任务仍存在不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18075 2026-02-11 cs.SE cs.ET 75%

A Context-Driven Approach for Co-Auditing Smart Contracts with The Support of GPT-4 code interpreter

基于GPT-4代码解释器的上下文驱动智能合约共审计方法

Mohamed Salah Bouafif, Chen Zheng, Ilham Ahmed Qasse, Ed Zulkoski, Mohammad Hamdaqa, Foutse Khomh

专题命中 评测与基准 :large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文提出一种基于GPT-4代码解释器的上下文驱动智能合约共审计方法,通过优化提示设计提升漏洞检测效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09079 2026-02-11 cs.LG 74%

Patient foundation model for risk stratification in low-risk overweight patients

患者风险分层的基础模型用于低风险超重患者

Zachary N. Flamholz, Dillon Tracy, Ripple Khera, Jordan Wolinsky, Nicholas Lee, Nathaniel Tann, Xiao Yin Zhu, Harry Phillips, Jeffrey Sherman

机构 * Zephyr AI, Inc.(Zephyr AI公司)

专题命中 评测与基准 :foundation model(title);分类 cs.LG

AI总结 PatientTPP通过整合临床知识和时间序列数据,为低风险超重患者的风险分层提供可解释的通用模型,有效提升心血管相关医疗成本的预测能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09346 2026-02-11 cs.CL 70%

Digital Linguistic Bias in Spanish: Evidence from Lexical Variation in LLMs

西班牙的数字语言偏见:来自LLMs词形变异的证据

Yoshifumi Kawasaki

机构 * University of Tokyo(东京大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本研究通过评估LLMs对西班牙语地域词形变异的识别能力,揭示了模型在不同方言表现上的系统性差异,并指出数据量之外的因素影响方言表征。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24654 2026-02-11 cs.CL 70%

Evolving Interactive Diagnostic Agents in a Virtual Clinical Environment

在虚拟临床环境中进化交互诊断代理

Pengcheng Qiu, Chaoyi Wu, Junwei Liu, Qiaoyu Zheng, Yusheng Liao, Haowen Wang, Yun Yue, Qianrui Fan, Shuai Zhen, Jian Wang, Jinjie Gu, Yanfeng Wang, Ya Zhang, Weidi Xie

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本研究提出DiagAgent,通过强化学习在虚拟临床环境中训练,实现多轮交互式诊断,显著提升诊断准确性和检查推荐效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16596 2026-02-11 cs.CV cs.AI 70%

SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense

SHIELD:通过偏差和脆弱性防御抑制LVLM编码器中的幻觉

Yiyang Huang, Liang Shi, Yitian Zhang, Yi Xu, Yun Fu

机构 * Department of Electrical and Computer Engineering, Northeastern University(电气与计算机工程系,东北大学) Khoury College of Computer Science, Northeastern University(计算机科学学院,东北大学)

专题命中 评测与基准 :LLM(abstract);language model(abstract);分类 cs.AI

AI总结 SHIELD通过减少统计偏差、对抗固有偏差和解决脆弱性,有效抑制LVLM编码器中的对象幻觉。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15935 2026-02-11 cs.DB cs.CL cs.CR 70%

MAPS: A Multilingual Benchmark for Agent Performance and Security

MAPS: 一种多语言评估基准,用于代理性能和安全

Omer Hofman, Jonathan Brokman, Oren Rachmil, Shamik Bose, Vikas Pahuja, Toshiya Shimizu, Trisha Starostina, Kelly Marchisio, Seraphina Goldfarb-Tarrant, Roman Vainshtein

机构 * Fujitsu Research of Europe(富士通欧洲研究部) Fujitsu Limited(富士通有限公司) Cohere

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 MAPS是一个多语言评估基准,用于评估代理AI在不同语言和任务中的性能与安全性。

Comments Accepted to EACL 2026 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14708 2026-02-11 cs.AI cs.DC cs.ET cs.HC 70%

Creation of AI-driven Smart Spaces for Enhanced Indoor Environments -- A Survey

构建AI驱动的智能空间以提升室内环境——综述

Aygün Varol, Naser Hossein Motlagh, Mirka Leino, Sasu Tarkoma, Johanna Virkki

机构 * Department of Computing Sciences(计算科学系) Tampere University(塔尔皮奥大学) Department of Computer Science(计算机科学系) University of Helsinki(赫尔辛基大学) Faculty of Technology(技术学院) Satakunta University of Applied Sciences(萨塔昆塔应用科学大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文综述了AI驱动智能空间的基础组件和关键技术,探讨了传统和新兴AI方法在提升室内环境中的应用与挑战,为未来研究提供了方向。

Comments 39 pages, 3 figures, 1 table, journal

Journal ref Internet of Things, Volume 36, 2026, 101876

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09772 2026-02-11 cs.RO 67%

Design and Evaluation of an Assisted Programming Interface for Behavior Trees in Robotics

行为树在机器人中的辅助编程接口设计与评估

Jonathan Styrud, Matteo Iovino, Rebecca Stower, Mart Kartašev, Mikael Norrlöf, Mårten Björkman, Christian Smith

专题命中 评测与基准 :large language model(abstract);language model(abstract)

AI总结 本文提出BETR-GUI,结合AI助手与拖放编辑器,通过整合多种技术提升机器人行为树编程效率,实验证明人类用户优于纯AI助手。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09210 2026-02-11 eess.SP cs.SD eess.AS 67%

AI-Driven Cardiorespiratory Signal Processing: Separation, Clustering, and Anomaly Detection

由AI驱动的心脏呼吸信号处理:分离、聚类和异常检测

Yasaman Torabi

专题命中 评测与基准 :large language model(abstract);language model(abstract)

AI总结 本文提出利用AI技术对心脏呼吸信号进行分离、聚类和异常检测,并探讨了新一代传感器在智能医疗诊断中的应用。

Comments PhD thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00185 2026-02-11 cs.AI cs.LG 62%

Chunking Strategies for Multimodal AI Systems

多模态AI系统中的分块策略

Shashanka B R, Mohith Charan R, Seema Banu F

专题命中 评测与基准 :foundation model(abstract);分类 cs.AI、cs.LG

AI总结 本文综述了多模态系统中分块策略的分类和技术分析,探讨了不同模态的数据处理方法及挑战。

Comments 50 pages, 5 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04072 2026-02-11 cs.AI cs.LG 62%

Among Us: A Sandbox for Measuring and Detecting Agentic Deception

Among Us: 一种用于衡量和检测代理欺骗的沙盒

Satvik Golechha, Adrià Garriga-Alonso

机构 * MATS FAR AI

专题命中 评测与基准 :LLM(abstract);分类 cs.AI、cs.LG

AI总结 Among Us通过开放的沙盒游戏评估LLM的欺骗能力,发现强化学习训练的模型更擅长生成欺骗而非检测,且基于激活的逻辑回归和SAEs在检测中表现优异。

Comments 21 pages, preprint

Journal ref NeurIPS 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04772 2026-02-11 cs.CV cs.AI cs.LG 62%

Federated Learning for Surgical Vision in Appendicitis Classification: Results of the FedSurg EndoVis 2024 Challenge

在阑尾炎分类中用于手术视觉的联邦学习:FedSurg EndoVis 2024挑战的结果

Max Kirchner, Hanna Hoffmann, Alexander C. Jenke, Oliver L. Saldanha, Kevin Pfeiffer, Weam Kanjo, Julia Alekseenko, Claas de Boer, Santhi Raj Kolamuri, Lorenzo Mazza, Nicolas Padoy, Sophia Bano, Annika Reinke, Lena Maier-Hein, Danail Stoyanov, Jakob N. Kather, Fiona R. Kolbinger, Sebastian Bodenstedt, Stefanie Speidel

机构 * Department of Translational Surgical Oncology, National Center for Tumor Diseases (NCT), NCT/UCC Dresden, a partnership between DKFZ, Faculty of Medicine and University Hospital Carl Gustav Carus, TUD Dresden University of Technology(转化外科肿瘤学部,肿瘤疾病国家中心(NCT),NCT/UCC德累斯顿,DKFZ、医学院和卡尔·戈斯瓦尔德·卡尔医院之间的合作) Centre for Tactile Internet with Human-in-the-Loop (CeTI), TUD Dresden University of Technology(人机协同触觉互联网中心(CeTI),德累斯顿技术大学) Department of Medicine I, Faculty of Medicine and University Hospital Carl Gustav Carus, TUD Dresden University of Technology(医学部I,医学院和卡尔·戈斯瓦尔德·卡尔医院,德累斯顿技术大学) Medical Oncology, National Center for Tumor Diseases (NCT), University Hospital Heidelberg(肿瘤医学部,肿瘤疾病国家中心(NCT),海德堡大学医院) Weldon School of Biomedical Engineering, Purdue University(生物医学工程学院,普渡大学) Else Kroener Fresenius Center for Digital Health, Faculty of Medicine and University Hospital Carl Gustav Carus, TUD Dresden University of Technology(Else Kroener Fresenius数字健康中心,医学院和卡尔·戈斯瓦尔德·卡尔医院,德累斯顿技术大学) IHU Strasbourg, Institute of Image-Guided Surgery(斯特拉斯堡IHU,影像引导手术研究所) University of Strasbourg, CNRS, INSERM, ICube, UMR7357(斯特拉斯堡大学,CNRS,INSERM,ICube,UMR7357) Dr. NTR University of Health Sciences(Dr. NTR健康科学大学) Department of Visceral, Thoracic and Vascular Surgery, Faculty of Medicine and University Hospital Carl Gustav Carus, TUD Dresden University of Technology(visceral、胸腔和血管外科部,医学院和卡尔·戈斯瓦尔德·卡尔医院,德累斯顿技术大学) Digital Technologies, Medtronic(数字技术,美敦力) UCL Hawkes Institute and Department of Computer Science, University College London(UCL Hawkes研究所和计算机科学系,伦敦大学学院) German Cancer Research Center (DKFZ) Heidelberg, Division of Intelligent Medical Systems(德国癌症研究中心(DKFZ)海德堡,智能医学系统部) German Cancer Research Center (DKFZ) Heidelberg, Helmholtz Imaging(德国癌症研究中心(DKFZ)海德堡,海德堡成像) National Center for Tumor Diseases (NCT) Heidelberg, a partnership between DKFZ and Heidelberg University Hospital(肿瘤疾病国家中心(NCT)海德堡,DKFZ和海德堡大学医院之间的合作) Faculty of Mathematics and Computer Science, Heidelberg University(数学和计算机科学学院,海德堡大学) Helmholtz Information and Data Science School for Health(海德堡信息与数据科学健康学院) Heidelberg University Hospital, Surgical Clinic, Surgical AI Research Group(海德堡大学医院,外科诊所,外科AI研究组) Krankenhaus St. Joseph-Stift Dresden GmbH(德累斯顿圣约瑟夫修道院医院)

专题命中 评测与基准 :foundation model(abstract);分类 cs.AI、cs.LG

AI总结 FedSurg挑战评估了联邦学习在手术视频分类中的性能,发现局部微调提升效果有限,但ViViT模型表现最佳,突显了架构选择和预处理的重要性。

Comments A challenge report pre-print (31 pages), including 7 tables and 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17684 2026-02-11 cs.CL cs.AI 62%

Can LLMs Automate Fact-Checking Article Writing?

LLMs能否自动化事实核查文章写作?

Dhruv Sahnan, David Corney, Irene Larraz, Giovanni Zagni, Ruben Miguez, Zhuohan Xie, Iryna Gurevych, Elizabeth Churchill, Tanmoy Chakraborty, Preslav Nakov

机构 * MBZUAI, UAE(阿联酋马布里克人工智能研究所) Full Fact, UK(英国Full Fact) Newtral, Spain(西班牙Newtral) Pagella Politica, Italy(意大利Pagella Politica) TU Darmstadt, Germany(德国图尔恩大学) IIT Delhi, India(印度德里印度理工学院)

专题命中 评测与基准 :LLM(abstract);分类 cs.CL、cs.AI

AI总结 本文提出QRAFT框架,通过LLM生成事实核查文章,旨在解决自动化事实核查中缺乏合理解释的问题。

Comments Accepted to TACL 2026, pre-MIT Press publication version

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09339 2026-02-11 cs.CL 57%

Understanding Risk and Dependency in AI Chatbot Use from User Discourse

理解AI聊天机器人使用中的风险与依赖性:从用户话语出发

Jianfeng Zhu, Karin G. Coifman, Ruoming Jin

专题命中 评测与基准 :LLM(abstract);分类 cs.CL

AI总结 本研究通过分析用户话语揭示AI聊天机器人使用中的心理风险维度,发现自我调节困难和对自主性、控制和技术风险的恐惧为主要特征。

Comments 21 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09336 2026-02-11 cs.CL 57%

FM SO.P: A Progressive Task Mixture Framework with Automatic Evaluation for Cross-Domain SOP Understanding

FM SO.P: 一种具有自动评估的渐进式任务混合框架用于跨域SOP理解

Siyuan Huang, Ziyu Wang, Chao Pan, Han Zhao

机构 * Amazon(亚马逊) Johns Hopkins University(约翰霍普金斯大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 评测与基准 :language model(abstract);分类 cs.CL

AI总结 FM SO.P提出了一种渐进式任务混合框架和自动评估系统,以提升跨领域SOP理解的准确性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06734 2026-02-11 cs.CR cs.LG 57%

Quantifying the Generalization Gap: A New Benchmark for Out-of-Distribution Graph-Based Android Malware Classification

量化泛化差距:面向分布外图基Android恶意软件分类的新基准

Ngoc N. Tran, Anwar Said, Waseem Abbas, Tyler Derr, Xenofon D. Koutsoukos

机构 * Vanderbilt University(范德比大学) The University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 评测与基准 :LLM(abstract);分类 cs.LG

AI总结 本研究提出一个新的基准测试套件,用于评估图基Android恶意软件分类在分布偏移下的泛化能力,并通过语义增强框架提升检测鲁棒性。

Comments 14 pages, 5 figures, 10 tables, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10382 2026-02-11 cs.LG 57%

Leveraging RAG-LLMs for Urban Mobility Simulation and Analysis

利用 RAG-LLMs 进行城市交通模拟与分析

Yue Ding, Conor McCarthy, Kevin O'Shea, Mingming Liu

机构 * Centre for Research Training in Machine Learning (ML-Labs)(机器学习研究培训中心(ML实验室)) Dublin City University(都柏林城市大学) Insight Centre for Data Analytics(数据分析洞察中心) School of Electronic Engineering(电子工程学院)

专题命中 评测与基准 :LLM(abstract);分类 cs.LG

AI总结 本文提出基于 RAG-LLMs 的共享电动交通平台,通过个性化路线推荐和 schema-level RAG 框架提升交通模拟与分析效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09531 2026-02-11 cs.CV 50%

DR.Experts: Differential Refinement of Distortion-Aware Experts for Blind Image Quality Assessment

DR.Experts: 基于失真感知专家的差分细化方法用于盲图像质量评估

Bohan Fu, Guanyi Qin, Fazhan Zhang, Zihao Huang, Mingxuan Li, Runze Hu

专题命中 评测与基准 :language model(abstract)

AI总结 DR.Experts通过引入失真感知专家和差分细化方法,有效整合失真先验知识,提升盲图像质量评估的准确性与泛化能力。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09439 2026-02-11 cs.CV 50%

Fine-T2I: An Open, Large-Scale, and Diverse Dataset for High-Quality T2I Fine-Tuning

Fine-T2I: 一个大规模、高质量且多样化的开放数据集用于高质量T2I微调

Xu Ma, Yitian Zhang, Qihua Dong, Yun Fu

机构 * Department of Electrical \& Computer Engineering, Northeastern University, Boston

专题命中 评测与基准 :pretraining(abstract)

AI总结 Fine-T2I是一个大规模高质量开放数据集,通过结合合成和真实图像,提升T2I微调的生成质量和指令遵循性。

Comments Dataset: https://huggingface.co/datasets/ma-xu/fine-t2i

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09411 2026-02-11 cs.CV 50%

K-Sort Eval: Efficient Preference Evaluation for Visual Generation via Corrected VLM-as-a-Judge

K-Sort Eval: 通过校正VLM作为评判者实现视觉生成的高效偏好评估

Zhikai Li, Jiatong Li, Xuewen Liu, Wangbo Zhao, Pan Du, Kaicheng Zhou, Qingyi Gu, Yang You, Zhen Dong, Kurt Keutzer

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of California, Berkeley(加州大学伯克利分校) University of California, Santa Barbara(加州大学圣巴巴拉分校) National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学)

专题命中 评测与基准 :language model(abstract)

AI总结 K-Sort Eval通过后验校正和动态匹配策略,实现高效可靠的视觉生成模型偏好评估。

Comments ICLR 2026. Code is available at: https://github.com/zkkli/K-Sort-Eval

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23429 2026-02-11 cs.CV 50%

Hunyuan-GameCraft-2: Instruction-following Interactive Game World Model

Hunyuan-GameCraft-2:基于指令的交互式游戏世界模型

Junshu Tang, Jiacheng Liu, Jiaqi Li, Longhuang Wu, Haoyu Yang, Penghao Zhao, Siruis Gong, Xiang Yuan, Shuai Shao, Linfeng Zhang, Qinglin Lu

机构 * Tencent Hunyuan(腾讯 Hunyuan)

专题命中 评测与基准 :foundation model(abstract)

AI总结 Hunyuan-GameCraft-2通过自然语言提示等多模态交互方式,实现更灵活的生成游戏世界建模,提升交互性和因果一致性。

Comments Technical Report, Project page:https://hunyuan-gamecraft-2.github.io/, Demo:https://hunyuan.tencent.com/game/game-craft

详情

展开后加载摘要…

URL PDF HTML 收藏