arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 1448 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 1448 篇

2604.05738 2026-06-25 cs.CL 版本更新 79%

MedLayBench-V: A Large-Scale Benchmark for Expert-Lay Semantic Alignment in Medical Vision Language Models

MedLayBench-V:面向医学视觉语言模型中专家与普通人语义对齐的大规模基准

Han Jang, Junhyeok Lee, Heeseong Eum, Kyu Sung Choi

机构 * Seoul National University(首尔国立大学) Seoul National University College of Medicine(首尔国立大学医学院) Department of Radiology, Seoul National University Hospital(首尔国立大学医院放射科) Healthcare AI Research Institute, Seoul National University Hospital(首尔国立大学医院健康人工智能研究所) The Advanced Imaging and Computational Neuroimaging (AICON) Laboratory(先进影像与计算神经影像实验室)

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL

AI总结 提出首个大规模多模态基准MedLayBench-V,通过结构化概念基础精炼管道实现专家-普通人语义对齐,用于训练和评估能弥合医患沟通鸿沟的医学视觉语言模型。

Comments Findings of ACL 2026. 9 pages, 5 figures, 11 tables, plus appendix

Journal ref Findings of the Association for Computational Linguistics: ACL 2026, pages 18375-18394

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04584 2026-06-25 cs.CL cs.SD eess.AS 版本更新 79%

Robustness assessment of large audio language models in multiple-choice evaluation

大型音频语言模型在多项选择评估中的鲁棒性评估

Fernando López, Santosh Kesiraju, Jordi Luque

机构 * Scientific Research, Telefónica Innovación Digital, Spain(Telefónica Innovación Digital科研部,西班牙) Universidad Autónoma de Madrid, Spain(马德里自治大学,西班牙) Brno University of Technology, Czech Republic(布拉格技术大学,捷克)

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL

AI总结 研究大型音频语言模型在多项选择问答中对选项顺序、问题措辞的敏感性,并提出考虑细微变化的评估协议和指标。

Comments Accepted in Interspeech 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17742 2026-06-16 eess.SP cs.AI cs.HC 版本更新 79%

EEG-FM-Bench: A Comprehensive Benchmark for the Systematic Evaluation and Diagnostic Analyses of EEG Foundation Models

EEG-FM-Bench:脑电图基础模型系统评估与诊断分析的综合基准

Wei Xiong, Jiangtong Li, Jie Li, Kun Zhu, Changjun Jiang

机构 * School of Computer Science and Technology, Tongji University, Shanghai, China(同济大学计算机科学与技术学院,上海,中国) Translational Research Center, Shanghai Yangzhi Rehabilitation Hospital (Shanghai Sunshine Rehabilitation Center), China(上海杨氏康复医院(上海阳光康复中心)转化研究中心,中国)

专题命中 评测与基准 :foundation model(title,abstract);分类 cs.AI

AI总结 提出EEG-FM-Bench统一基准,整合14个数据集和10种范式,通过多种微调策略和诊断分析揭示多任务学习可缓解过拟合、预训练效率受梯度冲突限制、模型规模非唯一决定因素等关键发现。

Comments 36 pages, 30 figures, Accepted by ICML2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17588 2026-06-16 cs.CV cs.CL 版本更新 79%

Dual-branch Prompting for Multimodal Machine Translation

双分支提示用于多模态机器翻译

Jie Wang, Zhendong Yang, Liansong Zong, Xiaobo Zhang, Dexian Wang, Ji Zhang

机构 * School of Computing and Artificial Intelligence, Southwest Jiaotong University(西南交通大学计算机与人工智能学院) School of Computer and Software Engineering, Xihua University(西华大学计算机与软件工程学院) School of Intelligent Medicine, Chengdu University of Traditional Chinese Medicine(成都中医药大学针灸推拿学院)

专题命中 评测与基准 :prompting(title,abstract);分类 cs.CL

AI总结 提出基于扩散模型的双分支提示框架D2P-MMT,利用重建图像过滤视觉噪声,通过分布对齐损失提升鲁棒翻译性能。

Comments This manuscript has been fully accepted and published by ACM Transactions on Multimedia Computing, Communications, and Applications (ACM TOMM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15411 2026-08-21 cs.CL cs.AI 版本更新 79%

HiFi-KPI: A Dataset for Hierarchical KPI Extraction from Earnings Filings

HiFi-KPI:用于从财报中提取层次化KPI的数据集

Rasmus T. Aavang, Giovanni Rizzi, Rasmus Tjalk-Bøggild, Alexandre Iolov, Mike Zhang, Johannes Bjerva

机构 * Department of Computer Science, Aalborg University(奥尔堡大学计算机科学系) ALIPES ApS(ALIPES公司) University of Copenhagen(哥本哈根大学) Pioneer Centre for AI(先锋人工智能中心)

专题命中 评测与基准 :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 针对财报中关键绩效指标(KPI)跨公司可迁移性差的问题,提出包含165万段落和19.8万层次化标签的HiFi-KPI数据集,并评估分类、提取和结构化提取三个任务。

Comments Camera-ready. Accepted at LREC 2026 (main conference)

Journal ref Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026), pages 441-455, Palma, Mallorca, Spain. ELRA, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.21071 2026-08-17 cs.CL cs.AI 版本更新 79%

Fine-grained Claim-level RAG Benchmark for Law

细粒度声明级法律RAG基准

Souvick Das, Sallam Abualhaija, Domenico Bianculli

机构 * University of Luxembourg(卢森堡大学)

专题命中 评测与基准 :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 提出一个支持法语和英语的细粒度声明级法律RAG数据集ClaimRAG-LAW,用于评估检索和生成性能,揭示法律领域RAG系统的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11616 2026-08-14 cs.AI cs.CV cs.LG 版本更新 79%

MBA: Multimodal Benchmark and Agents for Real-World Business Ideation

MBA:面向现实世界商业创意的多模态基准与智能体

Hojun Choi, Jaeyo Shin, Suin Lee, Hyunjung Shim

专题命中 评测与基准 :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 该研究推出首个多模态商业创意基准MBA-Bench,提出MBA-b和MBA-k两种智能体,经实验其性能显著优于相关基准,为多模态商业创意智能体研究提供了重要支撑。

Comments Project page: https://hchoi256.github.io/projects/mba/ Code: https://github.com/hchoi256/MBA

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20295 2026-08-10 cs.LG cs.AI cs.CV 版本更新 79%

Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription

通过封面判断书籍:调查多模态大语言模型用于多页手写文档转录

Benjamin Gutteridge, Matthew Thomas Jackson, Toni Kukurin, Xiaowen Dong

机构 * University of Oxford(牛津大学) QuantCo

专题命中 评测与基准 :LLM(abstract,abstract_cn);prompting(abstract);分类 cs.AI、cs.LG

AI总结 本文研究多模态大语言模型在多页手写文档转录中的应用,提出OCR+PAGE-1和OCR+PAGE-N策略,通过共享页面内容提升转录效果。

Comments 10 pages (36 including references and appendices), 11 figures, accepted at COLM 2026, earlier version accepted at AAAI 2025 Workshop on Document Understanding and Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14695 2026-08-07 cs.LG cs.CL 版本更新 79%

Persona-Pruner: Sculpting Lightweight Models for Role-Playing

Persona-Pruner: 为角色扮演雕琢轻量级模型

Jinsu Kim, Jihoon Tack, Noah Lee, Jongheon Jeong

机构 * Department of Artificial Intelligence, Korea University, Seoul, South Korea(韩国大学人工智能系) Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)

专题命中 评测与基准 :LLM(abstract,abstract_cn);language model(abstract);分类 cs.CL、cs.LG

AI总结 提出Persona-Pruner框架,通过从单个描述中隔离特定角色的子网络来剪枝语言模型,在保持角色扮演性能的同时大幅降低计算成本,性能下降比最强基线减少93.8%。

Comments 25 pages; ICML 2026; Code is available at https://github.com/jsu-kim/Persona-Pruner

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21692 2026-08-06 cs.CL cs.AI 版本更新 79%

Revisiting Generalization Across Difficulty Levels: It's Not So Easy

重新审视不同难度层级间的泛化:这并不容易

Yeganeh Kordi, Nihal V. Nayak, Max Zuo, Ilana Nguyen, Stephen H. Bach

机构 * Brown University(布朗大学) Harvard University(哈佛大学)

专题命中 评测与基准 :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文研究了LLMs在不同任务难度间泛化的能力,发现训练数据的难度对泛化效果影响有限,强调在训练和评估中需涵盖多种难度以避免风险。

Comments Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19827 2026-08-03 cs.CL cs.AI cs.IR 版本更新 79%

When Iterative RAG Beats Ideal Evidence: A Diagnostic Study in Scientific Multi-hop Question Answering

当迭代RAG优于理想证据:科学多跳问答中的诊断研究

Mahdi Astaraki, Mohammad Arshi Saloot, Ali Shiraee Kasmaee, Hamidreza Mahyar, Soheila Samiee

机构 * Faculty of Engineering, McMaster University, Canada(麦斯特大学工程学院,加拿大) BASF Canada Inc., Canada(巴斯夫加拿大公司,加拿大)

专题命中 评测与基准 :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 通过化学多跳问答数据集,诊断发现迭代检索-推理循环在科学领域显著优于静态RAG上限,揭示了阶段式检索的优势与失败模式。

Comments 51 pages, 29 figures, Published in Transactions on Machine Learning Research (05/2026). OpenReview: https://openreview.net/forum?id=pa5TnBdyDP

Journal ref Transactions on Machine Learning Research (05/2026), ISSN 2835-8856

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01153 2026-07-30 cs.CL cs.AI cs.SE 版本更新 79%

Adversarial Pragmatics for AI Safety Evaluation: A Diagnostic Framework and Seed Benchmark for Language-Mediated Control

面向AI安全评估的对抗语用学:指令冲突、嵌入命令与策略模糊性基准

Brett Reynolds

机构 * Humber Polytechnic(汉博理工学院) University of Toronto(多伦多大学)

专题命中 评测与基准 :LLM(abstract,abstract_cn);language model(abstract);分类 cs.CL、cs.AI

AI总结 提出对抗语用学基准和标注协议,通过语言学控制的分类法评估模型在指令冲突、嵌入命令等场景下的行为,为安全评估提供实证和方法论工具。

Comments 32-page main paper plus 13-page supplement; 6 figures and 17 tables total; code and data artifact available at the linked repository

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05710 2026-07-30 physics.ao-ph cs.AI cs.LG 版本更新 79%

The Rise of AI in Weather and Climate Information and its Impact on Global Inequality

人工智能在天气和气候信息中的兴起及其对全球不平等的影响

Amirpasha Mozaffari, Amanda Duarte, Lina Teckentrup, Stefano Materia, Gina E. C. Charnley, Lluis Palma, Eulalia Baulenas Serra, Dragana Bojovic, Paula Checchia, Aude Carreric, Francisco Doblas-Reyes

机构 * Catalan Institution for Research and Advanced Studies (ICREA)(加泰罗尼亚研究与高级研究机构)

专题命中 评测与基准 :large language model(abstract);language model(abstract);foundation model(abstract);分类 cs.AI、cs.LG

AI总结 人工智能在天气和气候信息中的兴起加剧了全球不平等,需通过数据为中心的发展、气候数字基础设施和知识共生产来解决不平等问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14707 2026-07-21 cs.CL cs.AI 版本更新 79%

Harnessing LLMs for Reliable Academic Supervision: A Comparative Study

利用大语言模型进行可靠的学术监督:一项比较研究

Akash Raj

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 该研究以学术监督为案例,比较无支架的GPT-5聊天机器人基线与多模块系统ASuS,ASuS将小模型GPT-4o-mini封装在LangGraph支架中,经评估发现ASuS在各维度表现更优,提取了七种模式,挑战“大模型更好”直觉。

Comments 16 pages, 4 tables, 1 figure. Code and data available at https://github.com/AkashRajSingh/Harnessing-LLMs-for-Reliable-Academic-Supervision

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17356 2026-07-21 cs.AI cs.LG 版本更新 79%

Comprehend, Divide, and Conquer: Feature Subspace Exploration via Multi-Agent Hierarchical Reinforcement Learning

理解、划分与征服:通过多智能体分层强化学习进行特征子空间探索

Weiliang Zhang, Xiaohan Huang, Yi Du, Ziyue Qiao, Qingqing Long, Zhen Meng, Yuanchun Zhou, Meng Xiao

机构 * Computer Network Information Center, Chinese Academy of Sciences(中国科学院计算机网络信息中心) University of Chinese Academy of Sciences(中国科学院大学) Great Bay University(Great Bay大学) Duke-NUS Medical School, National University of Singapore(新加坡国立大学杜克-奈素医学院)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 研究针对特征选择问题,提出HRLFS方法,先利用基于大语言模型的混合状态提取器捕捉特征特性并聚类,构建分层智能体,通过多智能体分层强化学习进行特征子空间探索,提升了下游机器学习性能并加速运行

Comments 25 pages, keywords: Automated Feature Engineering, Tabular Dataset, Multi-Agent Reinforcement Learning, Feature Selection, Accepted by ACM Transactions on Knowledge Discovery from Data

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26842 2026-07-17 cs.LG cs.AI cs.CV 版本更新 79%

VAN-AD: Visual Masked Autoencoder with Normalizing Flow For Time Series Anomaly Detection

VAN-AD:基于归一化流的视觉掩码自编码器用于时间序列异常检测

PengYu Chen, Shang Wan, Xiaohou Shi, Yuan Chang, Yan Sun, Sajal K. Das

机构 * School of Computer Science (National Pilot Software Engineering School)(计算机学院(国家级试点软件工程学院)) Beijing University of Posts and Telecommunications(北京邮电大学) China Telecom Research Institute Beijing(中国电信研究院北京) Department of Computer Science, Missouri University of Science and Technology(计算机科学系,密苏里科技大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);foundation model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出VAN-AD,通过归一化流模块和自适应分布映射模块改进视觉掩码自编码器,提升时间序列异常检测的泛化能力和局部感知能力。

Comments 15 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15239 2026-07-15 cs.CL cs.AI econ.GN q-fin.EC stat.ME 版本更新 79%

Modeling Story Expectations: A Generative Framework using LLMs

建模故事期望:一种使用大语言模型的生成框架

Hortense Fong, George Gui, Bo Yang

机构 * Columbia Business School(哥伦比亚商学院)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 研究针对非结构化叙事内容建模消费者故事期望的难题,利用大语言模型生成故事延续并提取特征,通过两种验证程序,将其应用于不同数据发现模型期望与人类信念及实际故事延续相关,为叙事内容信念建模提供可扩展方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06401 2026-07-14 cs.AI cs.CL 版本更新 79%

BizFinBench.v2: Towards Reliable LLMs in Finance via Real-User Data and Offline/Online Bilingual Evaluation

BizFinBench.v2:通过真实用户数据和离线/在线双语评估实现金融领域可靠的大语言模型

Xin Guo, Rongjunchen Zhang, Guilong Lu, Xuntao Guo, Shuai Jia, Zhi Yang, Liwen Zhang

机构 * HiThink Research(HiThink研究机构) Shanghai University of Finance and Economics(上海财经大学)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 研究针对大语言模型在金融应用中实际效果与报告性能差距大的问题,构建BizFinBench.v2基准,涵盖中美股票市场真实用户数据及多个任务,通过实验评估模型,揭示现有模型局限,为金融领域大语言模型部署提供基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20459 2026-07-03 cs.AI cs.CL 版本更新 79%

PreScience: A Dataset and Benchmark for Scientific Forecasting

PreScience:一个用于科学预测的数据集和基准

Anirudh Ajith, Amanpreet Singh, Jay DeYoung, Nadav Kunievsky, Austin C. Kozlowski, Oyvind Tafjord, James Evans, Daniel S. Weld, Tom Hope, Doug Downey

机构 * Allen Institute for Artificial Intelligence(Allen人工智能研究所) Knowledge Lab, University of Chicago(芝加哥大学知识实验室) School of Computer Science and Engineering, Hebrew University of Jerusalem(耶路撒冷希伯来大学计算机科学与工程学院) Northwestern University(西北大学)

专题命中 评测与基准 :LLM(abstract,abstract_cn);language model(abstract);分类 cs.CL、cs.AI

AI总结 提出PreScience数据集和基准,包含98K篇AI论文及关联数据,设计7项预测任务,开发基线方法及LACER评估指标,发现合成论文多样性低于人类。

Comments 11 pages (70 with bibliography and appendix), 3 figures (14 with appendix), 5 tables (18 with appendix), 1 algorithm in appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06676 2026-06-23 cs.CL cs.AI cs.HC 版本更新 79%

One Interaction Is Worth a Thousand Guesses: Benchmarking the Interactive Capabilities of Deep Research Agents

一次交互胜过千次猜测:深度研究代理的交互能力基准测试

Yingchaojie Feng, Qiang Huang, Xiaoya Xie, Zhaorui Yang, Jun Yu, Wei Chen, Anthony K. H. Tung

机构 * School of Computing, National University of Singapore(新加坡国立大学计算机学院) School of Intelligence Science and Engineering, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)智能科学与工程学院) Zhejiang University(浙江大学) State Key Lab of CAD&CG, Zhejiang University(浙江大学计算机辅助设计与图形学国家重点实验室)

专题命中 评测与基准 :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 提出IDRBench基准,通过交互式框架和用户模拟器评估深度研究代理的交互能力,实验表明交互能提升研究质量和鲁棒性。

Comments 17 pages, 9 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29676 2026-06-18 cs.AI cs.CL 版本更新 79%

Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems

符号至关重要:智能体AI系统中令牌优化格式的基准研究

Lorenz Kutschka, Bernhard Geiger

机构 * Know Center Research GmbH(知中心研究有限公司) Graz University of Technology(格拉茨技术大学) Graz Center for Machine Learning(格拉茨机器学习中心)

专题命中 评测与基准 :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本研究在四个智能体基准上评估了两种令牌优化格式TOON和TRON,发现TRON在保持准确率的同时最多减少27%的令牌,而TOON虽减少18%但存在多轮解析失败和并行工具调用输出崩溃的问题。

Comments 16 pages, 6 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18154 2026-06-12 cs.CL cs.AI cs.DB 版本更新 79%

FENCE: A Financial and Multimodal Jailbreak Detection Dataset

FENCE:一个金融和多模态越狱检测数据集

Mirae Kim, Seonghun Jeong, Youngjun Kwak

专题命中 评测与基准 :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 针对金融领域多模态越狱检测资源匮乏的问题,提出FENCE数据集,包含韩英双语文本和图像,用于训练和评估检测器,实验表明基线检测器准确率达99%。

Comments lrec 2026 accepted paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12306 2026-06-10 cs.LG cs.AI 版本更新 79%

GCA Framework: A GCC Countries-Grounded Dataset and Agentic Pipeline for Climate Decision Support

GCA框架:面向海湾合作委员会国家的数据集与气候决策支持智能体管道

Muhammad Umer Sheikh, Khawar Shehzad, Salman Khan, Fahad Shahbaz Khan, Muhammad Haris Khan

机构 * Mohamed Bin Zayed University of Artificial Intelligence (MBZUAI)(莫扎德人工智能大学) University of Missouri(密苏里大学) Australian National University(澳大利亚国立大学) Linköping University(林肯大学)

专题命中 评测与基准 :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 提出GCA框架,包含GCC国家多模态数据集GCA-DS和工具增强型智能体GCA,通过领域微调和工具集成提升气候决策可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22097 2026-06-09 cs.SE cs.AI cs.CL cs.CR 版本更新 79%

SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios

SecureVibeBench: 通过重建引入漏洞的场景来基准测试AI代理的安全振动编码

Junkai Chen, Huihui Huang, Yunbo Lyu, Junwen An, Jieke Shi, Chengran Yang, Ting Zhang, Haoye Tian, Yikun Li, Zhenhao Li, Xin Zhou, Xing Hu, David Lo

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 评测与基准 :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出SecureVibeBench,一个包含105个C/C++安全编码任务的基准测试,旨在评估AI代理在真实场景中生成安全代码的能力,发现现有方法在评估人类与AI代理对比时的不足。

Comments ACL 2026 Main Conference. Our code and data are on https://github.com/iCSawyer/SecureVibeBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18632 2026-08-21 cs.RO 版本更新 78%

ROBOSHACKLES: A Safety Dataset for Human-Injury Prevention in Embodied Foundation Models

ROBOSHACKLES: 面向具身基础模型中人体伤害预防的安全数据集

Zhuowen Yin, Chongyang Liu, Wenzhang Yang, Renjue Li, Yinxing Xue

机构 * Institute of Al for Industries, Chinese Academy of Sciences(工业人工智能研究所,中国科学院) University of Science and Technology of China(中国科学技术大学)

专题命中 评测与基准 :foundation model(title,abstract)

AI总结 为解决机器人伤害人类数据难以安全收集的问题,提出基于真实观测的安全数据构建流水线,生成包含1万条视频的ROBOSHACKLES数据集,涵盖直接和间接伤害类别,评估发现现有模型在安全关键场景下100%产生不安全动作。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17809 2026-08-20 cs.CV 版本更新 78%

Million-scale multimodal pollen microscopy with expert-guided foundation models

百万级多模态花粉显微镜图像与专家引导的基础模型

András Biricz, Björn Gedda, Donát Magyar, Antonio Spanu, János Fillinger, Péter Pollner, István Csabai

机构 * Department of Physics of Complex Systems, ELTE Eötvös Loránd University(ELTE罗兰大学复杂物理系) The Palynological Laboratory at the Swedish Museum of Natural History(瑞典自然历史博物馆孢粉学实验室) National Centre for Public Health and Pharmacy(国家公共卫生与药品中心) INRAE, UR 546 BioSP, Site Agroparc(法国国家农业、食品与环境研究院,UR 546 BioSP,阿格罗帕克园区) National Korányi Institute for Pulmonology(国家科拉尼肺病研究所) Health Data Science and AI Knowledge Centre, Health Services Management Training Centre, Faculty of Health and Public Administration, Semmelweis University(塞梅维什大学健康与公共管理学院卫生服务管理培训中心健康数据科学与人工智能知识中心) Department of Biological Physics, ELTE Eötvös Loránd University(ELTE罗兰大学生物物理系)

专题命中 评测与基准 :foundation model(title);language model(abstract)

AI总结 提出百万级多模态花粉显微镜数据集Pollen AI Atlas,结合专家引导的视觉-语言模型生成形态描述,实现跨区域、跨设置的高精度花粉识别与检索。

Comments 31 pages, 5 main figures, supplementary information included. Submitted to Scientific Reports. v2: clarified reporting of taxonomic scope, captioning settings, backbone configuration, and evaluation details; no changes to numerical results or conclusions

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16871 2026-08-18 cs.RO 版本更新 78%

SADP: Subgoal-Aware Diffusion Policy for Long-Horizon Manipulation Learned from Foundation Model Generated Demonstrations

SADP:基于基础模型生成示范的子目标感知扩散策略用于可解释机器人

Site Hu, Takato Horii

机构 * Department of Systems Innovation, Graduate School of Engineering Science, Osaka University(系统创新系,工学研究科,大阪大学)

专题命中 评测与基准 :foundation model(title,abstract)

AI总结 本文提出SADP,一种基于基础模型生成示范的子目标感知扩散策略,用于可解释机器人,通过自主生成子目标标注的示范数据,训练扩散策略,使机器人能够通过子目标结构和执行进度向用户解释决策过程,从而在长周期操作中实现更高的任务成功率和故障诊断能力。

Comments Revised manuscript with an updated title, evaluation protocol, and simulation results

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20271 2026-08-18 cs.CV 版本更新 78%

A Versatile Foundation Model for AI-enabled Mammogram Interpretation

一种用于AI辅助乳腺X线摄影解读的通用基础模型

Fuxiang Huang, Jiayi Zhu, Yunfang Yu, Yu Xie, Yuan Guo, Qingcong Kong, Mingxiang Wu, Xinrui Jiang, Shu Yang, Jiabo Ma, Ziyi Liu, Zhe Xu, Zhixuan Chen, Yujie Tan, Zifan He, Luhui Mao, Xi Wang, Junlin Hou, Lei Zhang, Qiong Luo, Zhenhui Li, Herui Yao, Hao Chen

机构 * Guangdong Provincial Key Laboratory of Malignant Tumor Epigenetics and Gene Regulation, Guangdong-Hong Kong Joint Laboratory for RNA Medicine, Department of Medical Oncology, Breast Tumor Centre, Phase I Clinical Trial Centre(广东省恶性肿瘤表观遗传与基因调控重点实验室,粤港澳RNA医学联合实验室,医学肿瘤科,乳腺肿瘤中心,I期临床试验中心) Guangdong Provincial Key Laboratory of Cancer Pathogenesis and Precision Diagnosis and Treatment, AI Big Data Laboratory, Department of Medical Oncology(广东省肿瘤发生与精准诊断与治疗重点实验室,人工智能大数据实验室,医学肿瘤科) Chongqing Key Laboratory of Bio-perception(重庆生物感知重点实验室)

专题命中 评测与基准 :foundation model(title,abstract)

AI总结 研究针对乳腺X线分析基础模型的临床转化缺陷,推出VersaMammo模型,构建含70余万张图像的多机构数据集,经两阶段预训练后在92项临床任务基准中表现顶尖,推动乳腺癌筛查诊断进展。

Comments 69 pages, 12 figures, 52 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10119 2026-08-17 cs.SE 版本更新 78%

IntrinTrans: LLM-based Intrinsic Code Translator for RISC-V Vector

IntrinTrans: 基于大语言模型的RISC-V向量内在函数翻译器

Liutong Han, Zhiyuan Tan, Hongbin Zhang, Pengcheng Wang, Chu Kang, Mingjie Xing, Yanjun Wu

专题命中 评测与基准 :LLM(title,abstract)

AI总结 本文提出IntrinTrans,利用大语言模型和编译反馈自动翻译跨架构内在函数,并通过寄存器使用信息优化生成代码,实验证明其性能接近原生实现。

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09302 2026-08-12 cs.CV 版本更新 78%

Bootstrapping Vision-Language Model for Hysteroscopic Surgical Scene Segmentation

用于宫腔镜手术场景分割的自举式视觉-语言模型

Jun Huang, Meiyi Chen, Zijie Yue, Yuhang Xiao, Fang Li, Hanli Wang, Xiaowen Tong, Yi Guo, Miaojing Shi

专题命中 评测与基准 :language model(title,abstract)

AI总结 本研究提出首个基于VLM的宫腔镜手术场景分割方法VLM-hyster,通过类别特定文本提示与掩码蒸馏分支提升性能,在自行构建的4020张图像数据集上表现优于现有模型,获多中心验证,具临床应用潜力。

Comments Accept by Biomedical Signal Processing and Control

详情

展开后加载摘要…

URL PDF HTML 收藏