arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7409 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7409 篇

2604.10480 2026-04-14 cs.AI 57%

Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs

追溯根源:一种多智能体框架,用于揭示训练后LLM的数据血缘

Yu Li, Xiaoran Shang, Qizhi Pei, Yun Zhu, Xin Gao, Honglin Lin, Zhanping Zhong, Zhuoshi Pan, Zheng Liu, Xiaoyang Wang, Conghui He, Dahua Lin, Feng Zhao, Lijun Wu

机构 * Shanghai Artificial Intelligence Laboratory, OpenDataLab(上海人工智能实验室,OpenDataLab) University of Science and Technology of China(中国科学技术大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文提出多智能体框架,用于揭示训练后LLM的数据血缘,识别数据集演化图谱中的结构模式与系统问题,构建血缘感知的多样化数据集。

Comments 27 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03447 2026-04-14 cs.CV 57%

CoPS: Conditional Prompt Synthesis for Zero-Shot Anomaly Detection

CoPS:用于零样本异常检测的条件提示合成

Qiyu Chen, Zhen Qu, Wei Luo, Haiming Yao, Yunkang Cao, Yuxin Jiang, Yinan Duan, Huiyuan Luo, Chengkan Lv, Zhengtao Zhang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) CASIVISION(中科视语) THU(清华大学) HNU(湖南大学) HUST(华中科技大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 本文提出CoPS框架,通过动态提示合成提升零样本异常检测性能,利用视觉特征生成适应性状态模型,并在13个工业和医疗数据集上取得分类AUROC和分割AUROC的提升。

Comments Accepted by CVPR 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09920 2026-04-14 cs.CV 57%

Does Your VFM Speak Plant? The Botanical Grammar of Vision Foundation Models for Object Detection

你的VFM说话植物吗?面向物体检测的视觉基础模型的植物语法

Lars Lundqvist, Earl Ranario, Hamid Kamangir, Heesup Yun, Christine Diepenbrock, Brian N. Bailey, J. Mason Earles

机构 * University of California, Davis(加州大学戴维斯分校)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 本文提出系统提示优化框架,评估四个开放词汇检测器在合成和真实农田图像中对豆科植物花和豆荚的检测性能,发现提示结构对不同模型响应各异,通过模型特定的组合提示显著提升检测精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09855 2026-04-14 cs.AI cs.CL cs.GT econ.GN q-fin.EC 57%

Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards

通过可验证奖励的强化学习指导大语言模型进行谈判

Shuze Daniel Liu, Claire Chen, Jiabao Sean Xiao, Lei Lei, Yuheng Zhang, Yisong Yue, David Simchi-Levi

机构 * Purdue University(普渡大学) California Institute of Technology(加州理工学院) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Massachusetts Institute of Technology(麻省理工学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文探讨了通过可验证奖励的强化学习训练大语言模型进行谈判的有效性,揭示了四阶段战略演化,展示了训练后的模型在谈判中提取剩余价值的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09711 2026-04-14 cs.CV cs.CL 57%

Head-wise Modality Specialization within MLLMs for Robust Fake News Detection under Missing Modality

在多模态大语言模型中实现模态专精以在缺失模态下实现鲁棒的虚假新闻检测

Kai Qian, Weijie Shi, Jiaqi Wang, Mengze Li, Hao Chen, Yue Cui, Hanghui Guo, Ziyi Liu, Jia Zhu, Jiajie Xu

机构 * Soochow University(苏州大学) Hong Kong University of Science and Technology(香港科技大学) Computer Science and Engineering, The Chinese University of Hong Kong(香港中文大学计算机科学与工程系) The Hong Kong University of Science and Technology(香港科技大学) FiT, Tencent(腾讯FiT) Alibaba Group(阿里巴巴集团) School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) School of Education, Zhejiang Normal University(浙江师范大学教育学院)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

AI总结 本文提出在多模态大语言模型中通过头级专精机制提升鲁棒性,以应对缺失模态下的虚假新闻检测,通过保留模态专精能力与利用稀疏单模注释提升检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09644 2026-04-14 cs.CY cs.AI 57%

Detecting Corporate AI-Washing via Cross-Modal Semantic Inconsistency Learning

通过跨模态语义不一致性学习检测企业AI洗钱

Zhanjie Wen, Jingqiao Guo

机构 * School of Economics and Trade, Guangdong University of Finance(广东金融学院经济贸易学院) Department of Computer Science, Faculty of Science, Hong Kong Baptist University(香港浸会大学理学院计算机科学系)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文提出AWASH框架,通过跨模态主张-证据推理检测企业AI洗钱,利用AW-Bench基准测试,实现高准确率的AI能力识别。

Comments 28 pages, 6 figures, Journal Submission (Finance/Accounting & Computer Science Interdiscipline), 6 tables, 40 references, trimodal benchmark (88,412 firm-quarter observations) and end-to-end multimodal detection framework for corporate AI-washing

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09571 2026-04-14 cs.HC cs.AI 57%

Tuning Qwen2.5-VL to Improve Its Web Interaction Skills

调优 Qwen2.5-VL 以提升其网页交互能力

Alexandra Yakovleva, Henrik Pärssinen, Harri Valpola, Juho Kannala, Alexander Ilin

机构 * Aalto University(阿尔托大学) University of Oulu(奥卢大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.AI

AI总结 本文通过调优 Qwen2.5-VL 模型,改进其在网页交互中的可靠性,解决元素定位不准确、指令敏感性和过度乐观偏差等问题,提升单次点击任务的成功率。

Comments Accepted to the Short Paper Track of ACM Web Conference 2026 (WWW 2026). The final version will appear in the ACM Digital Library

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14629 2026-04-14 cs.CV cs.CL 57%

VisText-Mosquito: A Unified Multimodal Dataset for Visual Detection, Segmentation, and Textual Explanation on Mosquito Breeding Sites

VisText-Mosquito:一种统一的多模态数据集,用于蚊虫繁殖地的视觉检测、分割和文本解释

Md. Adnanul Islam, Md. Faiyaz Abdullah Sayeedi, Md. Asaduzzaman Shuvo, Shahanur Rahman Bappy, Md Asiful Islam, Swakkhar Shatabda

机构 * United International University, Bangladesh(孟加拉国联合国际大学) University of Arizona, USA(美国亚利桑那大学) BRAC University, Bangladesh(孟加拉国布拉卡大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 本文提出VisText-Mosquito数据集,整合视觉和文本数据以支持自动化检测、分割和解释,通过YOLOv9s和YOLOv11n-Seg模型实现高精度检测与分割,并利用大视觉语言模型生成文本解释,强调预防胜于治疗的主题。

Comments Accepted at CVPRW 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16091 2026-04-14 cs.CV 57%

Near OOD Detection for Vision-Language Prompt Learning with Contrastive Logit Score

面向视觉-语言提示学习的近分布外检测

Myong Chol Jung, Joanna Dipnall, Belinda Gabbe, He Zhao

机构 * Monash University(莫纳什大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 本文提出对比logit分数方法,用于提升预训练视觉-语言提示学习模型的近分布外检测性能,无需修改模型结构或重新训练,实验显示在AUROC上提升11.67%。

Comments Published at International Journal of Computer Vision (IJCV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09037 2026-04-13 cs.CV cs.CL cs.HC 57%

SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos

SiMing-Bench:从连续交互评估临床技能视频中的过程正确性

Xiyang Huang, Jiawei Lin, Keying Wu, Jiaxin Huang, Kailai Yang, Renxiong Wei, Cheng zeng, Jiayi Xiang, Ziyan Kuang, Min Peng, Qianqian Xie, Sophia Ananiadou

机构 * School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院) Center for Language and Information Research, Wuhan University(武汉大学语言与信息研究中心) Southwest Jiaotong University(西南交通大学) MBZUAI(穆罕默德·本·扎耶德人工智能大学) The University of Manchester(曼彻斯特大学) Zhongnan Hospital of Wuhan University(武汉大学中南医院)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

AI总结 SiMing-Bench旨在评估多模态大语言模型在临床技能视频中通过连续交互更新过程状态的能力,通过医师标注的数据集和标准化评分标准,揭示模型在过程级判断上的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08597 2026-04-13 cs.DB cs.AI 57%

STIndex: A Context-Aware Multi-Dimensional Spatiotemporal Information Extraction System

STIndex:一种基于上下文的多维时空信息提取系统

Wenxiao Zhang, Yu Liu, Qiang sun, Yihao Ding, Sirui Li, Yanbing Liu, Jin B. Hong, Wei Liu

机构 * The University of Western Australia(西澳大利亚大学) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) University of Chinese Academy of Sciences(中国科学院大学) Murdoch University(莫道克大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 STIndex通过多维时空数据仓库提升无结构数据的提取效率,结合大语言模型实现上下文感知提取,提升时空实体识别精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27979 2026-04-13 cs.CV 57%

RetinexDualV2: Physically-Grounded Dual Retinex for Generalized UHD Image Restoration

RetinexDualV2:基于物理的双Retinex用于通用超高清图像恢复

Mohab Kishawy, Jun Chen

机构 * McMaster University(麦克马斯特大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 本文提出RetinexDualV2,一种统一的双分支框架,用于超高清图像恢复。通过任务特定物理模块提取退化感知先验,结合物理条件多头自注意力机制,实现鲁棒的反射和光照校正,展现卓越的通用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19704 2026-04-13 cs.CV 57%

RADSeg: Unleashing Parameter and Compute Efficient Zero-Shot Open-Vocabulary Segmentation Using Agglomerative Models

RADSeg: 通过凝聚模型提升参数和计算效率的零样本开放词汇分割

Omar Alama, Darshil Jariwala, Avigyan Bhattacharya, Seungchan Kim, Wenshan Wang, Sebastian Scherer

机构 * Carnegie Mellon University(卡内基梅隆大学) IIIT Hyderabad(海得拉巴国际信息技术学院)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 RADSeg利用凝聚视觉基础模型RADIO,同时提升mIoU、延迟和参数效率,实现更高效的零样本开放词汇分割。

Comments Accepted to CVPR'26 Findings Code at https://radseg-ovss.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07772 2026-04-10 cs.CV 57%

ESOM: Efficiently Understanding Streaming Video Anomalies with Open-world Dynamic Definitions

ESOM:基于开放世界动态定义的高效流视频异常理解

Zihao Liu, Xiaoyu Wu, Wenna Li, Jianqin Wu, Linlin Yang

机构 * State Key Laboratory of Media Convergence and Communication, Communication University of China(媒体融合与传播国家重点实验室,中国传媒大学)

专题命中 视觉定位与Grounding :MLLM(abstract);分类 cs.CV

AI总结 本文提出ESOM模型,通过定义规范化、帧间匹配、混合流内存和概率评分模块,实现高效流视频异常检测,同时引入OpenDef-Bench基准测试,在单GPU上实现实时效率和先进性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.18030 2026-04-10 cs.OS cs.AI cs.PL cs.SE 57%

Quine: Realizing LLM Agents as Native POSIX Processes

Quine:将LLM代理作为原生POSIX进程实现

Hao Ke

机构 * Independent Researcher(独立研究员)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 Quine通过将LLM代理作为原生POSIX进程实现,利用操作系统提供的隔离、调度和通信机制,继承内核的隔离、组合和资源控制能力,支持递归委托和上下文刷新。

Comments Minor revision clarifying exec semantics

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20524 2026-04-10 cs.CV 57%

AnomalyVFM -- Transforming Vision Foundation Models into Zero-Shot Anomaly Detectors

AnomalyVFM -- 将视觉基础模型转变为零样本异常检测器

Matic Fučka, Vitjan Zavrtanik, Danijel Skočaj

机构 * University of Ljubljana, Faculty of Computer and Information Science(卢布尔雅那大学计算机与信息科学学院)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 本文提出AnomalyVFM框架,通过合成数据生成与参数高效适应机制,使视觉基础模型在零样本异常检测中表现更优,实测在9个数据集上达到94.1%的AUROC。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16282 2026-04-10 cs.CL cs.AI 57%

Generating Literature-Driven Scientific Theories at Scale

大规模生成文献驱动的科学理论

Peter Jansen, Peter Clark, Doug Downey, Daniel S. Weld

机构 * Allen Institute for Artificial Intelligence(艾伦人工智能研究所) University of Arizona(亚利桑那大学) University of Washington(华盛顿大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 研究通过大规模文献数据生成科学理论,比较文献基础与参数知识生成对理论质量和预测能力的影响。

Comments 9 pages plus appendix, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17445 2026-04-10 cs.CV 57%

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents

LangDriveCTRL: 通过多模态代理实现自然语言可控的驾驶场景编辑

Yun He, Francesco Pittaluga, Ziyu Jiang, Matthias Zwicker, Manmohan Chandraker, Zaid Tasneem

机构 * University of Maryland, College Park(马里兰大学帕克分校) NEC Labs America(NEC美国实验室) UC San Diego(加州大学圣迭戈分校)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 LangDriveCTRL通过多模态代理实现驾驶视频的自然语言可控编辑,生成多样化的交通场景,提升真实感和交通场景的真实性。

Comments Project Page: https://yunhe24.github.io/langdrivectrl/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08637 2026-04-10 cs.CY cs.AI cs.CR 57%

How Do Data Owners Say No? A Case Study of Data Consent Mechanisms in Web-Scraped Vision-Language AI Training Datasets

数据所有者如何说不?Web抓取视觉-语言AI训练数据集中的数据同意机制案例研究

Chung Peng Lee, Rachel Hong, Harry H. Jiang, Aster Plotnik, William Agnew, Jamie Morgenstern

机构 * Princeton University(普林斯顿大学) University of Washington(华盛顿大学) Carnegie Mellon University(卡内基梅隆大学) University of Toronto(多伦多大学) Amazon AWS AI/ML(亚马逊AWS AI/ML)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.AI

AI总结 研究探讨了数据所有者在AI训练数据集中的同意表达方式,分析了DataComp数据集中样本层面和网络层面的信息,揭示了当前数据收集流程对数据同意的不尊重问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06765 2026-04-09 cs.CL cs.AI 57%

TeamLLM: A Human-Like Team-Oriented Collaboration Framework for Multi-Step Contextualized Tasks

TeamLLM: 一种类人团队导向的多步骤上下文任务协作框架

Xiangyu Wang, Jin Wu, Haoran Shi, Wei Xia, Jiarui Yu, Chanjin Zheng

机构 * Department of Educational Psychology, East China Normal University(华东师范大学教育心理学系) Lab of Artificial Intelligence for Education, East China Normal University(华东师范大学教育人工智能实验室) Shanghai Institute of Artificial Intelligence for Education, East China Normal University(华东师范大学上海教育人工智能研究院) School of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院) Cascadia Institute for Neurotechnology (CIN)(卡斯卡迪亚神经技术研究所)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 TeamLLM通过引入四类团队角色和三阶段多LLM协作机制,提升多步骤上下文任务性能,提出CGPST基准测试集并验证其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06696 2026-04-09 cs.AI 57%

AgentGate: A Lightweight Structured Routing Engine for the Internet of Agents

AgentGate: 一种轻量级的结构化路由引擎用于物联网代理

Yujun Cheng, Enfang Cui, Hao Qin, Zhiyuan Liang, Qi Xu

机构 * School of Artificial Intelligence, University of Science and Technology Beijing(北京科技大学人工智能学院) China Telecom Research Institute(中国电信研究院) Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(中国科学院大学杭州高等研究院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文提出AgentGate,一种轻量级结构化路由引擎,用于在资源受限条件下高效且隐私友好的代理系统,通过分阶段决策和结构化接地提升路由效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06279 2026-04-09 physics.plasm-ph cs.AI 57%

Plasma GraphRAG: Physics-Grounded Parameter Selection for Gyrokinetic Simulations

等离子体图RAG:用于磁流体动力学模拟的物理基础参数选择

Ruichen Zhang, Feda AlMuhisen, Chenguang Wan, Zhisong Qu, Kunpeng Li, Youngwoo Cho, Kyungtak Lim, Virginie Grandgirard, Xavier Garbet

机构 * Nanyang Technological University(南洋理工大学) CEA, IRFM(法国原子能委员会,磁聚变研究所) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院) School of Physical and Mathematical Sciences, Nanyang Technological University(南洋理工大学物理与数学科学学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文提出Plasma GraphRAG,结合图检索增强生成与大语言模型,实现自动化、物理基础的参数范围识别,提升模拟可靠性并加速科学发现。

Comments 9 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09677 2026-04-09 cs.AI 57%

Logics-Parsing-Omni Technical Report

逻辑-解析-全能技术报告

Xin An, Jingyi Cai, Xiangyang Chen, Huayao Liu, Peiting Liu, Peng Wang, Bei Yang, Xiuwen Zhu, Yongfan Chen, Yan Gao, Yuan Gao, Baoyu Hou, Guangzheng Hu, Shuzhao Li, Weixu Qiao, Weidong Ren, Yanan Wang, Boyu Yang, Fan Yang, Jiangtao Zhang, Lixin Zhang, Lin Qu, Hu Wei, Xiaoxiao Xu, Bing Zhao

机构 * Alibaba Group(阿里巴巴集团)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文提出Omni Parsing框架,通过三级层次结构实现多模态解析,整合整体检测、细粒度识别和多级解释,构建证据锚定机制以提升逻辑推理能力,释放Logics-Parsing-Omni模型并发布评估基准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18336 2026-04-08 cs.CV cs.GR 57%

PPISP: Physically-Plausible Compensation and Control of Photometric Variations in Radiance Field Reconstruction

PPISP: 光度变化在辐射场重建中的物理合理补偿与控制

Isaac Deutsch, Nicolas Moënne-Loccoz, Gavriel State, Zan Gojcic

机构 * NVIDIA(英伟达)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 本文提出PPISP模块,通过物理基础和可解释的转换分离相机固有和拍摄依赖效应,实现对多视角3D重建中光度不一致的合理补偿与控制,提升重建效果和评估公平性。

Comments For more details and updates, please visit our project website: https://research.nvidia.com/labs/sil/projects/ppisp/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05623 2026-04-08 cs.CV cs.CL cs.MM 57%

DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions

DetailVerifyBench:长图像描述中密集 hallucination 定位的基准

Xinran Wang, Yuxuan Zhang, Xiao Zhang, Haolong Yan, Muxi Diao, Songyu Xu, Zhonghao Yan, Hongbing Li, Kongming Liang, Zhanyu Ma

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

AI总结 本文提出 DetailVerifyBench,通过1000张高质量图像和密集的token级注释,评估长图像描述中hallucination的精确定位能力。

Comments 8 pages, 5 figures. The dataset and code are available at https://zyx-hhnkh.github.io/DetailVerifyBench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05475 2026-04-08 cs.CV 57%

A Synthetic Eye Movement Dataset for Script Reading Detection: Real Trajectory Replay on a 3D Simulator

用于脚本阅读检测的合成眼动数据集:在3D模拟器上的真实轨迹回放

Kidus Zewde, Yuchen Zhou, Dennis Ng, Neo Tiangratanakul, Tommy Duong, Ankit Raj, Yuxin Zhang, Xingyu Shen, Simiao Ren

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 本文提出一种生成合成眼动视频的 pipeline,通过回放真实人类虹膜轨迹生成12小时的合成数据,用于脚本阅读检测任务,验证了生成轨迹在时间动态上的有效性。

Comments Synthetic eye movement dataset generation via 3D eye simulator; iris trajectory replay; script reading detection; behavioral data augmentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05467 2026-04-08 cs.IR cs.CL cs.LG 57%

CUE-R: Beyond the Final Answer in Retrieval-Augmented Generation

CUE-R:超越检索增强生成的最终答案

Siddharth Jain, Venkat Narayan Vedam

机构 * Intuit

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

AI总结 CUE-R通过干预基于证据项的检索使用痕迹,评估单次检索增强生成中每个证据项的操作效用,揭示证据项对正确性、基础忠实度和置信度误差的影响。

Comments 6 figures, 14 tables; appendix includes bootstrap CIs, metric definitions, duplicate position sensitivity, prompt template, and reproducibility details

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05278 2026-04-08 cs.SE cs.AI cs.MA 57%

Spec Kit Agents: Context-Grounded Agentic Workflows

Spec Kit Agents:基于上下文的代理工作流

Pardis Taghavi, Santosh Bhavani

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文提出Spec Kit Agents,一种多代理的规范驱动开发流程,通过添加阶段级上下文绑定钩子,提升代码质量与测试兼容性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09365 2026-04-08 cs.CL cs.AI 57%

Frame of Reference: Addressing the Challenges of Common Ground Representation in Situational Dialogs

参考框架:解决情境对话中共同地面表示的挑战

Biswesh Mohapatra, Théo Charlot, Giovanni Duca, Mayank Palan, Laurent Romary, Justine Cassell

机构 * Inria(法国国家信息与自动化研究所) Nantes Université(南特大学) University of Trento(特伦托大学) VJTI Mumbai(孟买维杰扬特理工学院) Carnegie Mellon University(卡内基梅隆大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文研究情境对话中共同地面表示的挑战,通过动态共享环境中的关系参考建立共同地面,并提出基于强化学习改进表示方法的策略。

Comments Work accepted at ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10287 2026-04-08 cs.LG cs.CL 57%

OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models

OutSafe-Bench:一种用于大型语言模型多模态攻击内容检测的基准测试

Yuping Yan, Yuhan Xie, Yuanshuai Li, Yingchao Yu, Lingjuan Lyu, Yaochu Jin

机构 * School of Engineering, Westlake University(西湖大学工学院) Sony AI(索尼人工智能)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.LG

AI总结 本文提出OutSafe-Bench,首个多模态时代内容安全评估测试套件,包含四类模态大规模数据集及多维交叉风险评分指标,评估九种内容风险类别,揭示九种先进MLLM的安全漏洞。

详情

展开后加载摘要…

URL PDF HTML 收藏