arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7409 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7409 篇

2603.02896 2026-03-04 cs.CV 57%

3D-DRES: Detailed 3D Referring Expression Segmentation

3D-DRES: 详细3D指称表达分割

Qi Chen, Changli Wu, Jiayi Ji, Yiwei Ma, Liujuan Cao

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 3D-DRES通过引入DetailRefer数据集和DetailBase架构,提升3D视觉语言理解的细粒度分割能力。

Comments AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02688 2026-03-04 cs.AI cs.RO 57%

Retrieval-Augmented Robots via Retrieve-Reason-Act

通过检索-推理-行动实现增强型机器人

Izat Temiraliev, Diji Yang, Yi Zhang

机构 * University of California, Santa Cruz(加州大学圣克ruz分校)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文提出检索增强型机器人学(RAR)范式,通过检索-推理-行动循环,使机器人能从外部文档获取未见程序知识,提升复杂任务执行能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02626 2026-03-04 cs.AI 57%

See and Remember: A Multimodal Agent for Web Traversal

见与忆:一种用于网页浏览的多模态代理

Xinjun Wang, Shengyao Wang, Aimin Zhou, Hao Hao

机构 * Shanghai Institute of AI for Education(上海人工智能教育研究院) East China Normal University(华东师范大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 V-GEMS通过视觉 grounding 和显式记忆系统实现精确稳健的网页浏览,实验显示其在导航任务中性能提升28.7%

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19096 2026-03-03 cs.LG 57%

The Power of Decaying Steps: Enhancing Attack Stability and Transferability for Sign-based Optimizers

衰减步骤的力量:提升基于符号的优化器的攻击稳定性和可迁移性

Wei Tao, Yang Dai, Jincai Huang, Qing Tao

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

AI总结 本文提出MDCS方法,通过衰减步长提升基于符号优化器的攻击稳定性和可迁移性,理论证明其收敛率最优。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01545 2026-03-03 cs.CV 57%

Training-Free Spatio-temporal Decoupled Reasoning Video Segmentation with Adaptive Object Memory

无需训练的时空解耦推理视频分割与自适应对象记忆

Zhengtong Zhu, Jiaqing Fan, Zhixuan Liu, Fanzhang Li

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

AI总结 本文提出无需训练的时空解耦推理视频分割方法SDAM,通过自适应对象记忆模块和时空解耦技术实现稳定的时间传播,优于现有微调方法。

Comments Accept by AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01465 2026-03-03 cs.RO cs.AI 57%

Non-Markovian Long-Horizon Robot Manipulation via Keyframe Chaining

非马尔可夫长时间 horizon 机器人操作 via 关键帧链

Yipeng Chen, Wentao Tan, Lei Zhu, Fengling Li, Jingjing Li, Guoli Yang, Heng Tao Shen

机构 * Tongji University(同济大学) University of Technology Sydney(技术悉尼大学) University of Electronic Science and Technology of China(电子科学与技术大学) Advanced Institute of Big Data(大数据先进研究院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文提出Keyframe-Chaining VLA框架,通过提取和链接关键历史帧,解决长horizon机器人操作中的非马尔可夫依赖问题,提升任务成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20999 2026-03-03 cs.CV 57%

VII: Visual Instruction Injection for Jailbreaking Image-to-Video Generation Models

VII: 视觉指令注入用于劫持图像到视频生成模型

Bowen Zheng, Yongli Xiang, Ziming Hong, Zerong Lin, Chaojian Yu, Tongliang Liu, Xinge You

机构 * National Anti-Counterfeit Engineering Research Center, Huazhong University of Science and Technology(华中科技大学反伪工程研究中心) School of Electronic Information and Communications, Huazhong University of Science and Technology(华中科技大学电子信息与通信学院) Sydney AI Centre, The University of Sydney(悉尼大学AI中心)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 VII通过将恶意意图伪装为良性视觉指令,有效劫持图像到视频生成模型,实现高攻击成功率并降低拒绝率。

Comments Project page: https://Zbwwwwwwww.github.io/VII

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16145 2026-03-03 cs.AI cs.CL 57%

SpiroLLM: Finetuning Pretrained LLMs to Understand Spirogram Time Series with Clinical Validation in COPD Reporting

SpiroLLM:通过临床验证在COPD报告中微调预训练大语言模型以理解肺功能时间序列

Shuhao Mei, Yongchao Long, Xiaoyu Xiao, Shan Cao, Xiaobo Han, Shijia Geng, Jinbo Sun, Yuxi Zhou, Shenda Hong

机构 * Guangzhou Institute of Technology, Xidian University, Xi’an, China(广州科技研究院,西安电子科技大学,中国) Department of Computer Science, Tianjin University of Technology, Tianjin, China(天津理工大学计算机学院,天津,中国) Department of Respiratory, The Second Hospital of Tianjin Medical University, China(天津医科大学第二医院呼吸科,中国) College of Pulmonary and Critical Care Medicine, Chinese PLA General Hospital, Beijing, China(中国人民解放军总医院呼吸与危重症医学科,北京,中国) HeartVoice Medical Technology, Hefei, China(合肥心声医疗技术有限公司,中国) School of Life Science and Technology, Xidian University, Xi’an, China(西安电子科技大学生命科学与技术学院,中国) National Institute of Health Data Science, Peking University, Beijing, China(北京大学国家健康数据科学研究院,北京,中国) Institute for Artificial Intelligence, Peking University, Beijing, China(北京大学人工智能研究院,北京,中国)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.AI

AI总结 SpiroLLM通过融合生理信号与大语言模型,实现对肺功能时间序列的解读,并在COPD诊断中展现出高准确性和稳健性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23348 2026-03-03 cs.RO cs.CV 57%

Physically Ground Commonsense Knowledge for Articulated Object Manipulation with Analytic Concepts

为拟合物体操纵的物理常识知识而引入分析概念

Jiude Wei, Yuxuan Li, Cewu Lu, Jianhua Sun

机构 * School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院) School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机学院) Shanghai Innovation Institute(上海创新研究院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 本文提出通过引入分析概念,将语义级常识知识接地到物理世界,以提升机器人对关节物体的通用精确操纵能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16953 2026-03-03 cs.CV 57%

Towards Real Zero-Shot Camouflaged Object Segmentation without Camouflaged Annotations

面向无遮蔽标注的零样本遮蔽物分割

Cheng Lei, Jie Fan, Xinran Li, Tianzhu Xiang, Ao Li, Ce Zhu, Le Zhang

机构 * University of Electronic Science and Technology of China(电子科学与技术大学) Space42

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

AI总结 本文提出了一种无需遮蔽标注的零样本遮蔽物分割框架,通过结合MIM、M-LLM和MFA机制,实现高效分割与快速推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01072 2026-03-03 cs.CV 57%

DAWA: Dynamic Ambiguity-Wise Adaptation for Real-Time Domain Adaptive Semantic Segmentation

DAWA: 动态模糊度适应于实时领域自适应语义分割

Taorong Liu, Zhen Zhang, Liang Liao, Jing Xiao, Chia-Wen Lin

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) Hangzhou Institute of Technology, Xidian University(西安电子科技大学杭州研究院) School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院) Department of Electrical Engineering(电气工程系) the Institute of Communications Engineering, National Tsing Hua University(清华大学通信工程研究所)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 DAWA通过动态模糊度适应策略,在实时领域自适应语义分割中提升效率和效果,实现40 FPS的高速推理性能。

Comments PRCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01029 2026-03-03 cs.CV 57%

Vision-Language Feature Alignment for Road Anomaly Segmentation

视觉-语言特征对齐用于道路异常分割

Zhuolin He, Jiacheng Tang, Jian Pu, Xiangyang Xue

机构 * School of Computer Science, Fudan University(复旦大学计算机科学学院) Institute of Science and Technology for Brain-Inspired Intelligence, Fudan University(复旦大学脑启发智能科学与技术研究院)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 VL-Anomaly通过结合视觉-语言模型的语义先验,提出了一种新的道路异常分割框架,有效提升异常检测的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00540 2026-03-03 cs.AI 57%

LOGIGEN: Logic-Driven Generation of Verifiable Agentic Tasks

LOGIGEN: 逻辑驱动的可验证代理任务生成

Yucheng Zeng, Weipeng Lu, Linyun Liu, Shupeng Li, Zitian Qu, Chenghao Zhu, Shaofei Li, Zhengdong Tan, Mengyue Liu, Haotian Zhao, Zhe Zhou, Jianmin Wu

机构 * Baidu Inc.(百度公司) Tsinghua University(清华大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 LOGIGEN通过逻辑驱动和验证训练生成可验证的复杂任务,提升代理在复杂环境中的任务完成率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00486 2026-03-03 cs.CV 57%

Random Wins All: Rethinking Grouping Strategies for Vision Tokens

随机胜出:重新思考视觉token的分组策略

Qihang Fan, Yuang Ai, Huaibo Huang, Ran He

机构 * MAIS & NLPR, Institute of Automation, Chinese Academy of Sciences, Beijing, China(自动化研究所,中国科学院,北京) School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China(人工智能学院,中国科学院大学,北京)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 本文提出随机分组策略,通过简单方法提升视觉token处理效率,实验显示其在多种任务中表现优异。

Comments Accepted by CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00267 2026-03-03 cs.AI cs.IR cs.SI 57%

Multi-Sourced, Multi-Agent Evidence Retrieval for Fact-Checking

多源多智能体证据检索用于事实核查

Shuzhi Gong, Richard O. Sinnott, Jianzhong Qi, Cecile Paris, Preslav Nakov, Zhuohan Xie

机构 * The University of Melbourne(墨尔本大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 WKGFC通过多智能体证据检索和知识图谱增强,提升事实核查的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17050 2026-03-03 cs.CV 57%

Geodesic Prototype Matching via Diffusion Maps for Interpretable Fine-Grained Recognition

通过扩散映射的测地原型匹配用于可解释的细粒度识别

Junhao Jia, Yunyou Liu, Yifei Sun, Huangwei Chen, Feiwei Qin, Changmiao Wang, Yong Peng

机构 * Hangzhou Dianzi University(杭州电子科技大学) Zhejiang University(浙江大学) Shenzhen Research Institute of Big Data(深圳大数据研究院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 本文提出通过扩散映射的测地原型匹配方法,用于可解释的细粒度识别,通过内在几何结构提升原型的语义对应性。

Comments The paper has been accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10067 2026-03-03 cs.CV 57%

FiLo++: Zero-/Few-Shot Anomaly Detection by Fused Fine-Grained Descriptions and Deformable Localization

FiLo++: 通过融合细粒度描述和变形定位实现零/少样本异常检测

Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen, Ming Tang, Jinqiao Wang

机构 * Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(基础模型研究中心,自动化研究所,中国科学院) School of Artifcial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Wuhan AI Research(武汉AI研究所) Peng Cheng Laboratory(鹏城实验室) Guangdong Provincial Key Laboratory of Intellectual Property & Big Data, Guangdong Polytechnic Normal University(广东省知识产权与大数据重点实验室,广东工业大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 FiLo++通过融合细粒度描述和变形定位技术,提升零/少样本异常检测的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00156 2026-03-03 cs.CV 57%

BiCLIP: Bidirectional and Consistent Language-Image Processing for Robust Medical Image Segmentation

BiCLIP: 用于鲁棒医学图像分割的双向和一致语言-图像处理

Saivan Talaei, Fatemeh Daneshfar, Abdulhady Abas Abdullah, Mustaqeem Khan

机构 * Department of Computer Engineering, University of Kurdistan, Iran(伊朗库尔德大学计算机工程系) Artificial Intelligence and Innovation Centre, University of Kurdistan, Erbil, Iraq(伊拉克埃尔比尔库尔德大学人工智能与创新中心) College of Information Technology, United Arab Emirates University, UAE(阿联酋大学信息科技学院)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 BiCLIP通过双向多模态融合和一致性目标提升医学图像分割的鲁棒性,有效应对标注稀少和临床伪影挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00098 2026-03-03 stat.OT cs.CY cs.LG econ.GN math.PR q-fin.EC 57%

Profiling vs. Case-specific Evidence: A Probabilistic Analysis

案件 profiling 与个案证据:一种概率分析

Marcello Di Bello, Nicolò Cangiotti, Michele Loi

机构 * Arizona State University(亚利桑那州立大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

AI总结 本文通过概率分析探讨了 profiling 证据与个案证据的区别,指出 profiling 证据不能直接证明被告具体罪行,挑战了其在刑事审判中的有效性。

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24021 2026-03-02 cs.CV 57%

Steering and Rectifying Latent Representation Manifolds in Frozen Multi-modal LLMs for Video Anomaly Detection

在冻结的多模态大语言模型中引导和校正潜在表示流形以进行视频异常检测

Zhaolin Cai, Fan Li, Huiyu Duan, Lijun He, Guangtao Zhai

专题命中 视觉定位与Grounding :MLLM(abstract);分类 cs.CV

AI总结 SteerVAD通过引导和校正冻结多模态大语言模型的潜在表示流形,提升视频异常检测的性能,仅需1%训练数据即达最优效果。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23869 2026-03-02 cs.CV 57%

Open-Vocabulary Semantic Segmentation in Remote Sensing via Hierarchical Attention Masking and Model Composition

通过层次化注意力掩码和模型组合实现遥感开放词汇语义分割

Mohammadreza Heidarianbaei, Mareike Dorozynski, Hubert Kanyamahanga, Max Mehltretter, Franz Rottensteiner

机构 * Institute of Photogrammetry and GeoInformation(摄影测量与地理信息研究所)

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV

AI总结 ReSeg-CLIP通过层次化注意力掩码和模型组合方法,实现无需训练的遥感开放词汇语义分割,取得最优性能。

Comments Published in the proceedings of the British Machine Vision Conference Workshops 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23656 2026-03-02 cs.CL cs.AI 57%

TRIZ-RAGNER: A Retrieval-Augmented Large Language Model for TRIZ-Aware Named Entity Recognition in Patent-Based Contradiction Mining

TRIZ-RAGNER: 一种用于基于专利矛盾挖掘的TRIZ-aware命名实体识别的检索增强型大语言模型

Zitong Xu, Yuqing Wu, Yue Zhao

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 TRIZ-RAGNER通过整合TRIZ知识和检索增强技术,提升专利矛盾挖掘中命名实体识别的准确性和一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06940 2026-03-02 cs.DB cs.AI 57%

VISTA: Knowledge-Driven Vessel Trajectory Imputation with Repair Provenance

VISTA: 基于知识的船舶轨迹补全与修复溯源

Hengyu Liu, Tianyi Li, Haoyu Wang, Kristian Torp, Tiancheng Zhang, Yushuai Li, Christian S. Jensen

机构 * Department of Computer Science, Aalborg University, Denmark(计算机科学系,奥胡斯大学,丹麦) School of Computer Science and Engineering, Northeastern University, Shenyang, China(计算机科学与工程学院,东北大学,沈阳,中国)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 VISTA通过基于知识的可解释方法实现船舶轨迹补全,生成修复溯源以提升下游决策的可信度和效率。

Comments 24 pages, 14 figures, 4 algorithms, 8 tables. Code available at https://github.com/hyLiu1994/VISTA

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19364 2026-03-02 cs.AI cs.MA 57%

Integrating LLM in Agent-Based Social Simulation: Opportunities and Challenges

将大语言模型整合到基于代理的社会模拟中:机遇与挑战

Patrick Taillandier, Jean Daniel Zucker, Arnaud Grignard, Benoit Gaudou, Nghi Quang Huynh, Alexis Drogoul

机构 * UR MIAT, University of Toulouse, INRAE, Castanet-Tolosan, France(法国图卢兹大学UR MIAT,INRAE,Castanet-Tolosan分校) UMI 209 UMMISCO, IRD/Sorbonne University, Bondy, France(法国Bondy IRD/索邦大学UMI 209 UMMISCO) LMI ACROSS, Thuyloi University, Hanoi, Vietnam(越南胡志明大学LMI ACROSS) UMR 5505 IRIT, University of Toulouse Capitole, Toulouse, France(法国图卢兹大学Capitole UMR 5505 IRIT) CICT, Can Tho University, Can Tho, Vietnam(越南金瓯大学CICT)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文探讨了大语言模型在社会模拟中的应用,分析其潜力与局限,并提出混合宪法架构作为整合LLM与传统代理模型的可行方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23339 2026-02-27 cs.CV 57%

Retrieve and Segment: Are a Few Examples Enough to Bridge the Supervision Gap in Open-Vocabulary Segmentation?

检索与分割:在开放词汇分割中,几个示例是否足以弥合监督差距?

Tilemachos Aravanis, Vladan Stojnić, Bill Psomas, Nikos Komodakis, Giorgos Tolias

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 本文提出了一种基于检索增强的测试时间适配器,通过融合文本和视觉支持特征来实现开放词汇分割,有效缩小了零样本与监督分割间的差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22963 2026-02-27 cs.AI 57%

FactGuard: Agentic Video Misinformation Detection via Reinforcement Learning

FactGuard:通过强化学习进行代理视频虚假信息检测

Zehao Li, Hongwei Yu, Hao Jiang, Qiang Sheng, Yilong Xu, Baolong Bi, Yang Li, Zhenlong Yuan, Yujun Cai, Zhaoqi Wang

机构 * Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China(中国科学院计算技术研究所) University of Chinese Academy of Sciences, Beijing, China(中国科学院大学) University of Science and Technology Beijing(北京科技大学) The University of Queensland, Brisbane, Australia(昆士兰大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.AI

AI总结 FactGuard通过强化学习和代理框架,提升视频虚假信息检测的鲁棒性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11620 2026-02-27 cs.AI 57%

A Mind Cannot Be Smeared Across Time

意识无法跨越时间被抹平

Michael Timothy Bennett

机构 * Michael Timothy Bennett(独立研究者)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 作者提出意识统一性需要客观同时实例化,而非时间顺序实现,指出软件意识在严格顺序子系统中无法实现,强调硬件对意识的重要性。

Comments Forthcoming in the proceedings of the AAAI 2026 Spring Symposium on Machine Consciousness: Integrating Theory, Technology, and Philosophy

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06139 2026-02-27 cs.CV 57%

Deforming Videos to Masks: Flow Matching for Referring Video Segmentation

视频变形到掩码:用于指认视频分割的流匹配

Zanyi Wang, Dengyang Jiang, Liuzhuozheng Li, Sizhe Dang, Chengzu Li, Harry Yang, Guang Dai, Mengmeng Wang, Jingdong Wang

机构 * SGIT AI Lab, State Grid Corporation of China(国网信通研究院) University of California, San Diego(加州大学圣地亚哥分校) The Hong Kong University of Science and Technology(香港科技大学) The University of Tokyo(东京大学) University of Cambridge(剑桥大学) Zhejiang University of Technology(浙江工业大学) Baidu(百度)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 本文提出FlowRVS框架,将指认视频分割视为条件连续流问题,通过学习视频整体表示到目标掩码的直接语言引导变形,在多个基准上取得新突破。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22219 2026-02-27 cs.IR cs.AI cs.CL 57%

Comparative Analysis of Neural Retriever-Reranker Pipelines for Retrieval-Augmented Generation over Knowledge Graphs in E-commerce Applications

电商应用场景下知识图谱检索增强生成中神经检索-排序流水线的比较分析

Teri Rumble, Zbyněk Gazdík, Javad Zarrin, Jagdeep Ahluwalia

机构 * Abertay University(阿伯泰大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文研究电商场景下知识图谱检索增强生成中神经检索-排序流水线的设计与比较,通过实验验证了优化配置在提升检索性能上的有效性。

Comments This manuscript is under review at the Springer journal Knowledge and Information Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22208 2026-02-27 cs.CV 57%

Solaris: Building a Multiplayer Video World Model in Minecraft

Solaris:在Minecraft中构建多玩家视频世界模型

Georgy Savva, Oscar Michel, Daohan Lu, Suppakit Waiwitlikhit, Timothy Meehan, Dhairya Mishra, Srivats Poddar, Jack Lu, Saining Xie

机构 * New York University(纽约大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 Solaris通过多玩家数据系统和分阶段训练方法,在Minecraft中构建了能模拟多视角观测的视频世界模型,提升了多代理交互的建模能力。

Comments Project website: https://solaris-wm.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏