arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7409 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7409 篇

2602.11583 2026-02-13 cs.AI cs.LG 62%

The Five Ws of Multi-Agent Communication: Who Talks to Whom, When, What, and Why -- A Survey from MARL to Emergent Language and LLMs

多智能体通信的五个W:谁与谁沟通,何时沟通,沟通什么,为何沟通——从MARL到涌现语言和LLMs的综述

Jingdi Chen, Hanqing Yang, Zongjun Liu, Carlee Joe-Wong

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 本文综述了多智能体通信的五个W,探讨了从MARL到涌现语言和LLM的发展,分析了不同范式下的通信设计、权衡与挑战。

Comments Accepted at Transactions on Machine Learning Research (TMLR), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13779 2026-02-13 cs.CL cs.AI cs.LG 62%

NewsInterview: a Dataset and a Playground to Evaluate LLMs' Ground Gap via Informational Interviews

NewsInterview: 一个用于通过信息性访谈评估LLM地面间隙的数据集和游乐场

Alexander Spangher, Michael Lu, Sriya Jeslyn Kalyan, Hyundong Justin Cho, Weiyan Shi, Jonathan May

机构 * University of California, Berkeley(加州大学伯克利分校) University of Southern California(南加州大学) Information Sciences Institute(信息科学研究所) Northeastern University(东北大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 NewsInterview通过信息性访谈数据集和模拟环境,揭示LLMs在战略对话和多轮规划方面的能力不足,强调需提升其对话策略与说服力。

Comments Accepted at ACL 2025: https://aclanthology.org/2025.acl-long.1580/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10881 2026-02-12 cs.CL cs.AI cs.LG 62%

Diagnosing Structural Failures in LLM-Based Evidence Extraction for Meta-Analysis

诊断基于大型语言模型的证据提取在元分析中的结构性故障

Zhiyin Tan, Jennifer D'Souza

机构 * L3S Research Center, Leibniz University Hannover(莱比锡大学汉诺威分校L3S研究中心) TIB Leibniz Information Centre for Science(莱比锡信息中心科学与技术研究所)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 研究发现当前LLM在元分析中存在结构性缺陷,导致关联元组提取不可靠,需改进结构性和数值基础。

Comments Accepted at the 22nd Conference on Information and Research Science Connecting to Digital and Library Science (IRCDL 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10863 2026-02-12 cs.LG cs.AI 62%

ICA: Information-Aware Credit Assignment for Visually Grounded Long-Horizon Information-Seeking Agents

ICA: 信息感知的信用分配用于基于视觉的长周期信息寻求智能体

Cong Pang, Xuyu Feng, Yujie Yi, Zixuan Chen, Jiawei Hong, Tiankuo Yao, Nang Yuan, Jiapeng Luo, Lewei Lu, Xin Lou

机构 * Shanghai Jiao Tong University(上海交通大学) ShanghaiTech University(上海科技大学) Wuhan University(武汉大学) SenseTime Research(商汤科技研究院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 本文提出ICA方法,通过视觉快照和信息层面信用分配,提升开放网络环境中信息寻求智能体的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18936 2026-02-12 cs.LG cs.CV 62%

Revisit Visual Prompt Tuning: The Expressiveness of Prompt Experts

重新审视视觉提示调优:提示专家的表达性

Minh Le, Anh Nguyen, Huy Nguyen, Chau Nguyen, Anh Tran, Nhat Ho

机构 * Qualcomm AI Research(高通人工智能研究)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.LG

AI总结 本文提出VAPT,通过增强提示专家的表达性,在保持参数效率的同时,提升了视觉任务的性能表现。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00185 2026-02-11 cs.AI cs.LG 62%

Chunking Strategies for Multimodal AI Systems

多模态AI系统中的分块策略

Shashanka B R, Mohith Charan R, Seema Banu F

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 本文综述了多模态系统中分块策略的分类和技术分析,探讨了不同模态的数据处理方法及挑战。

Comments 50 pages, 5 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09040 2026-02-11 eess.AS cs.AI cs.LG cs.SD 62%

Soft Clustering Anchors for Self-Supervised Speech Representation Learning in Joint Embedding Prediction Architectures

用于联合嵌入预测架构中自监督语音表示学习的软聚类锚点

Georgios Ioannides, Adrian Kieback, Judah Goldfeder, Linsey Pang, Aman Chadha, Aaron Elkins, Yann LeCun, Ravid Shwartz-Ziv

机构 * Carnegie Mellon University(卡内基梅隆大学) New York University(纽约大学) James Silberrad Brown Center for AI(詹姆斯·西伯拉德·布朗人工智能中心) Columbia University(哥伦比亚大学) Northeastern University(东北大学) Stanford University(斯坦福大学) Amazon GenAI(亚马逊生成人工智能)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 GMM-Anchored JEPA通过软聚类锚点提升语音表示学习,实现ASR、情感识别和槽填充的性能提升。

Comments 15 pages, 5 figures. Code: github.com/gioannides/clustering-anchored-jepa

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22441 2026-02-10 cs.CV cs.AI 62%

Can NeRFs See without Cameras?

NeRF能否在没有摄像头的情况下看到?

Chaitanya Amballa, Sattwik Basu, Yu-Lin Wei, Zhijian Yang, Mehmet Ergezer, Romit Roy Choudhury

机构 * University of Illinois Urbana–Champaign(伊利诺伊大学厄巴纳-香槟分校) Amazon(亚马逊)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

AI总结 本文提出通过重新设计NeRF来利用多路径信号推断环境,应用于从稀疏Wi-Fi测量中推断室内平面图。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06592 2026-02-09 cs.CV cs.AI 62%

ProtoQuant: Quantization of Prototypical Parts For General and Fine-Grained Image Classification

ProtoQuant:用于通用和细粒度图像分类的原型部分量化

Mikołaj Janusz, Adam Wróbel, Bartosz Zieliński, Dawid Rymarczyk

机构 * Jagiellonian University, Faculty of Mathematics and Computer Science(杰洛尼莫夫大学数学与计算机科学系) Jagiellonian University, Doctoral School of Exact and Natural Sciences(杰洛尼莫夫大学精确与自然科学研究博士学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

AI总结 ProtoQuant通过潜在向量量化实现原型稳定性和可解释性,适用于通用和细粒度图像分类任务。

Comments Work under review. Code will be released upon acceptance

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04518 2026-02-05 cs.CY cs.AI cs.LG 62%

Learning the Value Systems of Agents with Preference-based and Inverse Reinforcement Learning

基于偏好和逆强化学习的学习代理价值系统

Andrés Holgado-Sánchez, Holger Billhardt, Alberto Fernández, Sascha Ossowski

机构 * Universidad Rey Juan Carlos(雷乌安卡洛斯大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种基于偏好和逆强化学习的方法,用于自动从观察和人类示范中学习代理的价值系统,以解决不同用户价值系统差异带来的协议一致性问题。

Comments 42 pages, 5 figures. Published in Journal of Autonomous Agents and Multi-Agent Systems

Journal ref Holgado-Sánchez, A., Billhardt, H., Fernández, A., Ossowski, S. Learning the value systems of agents with preference-based and inverse reinforcement learning. Autonomous Agents Multi-Agent Systems 40, 4 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26275 2026-02-05 cs.SE cs.AI cs.ET cs.LG cs.MA 62%

A Research Roadmap for Augmenting Software Engineering Processes and Software Products with Generative AI

生成人工智能增强软件工程过程和软件产品的研究路线图

Domenico Amalfitano, Andreas Metzger, Marco Autili, Tommaso Fulcini, Tobias Hey, Jan Keim, Patrizio Pelliccione, Vincenzo Scotti, Anne Koziolek, Raffaela Mirandola, Andreas Vogelsang

机构 * University of Naples "Federico II" Napoli Italy Ruhr Institute for Software Technology (paluno), University of Duisburg-Essen Essen Germany Karlsruhe Institute of Technology (KIT) Karlsruhe Germany Gran Sasso Science Institute (GSSI) L'Aquila Italy University of Naples "Federico II" Ruhr Institute for Software Technology (paluno), University of Duisburg-Essen Karlsruhe Institute of Technology (KIT) Gran Sasso Science Institute (GSSI)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 本文提出一个生成人工智能增强软件工程的路线图,通过多轮过程整合证据,系统分析生成人工智能对软件工程过程、方法和工具的影响,并确定未来研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01733 2026-02-05 cs.HC cs.AI cs.CV 62%

DISCOVER: Identifying Patterns of Daily Living in Human Activities from Smart Home Data

DISCOVER: 从智能家庭数据中识别日常生活的活动模式

Alexander Karpekov, Archith Iyer, Sourish Gunesh Dhekane, Sonia Chernova, Thomas Plötz

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

AI总结 DISCOVER通过自监督学习和可视化接口,从智能家庭数据中发现日常活动模式,以更细粒度和个性化的方式识别居民行为,为健康监测和认知退化研究提供新方法。

Comments v2: Re-submission. Under review at IMWUT

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15540 2026-02-04 cs.LG cs.AI cs.CL physics.data-an 62%

PRISM: Deriving a White-Box Transformer as a Signal-Noise Decomposition Operator via Maximum Coding Rate Reduction

PRISM:通过最大编码率减少原理推导出白盒变换器作为信号-噪声分解算子

Dongchen Huang

机构 * Institute of Physics, Chinese Academy of Sciences(中国科学院物理研究所)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 Prism通过几何构造实现白盒变换器,通过信号-噪声分解提升模型可解释性与性能

Comments 12 pages, 6 figures. Derives Transformer as a signal-noise decomposition operator via Maximizing Coding Rate Reduction. Identifies 'Attention Sink' as spectral resonance (Arnold Tongues) and proposes $π$-RoPE for dynamical stability

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20835 2026-02-02 cs.CV cs.AI 62%

Open-Vocabulary Functional 3D Human-Scene Interaction Generation

开放词汇功能3D人类-场景交互生成

Jie Liu, Yu Sun, Alpar Cseke, Yao Feng, Nicolas Heron, Michael J. Black, Yan Zhang

机构 * Meshcapade University of Amsterdam(阿姆斯特丹大学) Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究所) Stanford University(斯坦福大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 FunHSI通过功能驱动框架生成开放词汇任务下的功能正确3D人类-场景交互。

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01706 2026-02-02 cs.CL cs.AI cs.LG 62%

Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement

通过秩-2子空间解耦进行多步知识交互分析

Sekh Mainul Islam, Pepa Atanasova, Isabelle Augenstein

机构 * Department of Computer Science, University of Copenhagen, Copenhagen, Denmark(计算机科学系,哥本哈根大学,哥本哈根,丹麦)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 本文提出一种秩-2子空间方法,用于多步分析NLE中的知识交互,揭示PK和CK在不同生成中的对齐特性。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22164 2026-02-02 cs.CV cs.LG cs.RO 62%

Do Open-Vocabulary Detectors Transfer to Aerial Imagery? A Comparative Evaluation

开放词汇检测器能否迁移到航拍图像?一项比较评估

Christos Tsourveloudis

机构 * School of Electrical and Computer Engineering, National Technical University of Athens(电气与计算机工程学院,国家技术大学雅典)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.LG

AI总结 本文评估了开放词汇检测器在航拍图像上的迁移能力,发现语义混淆是主要瓶颈,且不同数据集表现差异显著。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00590 2026-01-30 cs.CL cs.AI cs.LG 62%

Wikontic: Constructing Wikidata-Aligned, Ontology-Aware Knowledge Graphs with Large Language Models

Wikontic: 构建与维基数据对齐、本体感知的知识图谱

Alla Chepurova, Aydar Bulatov, Mikhail Burtsev, Yuri Kuratov

机构 * Cognitive AI Systems Lab(认知人工智能系统实验室) Moscow Independent Research Institute of Artificial Intelligence(莫斯科独立人工智能研究所) London Institute for Mathematical Sciences(伦敦数学科学研究所)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 Wikontic通过多阶段流程构建维基数据对齐、本体感知的知识图谱,提升KG质量并实现高效构建。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21615 2026-01-30 cs.LG cs.AI 62%

Beyond Parameter Finetuning: Test-Time Representation Refinement for Node Classification

超越参数微调:用于节点分类的测试时间表示精调

Jiaxin Zhang, Yiqi Wang, Siwei Wang, Xihong Yang, Yu Shi, Xinwang Liu, En Zhu

机构 * National University of Defense Technology(国防科技大学) Intelligent Game and Decision Lab(智能游戏与决策实验室)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 本文提出TTReFT框架,通过改进测试时间适应方法,解决图神经网络在分布外测试中的性能下降问题,提升节点分类的准确性和实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13993 2026-01-30 eess.IV cs.AI cs.CV 62%

OrthoInsight: Rib Fracture Diagnosis and Report Generation Based on Multi-Modal Large Models

OrthoInsight:基于多模态大模型的肋骨骨折诊断与报告生成

Ningyong Wu, Jiangbo Zhang, Wenhong Zhao, Jinzhi Wang, Chenzhan Yu, Zhigang Xiu, Duwei Dai, Ziyu Xu, Yongli Yang

机构 * Organizational Management Department, School of Management, Xi’an Jiaotong University(管理学院组织管理部,西安交通大学) West China Longquan Hospital, Sichuan University(四川大学西部临床医学院) School of Electronic Science and Engineering, Xi’an Jiaotong University(西安交通大学电子科学与工程学院) Systems Engineering Institute, Xi’an Jiaotong University(西安交通大学系统工程研究院) Institute of Medical Artificial Intelligence, the Second Affiliated Hospital of Xi’an Jiaotong University(西安交通大学第二附属医院医学人工智能研究所) School of Human Settlements and Civil Engineering, Xi’an Jiaotong University(西安交通大学人居环境与土木工程学院) School of Life Science and Technology, Xi’an Jiaotong University(西安交通大学生命科学与技术学院)

专题命中 视觉定位与Grounding :LLaVA(abstract);分类 cs.CV、cs.AI

AI总结 OrthoInsight通过多模态大模型实现肋骨骨折的自动诊断与报告生成,结合CT图像分析与医学知识图谱,提升诊断效率和临床实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17673 2026-01-27 cs.CV cs.AI 62%

Uni-RS: A Spatially Faithful Unified Understanding and Generation Model for Remote Sensing

Uni-RS: 一种用于遥感的具有空间忠实性的统一理解和生成模型

Weiyu Zhang, Yuan Hu, Yong Li, Yu Liu

机构 * Institute of Remote Sensing and Geographic Information System, School of Earth and Space Sciences, Peking University(遥感与地理信息系统研究所,地球与空间科学学院,北京大学) Department of Civil and Environmental Engineering, The Hong Kong University of Science and Technology(土木与环境工程系,香港科学与技术大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

AI总结 Uni-RS通过空间布局规划、空间感知查询监督和图像描述空间布局变化,提升遥感文本到图像生成的空间忠实性,同时保持多模态理解任务的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14525 2026-01-22 cs.CL cs.AI cs.LG 62%

Towards Execution-Grounded Automated AI Research

面向执行导向的自动化AI研究

Chenglei Si, Zitong Yang, Yejin Choi, Emmanuel Candès, Diyi Yang, Tatsunori Hashimoto

机构 * Stanford University(斯坦福大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 本研究提出自动化执行器,通过执行反馈学习改进LLM预训练和后训练方法,验证了进化搜索在样本效率上的优势,但强化学习存在模式崩溃问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10111 2026-01-22 cs.CV cs.AI cs.CR 62%

Training-Free In-Context Forensic Chain for Image Manipulation Detection and Localization

无需训练的上下文取证链用于图像篡改检测与定位

Rui Chen, Bin Liu, Changtao Miao, Xinghao Wang, Yi Li, Tao Gong, Qi Chu, Nenghai Yu

专题命中 视觉定位与Grounding :MLLM(abstract);分类 cs.CV、cs.AI

AI总结 ICFC提出一种无需训练的多模态大语言模型框架,用于图像篡改检测与定位,通过可解释的推理流程实现高效且准确的图像分析。

Comments This version was uploaded in error and contains misleading information found in an early draft. The manuscript requires extensive and long-term revisions

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13075 2026-01-21 cs.LG cs.AI 62%

METIS: Mentoring Engine for Thoughtful Inquiry & Solutions

METIS:面向深入探究与解决方案的导师引擎

Abhinav Rajeev Kumar, Dhruv Trehan, Paras Chopra

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 METIS通过工具增强和阶段意识功能,帮助本科生从想法发展到论文,优于GPT-5和Claude Sonnet 4.5,在文档基础阶段表现更佳。

Comments 12 pages, 5 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09147 2026-01-21 cs.CV cs.AI 62%

SSVP: Synergistic Semantic-Visual Prompting for Industrial Zero-Shot Anomaly Detection

SSVP:协同语义-视觉提示用于工业零样本异常检测

Chenhao Fu, Han Fang, Xiuzheng Zheng, Wenbo Wei, Yonghua Li, Hao Sun, Xuelong Li

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 SSVP通过协同语义-视觉提示机制,提升工业零样本异常检测的细粒度感知能力,实现93.0%的图像AUROC和92.2%的像素AUROC。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12259 2026-01-21 cs.AI cs.CE cs.LG 62%

FutureX-Pro: Extending Future Prediction to High-Value Vertical Domains

FutureX-Pro: 将未来预测扩展到高价值垂直领域

Jiashuo Liu, Siyuan Chen, Zaiyuan Wang, Zhiyuan Zeng, Jiacheng Guo, Liang Hu, Lingyue Yin, Suozhi Huang, Wenxin Hao, Yang Yang, Zerui Cheng, Zixin Yao, Lingyue Yin, Haoxin Liu, Jiayi Cheng, Yuzhen Li, Zezhong Ma, Bingjie Wang, Bingsen Qiu, Xiao Liu, Zeyang Zhang, Zijian Liu, Jinpeng Wang, Mingren Yin, Tianci He, Yali Liao, Yixiao Tian, Zhenwei Zhu, Anqi Dai, Ge Zhang, Jingkai Liu, Kaiyuan Zhang, Wenlong Wu, Xiang Gao, Xinjie Chen, Zhixin Yao, Zhoufutu Wen, B. Aditya Prakash, Jose Blanchet, Mengdi Wang, Nian Si, Wenhao Huang

机构 * Hong Kong University of Science and Technology(香港科技大学) Georgia Institute of Technology(佐治亚理工学院) Stanford University(斯坦福大学) Princeton University(普林斯顿大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 FutureX-Pro通过扩展未来预测到金融、零售、公共健康和自然灾害等高价值垂直领域,评估代理LLMs在工业部署中的领域基础能力。

Comments 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12082 2026-01-21 cs.CV cs.AI 62%

Conditional Random Fields for Interactive Refinement of Histopathological Predictions

用于病理预测交互细化的条件随机场

Tiffanie Godelaine, Maxime Zanella, Karim El Khoury, Saïd Mahmoudi, Benoît Macq, Christophe De Vleeschouwer

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 本文提出HistoCRF框架,通过条件随机场细化病理预测,利用专家注释提升分类准确率,实验显示在无注释和少量注释情况下均取得显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11514 2026-01-19 cs.CV cs.LG 62%

ShapeR: Robust Conditional 3D Shape Generation from Casual Captures

ShapeR: 从随意捕获序列中生成鲁棒的条件3D形状

Yawar Siddiqui, Duncan Frost, Samir Aroudj, Armen Avetisyan, Henry Howard-Jenkins, Daniel DeTone, Pierre Moulon, Qirui Wu, Zhengqin Li, Julian Straub, Richard Newcombe, Jakob Engel

机构 * Meta Reality Labs Research(Meta现实实验室) Simon Fraser University(西蒙弗雷泽大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.LG

AI总结 ShapeR通过整合视觉-惯性SLAM、3D检测和视觉-语言模型,从随意捕获的序列中生成鲁棒的条件3D形状,实验显示其在Chamfer距离上优于现有方法2.7倍。

Comments Project Page: http://facebookresearch.github.io/ShapeR Video: https://www.youtube.com/watch?v=EbY30KAA55I

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03585 2026-01-19 cs.CV cs.AI cs.CL 62%

Causal-SAM-LLM: Large Language Models as Causal Reasoners for Robust Medical Segmentation

Causal-SAM-LLM:大型语言模型作为因果推理器用于稳健的医学分割

Tao Tang, Shijie Xu, Jionglong Su, Zhixiang Lu

机构 * City University of Hong Kong, Hong Kong, China(香港城市大学) Xi’an Jiaotong-liverpool University, Suzhou, China(西安交通大学利物浦大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 Causal-SAM-LLM利用大型语言模型作为因果推理器,提升医学图像分割的鲁棒性和泛化能力。

Comments Accepted by IEEE ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09858 2026-01-16 cs.CL cs.AI cs.LG 62%

OUTLINEFORGE: Hierarchical Reinforcement Learning with Explicit States for Scientific Writing

OUTLINEFORGE: 基于显式状态的分层强化学习用于科学写作

Yilin Bao, Ziyao He, Zayden Yang

机构 * UC San Diego(圣迭戈大学) Ohio State University(俄亥俄州立大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 OUTLINEFORGE通过分层强化学习和显式状态管理,提升科学写作的全局结构和引用一致性,改进文档规划和事实准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03910 2026-01-14 math.RT cs.AI cs.LG 62%

An Algebraic Representation Theorem for Linear GENEOs in Geometric Machine Learning

线性GENEOs在几何机器学习中的代数表示定理

Francesco Conti, Patrizio Frosini, Nicola Quercioli

机构 * Inria Sophia Antipolis(法国 Sophia Antipolis 研究所) University of Pisa(比萨大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 本文提出新的代数表示定理,用于描述在不同数据空间之间作用的线性GENEOs,通过广义T-置换测度实现,扩展了几何机器学习中对称性编码的应用。

详情

展开后加载摘要…

URL PDF HTML 收藏