arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-08-25 至 2026-08-25 共收录 53 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 53 篇

2608.23074 2026-08-25 cs.CV 新提交 92%

Grounding Isn't Knowing: Do VLMs Need Object Localization for Spatial Reasoning?

grounding 不等于知晓:视觉语言模型是否需要对象定位来进行空间推理?

Xiwei Liu, Yulong Li, Xinlin Zhuang, Xuhui Li, Zhixiang Lu, Haolin Yang, Imran Razzak, Yutong Xie

专题命中 视觉定位与Grounding :grounding(title,title_cn);LLaVA(summary_cn,abstract);vision-language model(abstract);分类 cs.CV

AI总结 本研究以 LLaVA-1.5 和 Qwen2.5-VL 为对象,通过多种可解释性工具揭示 VLMs 空间推理的机制,发现其空间关系预测无需精确对象定位,定位与空间推理共享早期通路但依赖部分不同通路。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21605 2026-08-25 cs.AI 新提交 90%

Generate in the Chart, Not on the Boundary: Function-Symbol Grounding for Hard Constraints in LTN-GANs

在图表中生成,而非在边界上:LTN-GAN中硬约束的函数符号 grounding

Nijesh Upreti, Vaishak Belle

机构 * The University of Edinburgh(爱丁堡大学)

专题命中 视觉定位与Grounding :grounding(title,title_cn);分类 cs.AI

AI总结 本研究针对LTN-GAN的硬约束问题,提出将公理以函数符号而非谓词 grounding 的方法,对比约束层发现其能保留余量分布,通过分辨率比R诊断 grounding 的学习能力,可生成真实有效样本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21744 2026-08-25 cs.HC 新提交 89%

CALM-BP: Observation-Matched Physiological Semantic Grounding for Non-Contact Blood Pressure Estimation

CALM-BP:用于非接触式血压估计的观测匹配生理语义 grounding

Haiyang Sun, Boyuan Gu, Yongjie Liu

专题命中 视觉定位与Grounding :grounding(title,title_cn)

AI总结 提出 CALM-BP 方法,构建观测匹配生理语义 grounding,基于 FlowBP-Set 数据集验证语言可通过组织生理观测助力非接触式血压估计。

Comments 17 pages, 3 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22950 2026-08-25 cs.CV 新提交 88%

WADE: A Reasoning-Annotated Benchmark for Multi-Instance Floating-Waste Grounding with Compact Vision-Language Models

WADE:面向紧凑视觉-语言模型的多实例漂浮垃圾定位推理标注基准

Md. Asaduzzaman Shuvo, Ahsan Farabi, Md. Abdul Ahad Minhaz, Mahedi Hasan, Israt Khandaker, Ibrahim Khalil Shanto, Muhammad Nomani Kabir

机构 * United International University(联合国际大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);分类 cs.CV

AI总结 针对内陆漂浮垃圾监测需求,研究推出WADE基准并评估6种VLMs,微调Qwen3-VL-2B可提升性能但仍有大量实例未被检测,为紧凑VLMs的漂浮垃圾定位提供挑战基准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16447 2026-08-25 cs.AI cs.RO 版本更新 87%

HaReCAP: Habitual-action Grounding for Recursive Large Language Model Agents

HaReCAP:面向递归大语言模型智能体的习惯性动作 grounding 方法

Shen Liu, Zhenguo Xu, Shaopu Wang, Yike Gao, Chunlei Wang

机构 * North China Institute of Computer System Engineering(华北计算机系统工程研究所) University of Science and Technology of China(中国科学技术大学) China Information Security Research Institute Co., Ltd.(中国信息安全研究院有限公司)

专题命中 视觉定位与Grounding :grounding(title,title_cn);分类 cs.AI

AI总结 HaReCAP 是针对 ReCAP 的低侵入性扩展,通过编译叶子反射规则减少长视距具身任务中 LLM 的重复调用,在 Robotouille 和 ALFWorld 上显著降低了 token 消耗。

Comments 15 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22586 2026-08-25 cs.CV cs.AI cs.CL 新提交 86%

Vision-Language Models for Occupational Physical Exposure Assessment: Estimating External Hand Forces in Manual Material Handling Tasks from RGB Video

用于职业体力暴露评估的视觉-语言模型:从RGB视频估算手工物料搬运任务中的外部手力

Mohammad Sadra Rajabi, Aanuoluwapo Ojelade, Sunwook Kim, Maury A. Nussbaum

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract,abstract_cn);分类 cs.CV、cs.AI

AI总结 本研究提出基于视觉-语言模型的流程,结合文本线索、视觉特征与已知负载质量,从RGB视频估算手工物料搬运的手力,验证了该方法的可行性及不同相机条件的性能差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19218 2026-08-25 cs.CL cs.AI cs.LG 版本更新 86%

Time-Series Retrieval for Grounding Multimodal Language Models in Remaining Useful Life Prediction

用于将多模态语言模型与剩余使用寿命关联的时间序列检索

Valeriu Dimidov, Raphaël Frank

机构 * University of Luxembourg(卢森堡大学)

专题命中 视觉定位与Grounding :grounding(title);MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出时间序列检索关联的多模态大语言模型框架,在C-MAPSS基准FD001分区实验显示,该方法可提升RUL预测的误差与稳定性,且效果随模型容量变化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22963 2026-08-25 cs.AI cs.CL 新提交 85%

Buried in Textual Debt: Context Pruning with Visual Evidence Preservation for MLLM Agents

被文本债务掩埋:面向多模态大语言模型智能体的保留视觉证据的上下文剪枝

Yuchen Huang, Sijia Li, Jun Zhang, Yi R. Fung

机构 * Hong Kong University of Science and Technology(香港科技大学)

专题命中 视觉定位与Grounding :MLLM(title,abstract_cn);grounding(abstract);multimodal large language model(abstract);分类 cs.AI

AI总结 针对多模态大语言模型智能体长轨迹中文本主导上下文、抑制视觉证据的问题,提出基于KL引导的SPARE框架,结合OPSD与SFT实现高效剪枝,在多步视觉工具使用基准中达到最优准确率与剪枝比例的权衡。

Comments 14 pages, 2 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22093 2026-08-25 cs.RO 新提交 85%

EndoNav: Semantic-to-Geometric Grounding for Language-Guided Robotic Endoscopic Examination

EndoNav:面向语言引导的机器人内窥镜检查的语义到几何的 grounding(语义几何关联)

Jecia Z. Y. Mao, Hisashi Ishida, Kathryn Jung, Masaru Ishii, Russell H. Taylor, Manish Sahu

专题命中 视觉定位与Grounding :grounding(title,title_cn)

AI总结 研究人员提出 EndoNav 框架,可将外科医生高级命令转换为患者特异性鼻窦解剖内的自主内窥镜可视化行为,其可视化效果接近住院外科医生水平,验证了该方法的可行性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22885 2026-08-25 cs.CV 新提交 84%

DRAgent: Discriminative Reasoning Agent for Referring Expression Segmentation

DRAgent:用于指代表达分割的判别推理智能体

Yujie Qi, Luyan Zhang

机构 * School of Computer Science and Technology, Hangzhou Dianzi University(杭州电子科技大学计算机科学与技术学院) Khoury College of Computer Sciences, Northeastern University(东北大学Khoury计算机科学学院)

专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 本文针对指代表达分割中MLLM单次坐标预测导致的定位偏差问题,提出DRAgent判别推理框架,通过两阶段目标选择与LoRA微调提升性能,在多个基准数据集上表现具竞争力。

Comments 5 pages, 3 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23474 2026-08-25 cs.CL cs.AI cs.CV 新提交 84%

What's the Catch? Evaluating Temporal Consistency in Vision-Language Models

有何玄机?评估视觉-语言模型的时间一致性

Marek Hradil, Danae Sánchez Villegas

机构 * University of Copenhagen(哥本哈根大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 本研究提出TimeCatch基准,发现视觉-语言模型可检测单帧异常但难整合跨帧信息,在时间异常检测上表现接近随机水平。

Comments 17 pages, ACL format

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22232 2026-08-25 cs.AI cs.CL cs.CV cs.MM 新提交 84%

Beyond What Meets the Eye: Unveiling Situational Illusions for Multimodal Large Language Models

超越表象:揭示多模态大语言模型的情境错觉

Zhiming Yang, Zhuoxi Xiong, Donglin Zhou, Wenjun Wei, Shiyao Cui, Jinqiao Shi

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 本文针对多模态大语言模型(MLLMs)面临的情境错觉问题,构建分类体系并推出MSIBench基准,发现27种模型配置均易受6类错觉影响,通过提示工程和监督微调最多提升性能20%。

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23302 2026-08-25 cs.CV 新提交 83%

Grounding Free-Form Instructions for Fashion Complementary Image Generation

基于自由形式指令的时尚互补图像生成的接地

Matteo Attimonelli, Claudio Pomo, Alessandro De Bellis, Danilo Danese, Dietmar Jannach, Tommaso Di Noia

机构 * Politecnico di Bari(巴里理工大学) Sapienza University of Rome(罗马大学) University of Klagenfurt(克拉根福大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV

AI总结 针对现有时尚互补图像生成基准的缺陷,该研究提出基于自由形式指令的新设置,用StyleFlow模型实现该任务,其生成的服装符合指令且风格连贯,同时降低了架构复杂度和推理成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23138 2026-08-25 cs.RO cs.AI cs.CV 新提交 81%

Pointing-VLA: Typed Spatial Grounding Interfaces for Vision-Language-Action Manipulation

Pointing-VLA:面向视觉-语言-动作操控的类型化空间定位接口

Xiwen Chen, Zelin Li, Zhiruo Zhou, Huiming Chen, Chenwei Wang, Xiaojun Zhu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

AI总结 Pointing-VLA是基于Embodied-R1的类型化空间读出接口,可提升VLA模型的机器人操控性能,在多任务评估中实现SOTA表现,还能高效迁移至其他机器人系统并提升真实机器人自主成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23065 2026-08-25 cs.CV cs.AI cs.CL cs.IR cs.MM 新提交 81%

Cultural Moment Benchmark: Evaluating Video Cultural Reasoning and Grounding in Southeast Asia

文化时刻基准:评估东南亚视频文化推理与定位

Burak Satar, Zhixin Ma, Cheng Yu-Tong, Huy Hoang Tran, Phuong Anh Nguyen, Chong-Wah Ngo

机构 * Singapore Management University(新加坡管理大学) VNU-HCM(胡志明市国家大学)

专题命中 视觉定位与Grounding :grounding(title);vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 该研究推出东南亚文化时刻基准CMB,通过三阶段评估视觉-语言模型的文化理解能力,发现模型跨阶段能力未完全级联,音频在非拉丁文字国家多为干扰,专家对邻国概念的识别也不及随机水平。

Comments Accepted to EMNLP 2026 Main Conference, this https URL (https://culturalmoment-benchmark.github.io/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21387 2026-08-25 cs.RO eess.SY 新提交 80%

Multimodal-Language-Model-Driven Interaction and Companionship for Service Robots in Elderly-Care Facilities

面向养老机构服务机器人的多模态语言模型驱动的交互与陪伴

Ching-Chieh Liu, Cong-Thanh Vu, Yen-Chen Liu

机构 * National Cheng Kung University (NCKU)(成功大学)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract)

AI总结 该研究针对养老机构服务机器人缺乏综合陪伴与安全能力的问题,提出整合主动视觉跟人、LLM语音交互及VLM安全监测的智能陪伴机器人系统,实验验证其跟随交互与跌倒检测的有效性。

Comments Accepted to the 2026 IEEE/ASME International Conference on Advanced Intelligent Mechatronics (AIM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23143 2026-08-25 cs.CV 新提交 79%

An end-to-end-trained vision-language model for native-language prostate pathology report generation

用于生成本土语言前列腺病理报告的端到端训练视觉-语言模型

Christian Grashei, Fabian Gülhan, Maximilian Legnar, Fabian Stögbauer, Cleo-Aron Weis, Carolin Mogler, Peter Schüffler

机构 * Technical University of Munich(慕尼黑工业大学) Munich Data Science Institute(慕尼黑数据科学研究所) Munich Center for Machine Learning(慕尼黑机器学习中心) University Hospital Heidelberg(海德堡大学医院) Heidelberg University(海德堡大学) Interdisciplinary Center for Scientific Computing (IWR)(跨学科科学计算中心(IWR))

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

AI总结 该研究提出语言独立的切片级视觉-语言框架,用自动化流水线生成17344对图像-文本对,实现德语前列腺病理报告生成,恶性肿瘤检测F1达96.2%,可支持机构用自有档案训练本土语言报告模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21878 2026-08-25 cs.CV 新提交 79%

ViSMoE: Visual-Aware Sparse Mixture-of-Experts for Embodied Referring Expression Grounding

ViSMoE:面向具身指代表达接地的视觉感知稀疏混合专家模型

Shuo Feng, Piji Li

机构 * College of Artificial Intelligence(人工智能学院) Nanjing University of Aeronautics and Astronautics(南京航空航天大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 针对现有具身指代表达接地方法无法区分视角与物体导致表示模糊的问题,提出带视觉感知路由策略的ViSMoE框架,在REVERIE和SOON数据集上性能优于现有SOTA方法。

Comments Accepted by ICANN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15382 2026-08-25 cs.AI cs.CL cs.IR q-bio.QM 版本更新 79%

Framework for Grounding Healthcare LLMs in a Causal Knowledge Graph: A Cardiovascular Example Pilot

将医疗大语言模型(LLM)基于因果知识图谱:框架、指标与心血管试点研究

Ummara Mumtaz, Aimen Noor, Awais Ahmed

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

AI总结 本研究提出以因果知识图谱为核心的医疗LLM评估框架,在心血管试点中验证其有效性,发现集成条件C4在因果推理相关指标上表现最优,未基于图的C1原始干预准确性最高但缺乏因果与证据基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22225 2026-08-25 cs.CV cs.AI 版本更新 79%

ExtrinSplat: Decoupling Geometry and Semantics for Open-Vocabulary Understanding in 3D Gaussian Splatting

ExtrinSplat:解耦几何与语义以实现3D高斯散射中的开放词汇理解

Jiayu Ding, Xinpeng Liu, Zhiyi Pan, Shiqiang Long, Ge Li

机构 * Guangdong Provincial Key Laboratory of Ultra High Definition Immersive Media Technology, Shenzhen Graduate School, Peking University(广东省超高清沉浸式媒体技术重点实验室,北京大学深圳研究生院) School of Computer Science and Technology, Tianjin University(天津大学计算机科学与技术学院) Guangdong Bohua UHD Innovation Center Co., Ltd.(广东博华超高清创新中心有限公司)

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 ExtrinSplat通过解耦几何与语义,利用视觉语言模型生成轻量文本假设,提升3D高斯散射中开放词汇物体选择和语义分割的性能和效率。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17177 2026-08-25 cs.SE 版本更新 71%

Grounding AI Agents in Contracts: An Empirical Evaluation of Spec-Driven Test Generation

将AI智能体建立在契约基础上:对规范驱动测试生成的实证评估

Michele Tufano, James McClure, José Cambronero, Runxiang Cheng, Sherry Y. Shi, Renyao Wei, Dorothy Chen, Franjo Ivančić, Livio Dalloro, Pat Rondon

专题命中 视觉定位与Grounding :grounding(title)

AI总结 本研究针对LLM智能体生成测试时易遗漏契约相关边界的问题,提出规范驱动测试生成方法,经Google生产缺陷评估,其缺陷检测率、分支覆盖率均优于基线,测试套件质量也显著更高。

Comments To appear in Proceedings of the 1st International Workshop on Specification-Driven Development Life Cycle (SpecOps 2026), co-located with SPLASH 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22883 2026-08-25 cs.CV 新提交 70%

FOVEA: Focused On-Demand Visual Evidence Adaptation for Cache-Friendly Multimodal Speculative Decoding

FOVEA:面向缓存友好型多模态推测解码的聚焦式按需视觉证据适配

Hengjie Zhu, Dayan Wu, Zihao Zhang, Xinze Liu, Jingxuan Yu, Peng Fu, Zheng Lin, Weiping Wang, Ding Wang

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

AI总结 FOVEA是一种缓存友好型多模态推测解码方法,通过动态检索视觉记忆子集提升草稿接受率,使多模态生成速度最高提升2.13倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21476 2026-08-25 cs.SE cs.CV 新提交 70%

From Subjective Judgments to Auditable Standards:Protocol-Guided AI Auditing of Website Redundancy

从主观判断到可审计标准:协议引导的网站冗余度AI审计

Ge Kong, Yongtong Cao

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

AI总结 本文提出CORA审计方法,通过多维度测量分离网站冗余的不同含义,在测试台上表现优于基线,对不达标工具弃权自动评分,是受控基准的可审计候选程序。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31192 2026-08-25 cs.CV 版本更新 70%

Specialist-Generalist Fusion with Outcome-Supervised Rationales for Deepfake Detection

语言训练深度伪造检测器的正则化能力

Benedikt Hopf, Zongwei Wu, Radu Timofte

机构 * Computer Vision Lab, CAIDAS, University of Würzburg(计算机视觉实验室,CAIDAS,乌尔姆大学)

专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 提出利用多模态大语言模型的双编码器架构和两阶段训练,通过语言正则化缓解过拟合,提升深度伪造检测的泛化性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22651 2026-08-25 cs.CL 新提交 67%

Iteration Without Elaboration: A Simple ReAct Architecture Suffices for Text-to-SQL Generation

无需复杂设计的迭代:简单的ReAct架构足以完成文本转SQL生成

Jian Lu, Haiwei Yu, Raymond M Xiong, Anru Zhang, Danyang Zhuo

机构 * Duke University(杜克大学)

专题命中 视觉定位与Grounding :grounding(abstract,abstract_cn)

AI总结 该研究针对现有文本转SQL系统复杂且延迟高的问题,提出简单的ReAct-SQL框架,在两个数据集上达到高准确率且运行速度提升8倍,验证了其有效性。

Comments 19 pages; preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21864 2026-08-25 cs.LG cs.AI cs.CV 新提交 67%

BioMed-Agent-RL: A Meta Learning, All You Need for Biomedical Applications

BioMed-Agent-RL:元学习,生物医学应用的一切所需

Md Asaduzzaman Jabin, Zihao Wu, Tianming Liu

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 该研究针对临床视觉大语言模型的缺陷,提出BioMed-Agent-RL医学智能体,整合多模态元学习与强化学习等技术,在多基准实验中准确率达约73%,较现有模型提升约5%,为临床智能体系统建立新标准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21067 2026-08-25 cs.CV cs.AI 版本更新 62%

AT-ViT: Area-Targeted Multi-View Vision Transformer with Cross-Attention and Multi-Scale Patching for Plant Trait Recognition in Herbarium Images

AT-ViT:面向标本馆图像植物性状识别的区域定向多视图视觉Transformer,融合交叉注意力与多尺度分块技术

Amani Sedrat, Takieddine Chehhat, Youcef Sklab, Hanane Ariouat, Abderrazak Sebaa, Eric Chenin, Jean-Daniel Zucker, Edi Prifti

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

AI总结 该研究针对标本馆图像植物性状识别的背景干扰问题,提出双分支视觉Transformer模型AT-ViT,通过交叉注意力与掩码引导分块加权机制提升植物区域注意力,在多任务中实现准确率与鲁棒性提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02592 2026-08-25 cs.CV cs.LG 版本更新 62%

H-OPD: Confidence Aware Heterogeneous Multi-Teacher Multimodal On-policy Distillation

H-OPD:置信感知异构多教师多模态在线策略蒸馏

Qixiang Yin, Huanjin Yao, Yuchen Cai, Jianghao Chen, Ziyi Wang, Min Yang, Fei Su, Zhicheng Zhao

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) ByteDance(字节跳动) USTC(中国科学技术大学) Beijing Key Laboratory of Network System and Network Culture(北京网络系统与网络文化重点实验室) Key Laboratory of Interactive Technology and Experience System, Ministry of Culture and Tourism(文化和旅游部互动技术与体验系统重点实验室) Zhongguancun Academy(中关村科学城)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.LG

AI总结 研究多模态推理的在线策略蒸馏问题,提出H-OPD框架,通过验证异构教师互补性,以令牌级教师仲裁取代任务或样本级路由,结合视觉与文本教师,在多基准测试中性能优越。

Comments EMNLP2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01101 2026-08-25 cs.LG cs.AI 版本更新 62%

Soft-NBCE: Entropy-Weighted Chunk Fusion for Long-Context

Soft-NBCE: 基于熵加权分块融合的长上下文处理

Shihao Ji, Mingyu Li, Zihui Song

机构 * Beijing Normal University(北京师范大学) Chunjiang Intelligence(春江智能)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 针对长上下文推理中硬选择策略导致语义碎片化的问题,提出Soft-NBCE,通过熵加权软融合和一致性蒸馏,在保持检索精度的同时提升多跳推理性能。

Comments Withdrawn due to serious concerns regarding the authenticity and accuracy of the listed authorship. The identity of one or more listed authors cannot presently be verified, and the author list may not represent distinct contributors. The manuscript is withdrawn pending institutional review

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06770 2026-08-25 cs.CV cs.AI 版本更新 62%

FlowExtract: Procedural Knowledge Extraction from Maintenance Flowcharts

FlowExtract:从维护流程图中提取过程知识

Guillermo Gil de Avalle, Laura Maruster, Eric Sloot, Christos Emmanouilidis

机构 * University of Groningen(格罗宁根大学) Philips Consumer Lifestyle B.V.(飞利浦消费生活方式有限公司)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 本文提出FlowExtract,通过分离元素检测与连接性重建,利用YOLOv8和EasyOCR提取标准流程图节点,结合新边检测方法实现高精度连接提取,优于视觉语言模型基线,为可查询的过程知识表示提供实用路径。

Comments Accepted to the 45th IFIP WG 5.7 International Conference on Advances in Production Management Systems (APMS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏