arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7473 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7473 篇

2512.21347 2025-12-29 cs.SE 50%

Understanding the Role of Large Language Models in Software Engineering: Evidence from an Industry Survey

理解大型语言模型在软件工程中的作用:来自行业调查的证据

Vítor Mateus de Brito, Kleinner Farias

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本文通过行业调查揭示大型语言模型在软件工程中的应用现状与影响,探讨其积极效果与潜在风险,提出需批判性使用LLM工具以提升开发效率与安全性。

Comments 4 Figures, 8 Tables, Text in Portuguese

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19959 2025-12-24 cs.CG 50%

A Comprehensive Guide to Mesh Simplification using Edge Collapse

基于边折叠的网格简化全面指南

Purva Kulkarni, Aravind Shankara Narayanan

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本文提供基于边折叠的网格简化方法的全面指南,涵盖理论基础、实现细节及实用技巧,帮助从业者和研究人员理解和应用该技术。

Comments 46 pages, 23 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19466 2025-12-23 cs.CY cs.CL cs.HC 50%

Epistemological Fault Lines Between Human and Artificial Intelligence

人类与人工智能之间的知识论断层

Walter Quattrociocchi, Valerio Capraro, Matjaž Perc

机构 * Department of Computer Science, Sapienza University of Rome, Rome, Italy Department of Psychology, University of Milan Bicocca, Milan, Italy Faculty of Natural Sciences Mathematics, University of Maribor, Maribor, Slovenia Community Healthcare Center Dr. Adolf Drolc Maribor, Maribor, Slovenia University College, Korea University, Seoul, Republic of Korea Department of Physics, Kyung Hee University, Seoul, Republic of Korea

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本文揭示大型语言模型与人类认知在知识生成机制上的结构性差异,指出LLMs是随机模式完成系统而非知识代理,并识别七种知识断层,对社会评估、治理及知识素养提出影响。

Comments 16 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22876 2025-12-23 cs.MA 50%

Cooperation as a Black Box: Conceptual Fluctuation and Diagnostic Tools for Misalignment in MAS

协作作为黑箱:多智能体系统中概念波动与对齐偏差的诊断工具

Shayak Nandi, Fernanda M. Eliott

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本文提出马赛克框架,用于诊断多智能体系统中概念波动与对齐偏差,强调术语一致性与道德基础以确保系统技术与道德对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18225 2025-12-23 cs.CL cs.SI 50%

GeoSense-AI: Fast Location Inference from Crisis Microblogs

GeoSense-AI:从危机推文快速推断位置

Deepit Sapru

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 GeoSense-AI通过领域调优的NLP和知识基础,从危机推文快速推断位置,提升应急响应效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17992 2025-12-23 cs.RO 50%

Unifying Deep Predicate Invention with Pre-trained Foundation Models

统一深度谓词发明与预训练基础模型

Qianwei Wang, Bowen Li, Zhanpeng Luo, Yifan Xu, Alexander Gray, Tom Silver, Sebastian Scherer, Katia Sycara, Yaqi Xie

机构 * Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所) Computer Science and Engineering Division, University of Michigan(密歇根大学计算机科学与工程系) Department of Computer Science, University of Pittsburgh(匹兹堡大学计算机科学系) Centaur AI Institute(Centaur人工智能研究所) Department of Electrical and Computer Engineering, Princeton University(普林斯顿大学电气与计算机工程系)

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 UniPred通过双层学习框架统一深度谓词发明与预训练基础模型,提升机器人任务的可扩展性和灵活性。

Comments 18 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00877 2025-12-19 eess.SY cs.CL cs.SY 50%

Verifiable Natural Language to Linear Temporal Logic Translation: A Benchmark Dataset and Evaluation Suite

可验证的自然语言到线性时序逻辑翻译:一个基准数据集和评估套件

William H English, Chase Walker, Dominic Simon, Sumit Kumar Jha, Rickard Ewetz

机构 * University of Florida(佛罗里达大学) Florida International University(佛罗里达国际大学)

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本文提出VLTL-Bench,一个用于评估可验证自然语言到线性时序逻辑翻译的基准数据集和评估套件。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03197 2025-12-18 cs.CL 50%

Explain with Visual Keypoints Like a Real Mentor! A Benchmark for Multimodal Solution Explanation

像真实导师一样用视觉关键点解释吧!一个多模态解决方案解释基准

Jaewoo Park, Jungyang Park, Dongju Jang, Jiwan Chung, Byungwoo Yoo, Jaewoo Shin, Seonjoon Park, Taehyeong Kim, Youngjae Yu

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本文提出多模态解决方案解释任务和ME2数据集,旨在评估模型识别视觉关键点并生成相关解释的能力,揭示当前LLM在数学视觉推理和教育应用中的不足。

Comments 14 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05310 2025-12-17 cs.CL cs.SI 50%

Listening Between the Lines: Decoding Podcast Narratives with Language Modeling

在字里行间倾听:利用语言模型解码播客叙事

Shreya Gupta, Ojasva Saxena, Arghodeep Nandi, Sarah Masud, Kiran Garimella, Tanmoy Chakraborty

机构 * Indian Institute Of Technology Delhi(印度理工学院德里分校) University of Copenhagen(哥本哈根大学) Rutgers University(罗格斯大学)

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本文提出一种基于 BERT 的方法,通过标注叙事框架与对话实体的关系,揭示播客中主题与框架的系统关联,提升对数字媒体影响的分析能力。

Comments 10 pages, 6 Figures, 5 Tables. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11297 2025-12-16 cs.CL 50%

LegalRikai: Open Benchmark -- Benchmark for Complex Japanese Corporate Legal Tasks

LegalRikai:开放基准——复杂日本企业法律任务的基准

Shogo Fujita, Yuji Naraki, Yiqing Zhu, Shinsuke Mori

机构 * LegalOn Technologies, Inc.(LegalOn技术公司) Kyoto University(京都大学)

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 LegalRikai基准通过四个复杂任务评估日本企业法律实践,揭示模型在文档编辑中的不足,并提出数据集评估框架以推动法律领域实践研究。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01948 2025-12-16 cs.CL 50%

How Far Are We from Genuinely Useful Deep Research Agents?

我们距离真正有用的深度研究代理还有多远?

Dingling Zhang, He Zhu, Jincheng Ren, Kangqi Song, Xinran Zhou, Boyu Feng, Shudong Liu, Jiabin Luo, Weihao Xie, Zhaohui Wang, Tianrui Qin, King Zhu, Yuqing Wang, Qianben Chen, Yuchen Eleanor Jiang, Wei Wang, Jiaheng Liu, Wangchunshu Zhou

机构 * OPPO AI Agent Team(OPPO 人工智能代理团队)

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本文提出FINDER基准和DEFT分类,揭示当前深度研究代理在证据整合与推理鲁棒性规划方面的不足。

Comments 34 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11039 2025-12-15 cs.SD 50%

Listening Between the Frames: Bridging Temporal Gaps in Large Audio-Language Models

在帧之间倾听:在大型音频-语言模型中弥合时间缺口

Hualei Wang, Yiming Li, Shuo Ma, Hong Liu, Xiangdong Wang

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 TimeAudio通过引入时间标记和绝对时间编码,提升大型音频-语言模型在时间定位和长音频理解上的能力,有效解决时间戳表示、架构和数据限制问题。

Comments Accepted by The Fortieth AAAI Conference on Artificial Intelligence (AAAI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17353 2025-12-11 cs.CE 50%

RoadBench: A Vision-Language Foundation Model and Benchmark for Road Damage Understanding

RoadBench: 一种用于道路损伤理解的视觉-语言基础模型和基准

Xi Xiao, Yunbei Zhang, Janet Wang, Lin Zhao, Yuxiang Wei, Hengjia Li, Yanshu Li, Xinyuan Song, Xiao Wang, Swalpa Kumar Roy, Hao Xu, Tianyang Wang

专题命中 视觉定位与Grounding :vision language model(abstract)

AI总结 RoadBench通过整合视觉与文本信息,提出RoadCLIP模型,显著提升道路损伤识别性能,为基础设施监测提供新基准。

Comments Accepted by WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05537 2025-12-08 cs.CL 50%

Automated Identification of Incidentalomas Requiring Follow-Up: A Multi-Anatomy Evaluation of LLM-Based and Supervised Approaches

自动识别需要随访的偶发瘤:基于LLM和监督方法的多解剖评估

Namu Park, Farzad Ahmed, Zhaoyi Sun, Kevin Lybarger, Ethan Breinhorst, Julie Hu, Ozlem Uzuner, Martin Gunn, Meliha Yetisgen

机构 * Department of Biomedical Informatics and Medical Education, University of Washington, Seattle, WA, USA(生物医学信息学与医学教育系,华盛顿大学,西雅图,华盛顿州,美国) Department of Information Sciences and Technology, George Mason University, Fairfax, VA, USA(信息科学与技术系,乔治·马歇尔大学,弗吉尼亚州,美国) Department of Radiology, Te Whatu Ora Health New Zealand, Te Toka Tumai Auckland, Auckland, New Zealand(放射学系,新西兰Te Whatu Ora健康机构,奥克兰,新西兰) Department of Radiology, School of Medicine, University of Washington, Seattle, WA, USA(放射学系,医学院,华盛顿大学,西雅图,华盛顿州,美国)

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本文提出了一种基于LLM和监督方法的多解剖评估,通过结构化病变标记和解剖学上下文提升偶发瘤检测性能,达到与人类专家相当的水平。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05428 2025-12-08 cs.SE 50%

Bita: A Conversational Assistant for Fairness Testing

Bita:一个用于公平性测试的对话助手

Keeryn Johnson, Cleyton Magalhaes, Ronnie de Souza Santos

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 Bita是一个基于大型语言模型的对话助手,旨在通过公平性视角帮助软件测试人员检测偏见、评估测试计划并生成公平性导向的测试章程,提供可重复的公平性测试解决方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03501 2025-12-04 cs.CY 50%

SocraticAI: Transforming LLMs into Guided CS Tutors Through Scaffolded Interaction

SocraticAI: 通过支架式交互将LLMs转化为指导性计算机科学导师

Karthik Sunil, Aalok Thakkar

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 SocraticAI通过结构化约束将LLMs整合到计算机科学教育中,培养学生战略性的人工智能交互能力。

Comments Presented in the Best Practices Track of COMPUTE 2025 (arXiv:2512.02349)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03289 2025-12-04 cs.HC 50%

DAWZY: A New Addition to AI powered "Human in the Loop" Music Co-creation

DAWZY:一种新的AI驱动的“人机协同”音乐共创工具

Aaron C Elkins, Sanchit Singh, Adrian Kieback, Sawyer Blankenship, Uyiosa Philip Amadasun, Aman Chadha

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 DAWZY是一种基于AI的开源音乐创作工具,通过自然语言交互和LLM代码生成,提升音乐制作效率与用户体验。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02710 2025-12-03 cs.CE 50%

Beyond N-grams: A Hierarchical Reward Learning Framework for Clinically-Aware Medical Report Generation

超越n-gram:一种面向临床意识的医疗报告生成层次奖励学习框架

Yuan Wang, Shujian Gao, Jiaxiang Liu, Songtao Jiang, Haoxiang Xia, Xiaotian Zhang, Zhaolu Kang, Yemin Wang, Zuozhu Liu

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本文提出HiMed-RL框架,通过层次化奖励学习提升医疗报告生成的临床质量与事实准确性。

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22479 2025-12-03 cs.CL 50%

NeLLCom-Lex: A Neural-agent Framework to Study the Interplay between Lexical Systems and Language Use

NeLLCom-Lex: 一种神经代理框架用于研究词汇系统与语言使用之间的相互作用

Yuqing Zhang, Ecesu Ürker, Tessa Verhoef, Gemma Boleda, Arianna Bisazza

机构 * Center for Language and Cognition, University of Groningen(语言与认知中心,格罗宁根大学) Department of Translation and Language Sciences, Universitat Pompeu Fabra(翻译与语言科学系,庞培法拉大学) Leiden Institute of Advanced Computer Science, Leiden University(莱顿高级计算机科学研究所,莱顿大学) Catalan Institution for Research and Advanced Studies (ICREA)(加泰罗尼亚研究与高级科学研究所(ICREA))

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 NeLLCom-Lex通过神经代理框架模拟词汇系统演变,研究词汇与语言使用间的相互作用及语义变化机制。

Comments Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05036 2025-12-03 cs.CL 50%

From Word Vectors to Multimodal Embeddings: Techniques, Applications, and Future Directions For Large Language Models

从词向量到多模态嵌入:大型语言模型的技术、应用与未来方向

Charles Zhang, Benji Peng, Xintian Sun, Qian Niu, Junyu Liu, Keyu Chen, Ming Li, Pohsun Feng, Ziqian Bi, Ming Liu, Yichao Zhang, Xinyuan Song, Cheng Fei, Caitlyn Heqi Yin, Lawrence KQ Yan, Hongyang He, Tianyang Wang

机构 * Georgia Institute of Technology(佐治亚理工学院) Simon Fraser University(西蒙弗雷泽大学) Kyoto University(京都大学) National Taiwan Normal University(台湾师范大学) Purdue University(普渡大学) The University of Texas at Dallas(德克萨斯大学达拉斯分校) Emory University(埃默里大学) Cornell University(康奈尔大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) The Hong Kong University of Science(香港科学大学) University of Liverpool(利物浦大学) University of Warwick(沃里克大学)

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本文综述了从词向量到多模态嵌入的发展,探讨了大型语言模型的技术、应用及未来方向。

Comments 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00657 2025-12-02 cs.LO 50%

Computational Paths Form a Weak ω-Groupoid

计算路径构成一个弱ω-群oid

Arthur F. Ramos, Tiago M. L. de Veras, Ruy J. G. B. de Queiroz, Anjolina G. de Oliveira

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本文通过计算路径构建弱 ω-群oid,提供显式的协变数据和归一化算法验证。

Comments 24 pages. Formalized in Lean 4

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00461 2025-12-02 cs.CY cs.CL 50%

Whose Personae? Synthetic Persona Experiments in LLM Research and Pathways to Transparency

谁的人设?LLM研究中的人设实验及透明化路径

Jan Batzner, Volker Stocker, Bingjun Tang, Anusha Natarajan, Qinhao Chen, Stefan Schmid, Gjergji Kasneci

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本文探讨了LLM研究中合成人设实验的代表性与生态效度问题,提出透明化检查表以提升评估的严谨性和实证性。

Comments Published at AAAI/ACM AIES 2025. Presented at NeurIPS 2025 Workshop Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling

Journal ref Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 8(1), 2025, 343-354

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22426 2025-12-01 physics.optics 50%

Universal convolution from wave dynamics: photonic processing and encryption in synthetic dimension

通用卷积来自波动力学:在合成维度中光子处理与加密

Xiaolong Su, Weiwei Liu, Ruiqian Cheng, Haoru Zhang, Xinyao Guo, He Huang, Chengzhi Qin, Peixiang Lu, Bing Wang

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本研究通过光子合成晶格中的波动力学实现通用卷积,提出了一种基于卷积的光学加密方法,并展示了高吞吐量的图像处理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21157 2025-11-27 cs.HC 50%

QuadStretcher: A Forearm-Worn Skin Stretch Display for Bare-Hand Interaction in AR/VR

QuadStretcher:一种可穿戴皮肤拉伸显示装置用于AR/VR中的裸手交互

Taejun Kim, Youngbo Aram Shim, Youngin Kim, Sunbum Kim, Jaeyeon Lee, Geehyuk Lee

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 QuadStretcher通过前臂皮肤拉伸显示技术,实现了裸手交互中的触觉反馈,提升了AR/VR沉浸感。

Comments ACM CHI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19226 2025-11-25 physics.med-ph 50%

In-vivo imaging with a low-cost MRI scanner and cloud data processing in low-resource settings

在低资源环境中使用低成本MRI扫描仪和云数据处理进行体内成像

Teresa Guallart-Naval, Robert Asiimwe, Patricia Tusiime, Mary A. Nassejje, Leo Kinyera, Lemi Robin, Maureen Nayebare, Luiz G. C. Santos, Marina Fernández-García, Lucas Swistunow, José M. Algarín, John Stairs, Michael Hansen, Ronald Amodoi, Andrew Webb, Joshua Harper, Steven J. Schiff, Johnes Obungoloch, Joseba Alonso

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本研究展示了在低资源环境中通过改进硬件和软件实现低成本MRI系统获得临床相关图像质量的可行性。

Comments 10 pages, 10 figures, comments welcome

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01588 2025-11-25 cs.SD eess.AS eess.SP 50%

Learning Perceptually Relevant Temporal Envelope Morphing

学习具有感知相关性的时序包络变形

Satvik Dixit, Sungjoon Park, Chris Donahue, Laurie M. Heller

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本文提出了一种基于感知指导的学习方法,用于生成具有时间中间性的音频包络变形,通过人类听觉研究和大规模数据集训练,优于现有方法。

Comments Accepted at WASPAA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02977 2025-11-25 cond-mat.stat-mech 50%

Open Problems within Nonextensive Statistical Mechanics

非广延统计力学中的开放问题

Kenric P. Nelson

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本文综述了非广延统计力学中的开放问题,包括熵与泛化熵的差异量化、参数q的物理解释、泛化乘积的定义改进、广义傅里叶变换的应用以及非广延熵的归一化重新审视,并提出形状参数可能用于定义系统统计复杂性。

Comments 19 pages; 1 figure; 2 tables; recommended for publication in Entropy, special issue Nonadditive Entropies and Nonextensive Statistical Mechanics

Journal ref Entropy 2024, 26(2), 118

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16677 2025-11-24 physics.ins-det cs.SY eess.SY 50%

Electromagnetic transients and failed upward leaders observed during lightning activity in an onshore wind farm

风力发电场中雷电活动观测到的电磁瞬变及失败的向上导体

Franjo Vukovic, Bozidar Filipovic-Grcic, Nina Stipetic, Bojan Franc

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 研究风力发电场中雷电活动引发的电磁瞬变及失败向上导体的观测与分析。

Comments 10 pages

Journal ref Electric power systems research, 251 (2026), 1-10

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15272 2025-11-20 quant-ph cs.CR cs.NI 50%

QADR: A Scalable, Quantum-Resistant Protocol for Anonymous Data Reporting

Nilesh Vyas, Konstantin Baier

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19594 2025-11-13 cs.CL 50%

Understanding and Leveraging the Expert Specialization of Context Faithfulness in Mixture-of-Experts LLMs

Jun Bai, Minghao Tong, Yang Liu, Zixia Jia, Zilong Zheng

机构 * State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI) School of Computer Science, Wuhan University(武汉大学计算机学院)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏