arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-01-21 至 2026-01-21 共收录 31 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 31 篇

2601.13440 2026-01-21 cs.CV 83%

Analyzing VLM-Based Approaches for Anomaly Classification and Segmentation

分析基于视觉语言模型的异常分类和分割方法

Mohit Kakda, Mirudula Shri Muthukumaran, Uttapreksha Patel, Lawrence Swaminathan Xavier Prince

机构 * Northeastern University(东北大学)

专题命中 视觉定位与Grounding :VLM(title,abstract);vision-language model(abstract);分类 cs.CV

AI总结 本文分析了基于视觉语言模型的异常分类和分割方法,探讨了其架构范式、对齐策略及性能评估,为工业质量控制提供了方法选择和未来研究方向的指导。

Comments 10 pages,4 images

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07413 2026-01-21 cs.CV 83%

REF-VLM: Triplet-Based Referring Paradigm for Unified Visual Decoding

基于三元组的引用范式用于统一视觉解码

Yan Tai, Luhao Zhu, Yunan Ding, Yiying Dong, Guangtao Zhai, Xiaohong Liu, Guodong Guo

机构 * School of Computer Science, Shanghai Jiao Tong University, Shanghai, 200240, China(上海交通大学计算机科学学院) Ningbo Institute of Digital Twin, Eastern Institute of Technology, Ningbo, China(宁波数字孪生研究院) School of Information Science and Electronic Engineering, Shanghai Jiao Tong University, Shanghai, 200240, China(上海交通大学信息科学与电子工程学院)

专题命中 视觉定位与Grounding :VLM(title,abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 REF-VLM通过引入基于三元组的引用范式,提升多任务视觉解码的性能和适应性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19825 2026-01-21 cs.LG cs.AI cs.DB 81%

Position: Foundation Models for Tabular Data within Systemic Contexts Need Grounding

位置:在系统性情境中为表格数据构建的基础模型需要接地

Tassilo Klein, Johannes Hoffart

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出语义链接表和SLT基础模型,通过双阶段训练实现表格数据在操作上下文中的接地,强调操作知识对自主代理的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13634 2026-01-21 cs.AI 79%

V2P: Visual Attention Calibration for GUI Grounding via Background Suppression and Center Peaking

V2P: 通过背景抑制和中心峰值校准实现GUI定位的视觉注意力校准

Jikai Chen, Long Chen, Dong Wang, Qinglin Su, Zhixuan Chu, Bingguang Hao, Leilei Gan, Chenyi Zhuang, Jinjie Gu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

AI总结 V2P通过背景抑制和中心峰值校准提升GUI元素定位精度,解决注意力漂移和中心边缘区分问题,实现高精度GUI接地任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12260 2026-01-21 cs.AI 77%

Docs2Synth: A Synthetic Data Trained Retriever Framework for Scanned Visually Rich Documents Understanding

Docs2Synth: 一种用于扫描视觉丰富文档理解的合成数据训练检索框架

Yihao Ding, Qiang Sun, Puzhen Wu, Sirui Li, Siwen Luo, Wei Liu

机构 * University of Western Australia(西澳大学) The University of Hong Kong(香港大学) Murdoch University(默多克大学)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.AI

AI总结 Docs2Synth通过合成监督框架实现私有和低资源领域文档理解,利用检索引导推理提升接地能力和领域泛化,无需人工标注。

Comments Accepted at WWW 2026 Demo Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14052 2026-01-21 cs.CV 74%

Vision Also You Need: Navigating Out-of-Distribution Detection with Multimodal Large Language Model

视觉也需要:利用多模态大语言模型进行分布外检测导航

Haoran Xu, Yanlin Liu, Zizhao Tong, Jiaze Li, Kexue Fu, Yuyang Zhang, Longxiang Gao, Shuaiguang Li, Xingyu Li, Yanran Xu, Changwei Wang

机构 * Zhejiang University(浙江大学) Tsinghua University(清华大学) University of Chinese Academy of Sciences(中国科学院大学) Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences)(教育部计算电力网络与信息安全重点实验室,山东计算机科学中心(国家超算中心济南中心),齐鲁工业大学(山东科学院)) Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science(山东省计算电力互联网与服务计算重点实验室,山东省计算机科学基础研究中心) University of Electronic Science and Technology of China(电子科技大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) RWTH Aachen University(亚琛工业大学)

专题命中 视觉定位与Grounding :multimodal large language model(title);分类 cs.CV

AI总结 本文提出MM-OOD方法,利用多模态大语言模型的推理能力,通过多轮对话增强分布外检测,提升近远OOD任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13142 2026-01-21 cs.CV cs.AI cs.CL 73%

TVWorld: Foundations for Remote-Control TV Agents

TVWorld: 电视遥控的基石

Zhantao Ma, Quanfeng Lu, Shuai Zhong, Dahai Yu, Ping Luo, Michael K. Ng

机构 * The University of Hong Kong(香港大学) Hong Kong Baptist University(香港 Baptist 大学) TCL Corporate Research (Hong Kong) Co., Ltd(TCL 香港企业研究有限公司)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 TVWorld提出了一种基于图的电视导航抽象,开发了TVWorld-N和TVWorld-G两个基准测试,通过拓扑感知训练框架TVTheseus实现了68.3%的成功率,超越现有基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06243 2026-01-21 cs.CL cs.AI 72%

CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning

CoT Referring: 通过 grounded 推理改进指称表达任务

Qihua Dong, Luis Figueroa, Handong Zhao, Kushal Kafle, Jason Kuen, Zhihong Ding, Scott Cohen, Yun Fu

机构 * Adobe Research(Adobe研究院) Northeastern University(东北大学)

专题命中 视觉定位与Grounding :MLLM(abstract,comments);multimodal large language model(abstract);分类 cs.AI

AI总结 通过 grounded 推理改进指称表达任务,提出CoT Referring方法,提升多模态大语言模型在复杂指称场景中的性能。

Comments MLLM, Referring Expression Segmentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14259 2026-01-21 cs.CV 70%

ManipShield: A Unified Framework for Image Manipulation Detection, Localization and Explanation

ManipShield: 一种用于图像篡改检测、定位和解释的统一框架

Zitong Xu, Huiyu Duan, Xiaoyu Wang, Zhaolin Cai, Kaiwei Zhang, Qiang Hu, Jing Liu, Xiongkuo Min, Guangtao Zhai

机构 * Institute of Image Communication and Network Engineering, Shanghai Jiao Tong University(上海交通大学图像通信与网络工程研究所) University of Electronic and Science Technology of China(电子科技大学) Tianjin University(天津大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

AI总结 ManipShield基于多模态大语言模型,通过对比学习LoRA微调和任务特定解码器,实现图像篡改的统一检测、定位和解释。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14227 2026-01-21 cs.SD 67%

Transformer Architectures for Respiratory Sound Analysis and Multimodal Diagnosis

用于呼吸声分析和多模态诊断的Transformer架构

Theodore Aptekarev, Vladimir Sokolovsky, Gregory Furman

机构 * Ben Gurion University of the Negev(本· Gurion 农业大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract)

AI总结 本文提出基于Transformer的AST和VLM模型,用于呼吸声分析和多模态诊断,AST在哮喘检测中达到97%准确率,VLM整合临床信息提升诊断能力。

Comments 7 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12486 2026-01-21 cs.HC 67%

A Multimodal Assistive System for Product Localization and Retrieval for People who are Blind or have Low Vision

一种多模态辅助系统,用于帮助盲人或低视力者进行产品定位与检索

Ligao Ruan, Giles Hamilton-Fletcher, Mahya Beheshti, Todd E Hudson, Maurizio Porfiri, John-Ross Rizzo

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract)

AI总结 本文提出一种多模态辅助系统,通过目标检测与视觉-语言模型结合,帮助盲人或低视力者自主完成产品定位与检索,提升其自主性和控制感。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11968 2026-01-21 cs.MM cs.SD eess.AS 67%

MuseAgent-1: Interactive Grounded Multimodal Understanding of Music Scores and Performance Audio

MuseAgent-1: 交互式 grounded 多模态理解音乐谱面与表演音频

Qihao Zhao, Yunqi Cao, Yangyu Huang, Hui Yi Leong, Fan Zhang, Kim-Hui Yap, Wei Hu

机构 * Nanyang Technological University(南洋理工大学) Beijing University of Chemical Technology(北京化工大学) Microsoft(微软公司) University of Chicago(芝加哥大学)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract)

AI总结 MuseAgent-1 是一个专注于音乐的多模态代理,通过结构化符号表示和多步骤推理,提升对音乐谱面和表演音频的交互式理解能力。

Comments Tech Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13075 2026-01-21 cs.LG cs.AI 62%

METIS: Mentoring Engine for Thoughtful Inquiry & Solutions

METIS:面向深入探究与解决方案的导师引擎

Abhinav Rajeev Kumar, Dhruv Trehan, Paras Chopra

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 METIS通过工具增强和阶段意识功能,帮助本科生从想法发展到论文,优于GPT-5和Claude Sonnet 4.5,在文档基础阶段表现更佳。

Comments 12 pages, 5 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09147 2026-01-21 cs.CV cs.AI 62%

SSVP: Synergistic Semantic-Visual Prompting for Industrial Zero-Shot Anomaly Detection

SSVP:协同语义-视觉提示用于工业零样本异常检测

Chenhao Fu, Han Fang, Xiuzheng Zheng, Wenbo Wei, Yonghua Li, Hao Sun, Xuelong Li

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 SSVP通过协同语义-视觉提示机制,提升工业零样本异常检测的细粒度感知能力,实现93.0%的图像AUROC和92.2%的像素AUROC。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12259 2026-01-21 cs.AI cs.CE cs.LG 62%

FutureX-Pro: Extending Future Prediction to High-Value Vertical Domains

FutureX-Pro: 将未来预测扩展到高价值垂直领域

Jiashuo Liu, Siyuan Chen, Zaiyuan Wang, Zhiyuan Zeng, Jiacheng Guo, Liang Hu, Lingyue Yin, Suozhi Huang, Wenxin Hao, Yang Yang, Zerui Cheng, Zixin Yao, Lingyue Yin, Haoxin Liu, Jiayi Cheng, Yuzhen Li, Zezhong Ma, Bingjie Wang, Bingsen Qiu, Xiao Liu, Zeyang Zhang, Zijian Liu, Jinpeng Wang, Mingren Yin, Tianci He, Yali Liao, Yixiao Tian, Zhenwei Zhu, Anqi Dai, Ge Zhang, Jingkai Liu, Kaiyuan Zhang, Wenlong Wu, Xiang Gao, Xinjie Chen, Zhixin Yao, Zhoufutu Wen, B. Aditya Prakash, Jose Blanchet, Mengdi Wang, Nian Si, Wenhao Huang

机构 * Hong Kong University of Science and Technology(香港科技大学) Georgia Institute of Technology(佐治亚理工学院) Stanford University(斯坦福大学) Princeton University(普林斯顿大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 FutureX-Pro通过扩展未来预测到金融、零售、公共健康和自然灾害等高价值垂直领域,评估代理LLMs在工业部署中的领域基础能力。

Comments 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12082 2026-01-21 cs.CV cs.AI 62%

Conditional Random Fields for Interactive Refinement of Histopathological Predictions

用于病理预测交互细化的条件随机场

Tiffanie Godelaine, Maxime Zanella, Karim El Khoury, Saïd Mahmoudi, Benoît Macq, Christophe De Vleeschouwer

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 本文提出HistoCRF框架,通过条件随机场细化病理预测,利用专家注释提升分类准确率,实验显示在无注释和少量注释情况下均取得显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18817 2026-01-21 cs.CV 57%

Disc3D: Automatic Curation of High-Quality 3D Dialog Data via Discriminative Object Referring

Disc3D:通过判别性对象指称自动校准高质量3D对话数据

Siyuan Wei, Chunjie Wang, Xiao Liu, Xiaosheng Yan, Zhishan Zhou, Rui Huang

机构 * PICO, ByteDance, Beijing(字节跳动北京研究院) Tsinghua University(清华大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 Disc3D通过自动化流程生成高质量3D对话数据,解决视角和对象指称模糊性问题,提升多模态大语言模型性能。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13207 2026-01-21 cs.CV 57%

GTPred: Benchmarking MLLMs for Interpretable Geo-localization and Time-of-capture Prediction

GTPred:评估多模态大语言模型在可解释地理定位和拍摄时间预测中的基准测试

Jinnao Li, Zijian Chen, Tingzhu Chen, Changbo Wang

机构 * School of Computer Science and Technology, East China Normal University(东华大学计算机科学与技术学院) Institute of Image Communication and Information Processing, Shanghai Jiao Tong University(上海交通大学图像通信与信息处理研究院) School of Humanities, Shanghai Jiao Tong University(上海交通大学人文学院) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 视觉定位与Grounding :MLLM(abstract);分类 cs.CV

AI总结 GTPred是一个新的地理-时间预测基准测试,通过评估多模态大语言模型在可解释地理定位和拍摄时间预测中的表现,揭示了当前模型在世界知识和时空推理方面的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12939 2026-01-21 cs.RO cs.AI eess.SP 57%

Active Inference-Driven World Modeling for Adaptive UAV Swarm Trajectory Design

基于主动推断的世界建模用于自适应无人机蜂群轨迹设计

Kaleem Arshid, Ali Krayani, Lucio Marcenaro, David Martin Gomez, Carlo Regazzoni

机构 * University of Genoa(热那亚大学) Carlos III University Madrid(马德里卡洛斯三世大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文提出基于主动推断的世界建模框架,用于自适应无人机蜂群轨迹设计,通过整合概率推理和自我学习实现动态环境下的适应性响应。

Comments This paper has been accepted for presentation at the 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (IEEE ICASSP 2026) Workshop: 'Multi-Modal Signal Processing and AI for Communications and Sensing in 6G and Beyond (MuSiC-6GB)'

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12839 2026-01-21 cs.LG q-fin.RM 57%

Knowledge-Integrated Representation Learning for Crypto Anomaly Detection under Extreme Label Scarcity; Relational Domain-Logic Integration with Retrieval-Grounded Context and Path-Level Explanations

基于知识整合的表示学习用于在极端标签稀缺下的加密异常检测;关系领域逻辑整合与检索基础的上下文和路径级解释

Gyuyeon Na, Minjung Park, Soyoun Kim, Jungbin Shin, Sangmi Chai

机构 * AI and Business Analytics, Ewha Womans University, Seoul, Republic of Korea(爱媛女子大学人工智能与商务分析系) Department of Business Administration, Kumoh National Institute of Technology, Gumi, Republic of Korea(Kumoh国立技术大学商业管理系) Coretrustlink, Seoul, Republic of Korea(Coretrustlink)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

AI总结 本文提出RDLI框架,通过整合领域逻辑与检索上下文,提升加密异常检测在极端标签稀缺下的准确性和可解释性。

Comments Gyuyeon Na, Minjung Park, Soyoun Kim contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12468 2026-01-21 cs.CV 57%

DCAC: Dynamic Class-Aware Cache Creates Stronger Out-of-Distribution Detectors

DCAC: 动态类感知缓存创建更强的分布外检测器

Yanqi Wu, Qichao Chen, Runhe Lai, Xinhua Lu, Jia-Xin Zhuang, Zhilin Zhao, Wei-Shi Zheng, Ruixuan Wang

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 DCAC通过动态类感知缓存提升分布外检测性能,减少误报率

Comments 9 pages, 9 figures, Accepted by AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12346 2026-01-21 cs.CV 57%

MMDeepResearch-Bench: A Benchmark for Multimodal Deep Research Agents

MMDeepResearch-Bench: 一个多模态深度研究代理的基准

Peizhou Huang, Zixuan Zhong, Zhongwei Wan, Donghao Zhou, Samiul Alam, Xin Wang, Zexin Li, Zhihao Dou, Li Zhu, Jing Xiong, Chaofan Tao, Yan Xu, Dimitrios Dimitriadis, Tuo Zhang, Mi Zhang

机构 * OSU(俄亥俄州立大学) Amazon(亚马逊公司) UMich(密歇根大学) UCL(伦敦大学学院) CUHK(香港中文大学) UCR(加州大学尔湾分校) CWRU(克里夫兰医学中心) HKU(香港大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 MMDeepResearch-Bench提出一个多模态深度研究代理的基准,强调报告式合成与引用证据的结合,揭示多模态完整性对深度研究代理的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10369 2026-01-21 cs.CV 57%

Fine-Grained Human Pose Editing Assessment via Layer-Selective MLLMs

通过层选择性MLLMs实现细粒度人体姿态编辑评估

Ningyu Sun, Zhaolin Cai, Zitong Xu, Peihang Chen, Huiyu Duan, Yichao Yan, Xiongkuo Min, Xiaokang Yang

机构 * Institute of Image Communication and Network Engineering, Shanghai Jiao Tong University(图像通信与网络工程研究所,上海交通大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

AI总结 本文提出HPE-Bench基准和基于层选择性MLLMs的统一框架,用于细粒度评估人体姿态编辑的真实性和质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12227 2026-01-21 cs.LG 57%

Learning Longitudinal Health Representations from EHR and Wearable Data

从电子健康记录和可穿戴数据中学习纵向健康表示

Yuanyun Zhang, Han Zhou, Li Feng, Yilin Hong, Shi Li

机构 * University of the Chinese Academy of Sciences, Columbia University(中国科学院大学,哥伦比亚大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

AI总结 本文提出一种多模态基础模型,通过联合表示电子健康记录和可穿戴数据,提升纵向健康预测的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12010 2026-01-21 cs.CV 57%

SMc2f: Robust Scenario Mining for Robotic Autonomy from Coarse to Fine

SMc2f: 从粗到细的机器人自主性鲁棒场景挖掘

Yifei Chen, Ross Greer

机构 * Department of Computer Science at Xi’an University of Technology(西安理工大学计算机科学系) department of Computer Science & Engineering at the University of California, Merced(加州大学默塞德分校计算机科学与工程系)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 SMc2f通过从粗到细的流程,利用视觉语言模型和文本-轨迹对比学习提升机器人场景挖掘的鲁棒性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10360 2026-01-21 cs.CV 57%

FaceXBench: Evaluating Multimodal LLMs on Face Understanding

FaceXBench: 评估多模态大语言模型在面部理解上的能力

Kartik Narayan, Vibashan VS, Vishal M. Patel

机构 * Department of Computer Science, Johns Hopkins University(计算机科学系,约翰霍普金斯大学) Department of Electrical and Computer Engineering, Johns Hopkins University(电气与计算机工程系,约翰霍普金斯大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

AI总结 FaceXBench通过5000个多模态问题评估MLLMs在面部理解上的能力,揭示了现有模型在复杂任务中的不足。

Comments Accepted in IEEE T-BIOM. Project Page: https://kartik-3004.github.io/facexbench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14081 2026-01-21 cs.SE 50%

Feature-Aware Test Generation for Deep Learning Models

面向深度学习模型的特征感知测试生成

Xingcheng Chen, Oliver Weissl, Andrea Stocco

专题命中 视觉定位与Grounding :vision-language model(abstract)

AI总结 本文提出Detect框架,通过特征感知扰动生成高质量测试用例,揭示模型中的捷径行为和未被准确度指标捕捉的bug,提升深度学习模型的可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14007 2026-01-21 cs.CL 50%

BACH-V: Bridging Abstract and Concrete Human-Values in Large Language Models

BACH-V: 联结抽象与具体的人类价值观在大语言模型中

Junyu Zhang, Yipeng Kang, Jiong Guo, Jiayu Zhan, Junqi Wang

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 BACH-V研究通过探测和引导方法揭示大语言模型中抽象与具体价值观的联结机制,发现其能稳定锚定抽象价值观以影响具体决策。

Comments 34 pagess, 16 figures, 6 tables, submitted to ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13882 2026-01-21 cs.CL 50%

OpenLearnLM Benchmark: A Unified Framework for Evaluating Knowledge, Skill, and Attitude in Educational Large Language Models

OpenLearnLM基准:一个评估教育大型语言模型知识、技能和态度的统一框架

Unggi Lee, Sookbun Lee, Heungsoo Choi, Jinseo Lee, Haeun Park, Younghoon Jeon, Sungmin Cho, Minju Kang, Junbo Koh, Jiyeong Bae, Minwoo Nam, Juyeon Eun, Yeonji Jung, Yeil Jeong

机构 * Chosun University(chosun大学) Korea University(韩国大学) Ewha Womans University(成均馆大学) Korea Institute for Curriculum and Evaluation(韩国课程评价院) Seoul National University(首尔国立大学) Texas A&M University(德克萨斯农工大学) Indiana University Bloomington(印第安纳大学布卢明顿分校)

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 OpenLearnLM基准通过统一框架评估教育大型语言模型的知识、技能和态度,揭示不同模型在多维度上的能力差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13297 2026-01-21 q-bio.NC 50%

Multifaceted neural representation of words in naturalistic language

自然语言中词语的多维神经表征

Xuan Yang, Chuanji Gao, Cheng Xiao, Nicholas Riccardi, Rutvik H. Desai

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本研究通过心理语言学建模与fMRI结合,揭示了词语多维属性在自然语言理解中的神经表征,发现八个潜在维度支持不同的认知功能。

Comments 65 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏