Seeing isn't Hearing: Benchmarking Vision Language Models at Interpreting Spectrograms
Tyler Loakman, Joseph James, Chenghua Lin
机构
*
Department of Computer Science, University of Sheffield, UK(谢菲尔德大学计算机科学系)
;
Department of Computer Science, University of Manchester, UK(曼彻斯特大学计算机科学系)
专题命中
评测与基准
:language model(title,abstract);large language model(abstract);分类 cs.CL
Integrating Genomics into Multimodal EHR Foundation Models
Jonathan Amar, Edward Liu, Alessandra Breschi, Liangliang Zhang, Pouya Kheradpour, Sylvia Li, Lisa Soleymani Lehmann, Alessandro Giulianelli, Matt Edwards, Yugang Jia, David Nola, Raghav Mani, Pankaj Vats, Jesse Tetreault, T. J. Chen, Cory Y. McLean
机构
*
Verily Life Sciences(Verily生命科学公司)
;
Nvidia(英伟达公司)
;
Google(谷歌公司)
An Evaluation of Representation Learning Methods in Particle Physics Foundation Models
Michael Chen, Raghav Kansal, Abhijith Gandrakota, Zichun Hao, Jennifer Ngadiuba, Maria Spiropulu
机构
*
Division of Physics, Mathematics and Astronomy(物理、数学和天文系)
;
California Institute of Technology(加州理工学院)
;
Particle Physics Division(粒子物理部)
;
Fermi National Accelerator Laboratory(费米国家加速器实验室)
;
Bexorg, Inc.(Bexorg公司)
NLP Methods May Actually Be Better Than Professors at Estimating Question Difficulty
Leonidas Zotos, Ivo Pascal de Jong, Matias Valdenegro-Toro, Andreea Ioana Sburlea, Malvina Nissim, Hedderik van Rijn
机构
*
Center for Language and Cognition, University of Groningen(语言与认知中心,格罗宁根大学)
;
Bernoulli Institute, University of Groningen(伯努利研究所,格罗宁根大学)
;
Department of Experimental Psychology, University of Groningen(实验心理学系,格罗宁根大学)
专题命中
评测与基准
:LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
Comments10 pages, 2 figures, presented at ECAI 2025 at the 2nd International Workshop on AI in Society, Education and Educational Research (AISEER)
Competence-Aware AI Agents with Metacognition for Unknown Situations and Environments (MUSE)
Rodolfo Valiente, Praveen K. Pilly
机构
*
Intelligent Systems Center, HRL Laboratories(智能系统中心,HRL实验室)
专题命中
评测与基准
:LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
CommentsReplaced all references to "self-awareness" with the more accurate term "self-assessment"; Updated Figure 2; Added recent pertinent work from the cognitive computational neuroscience literature; Removed the non-apples-to-apples comparison with Dreamer-v3 for self-assessment; Added additional experiments to validate the role of accurate self-assessment in effective self-regulation
EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
Wayne Chi, Valerie Chen, Ryan Shar, Aditya Mittal, Jenny Liang, Wei-Lin Chiang, Anastasios Nikolas Angelopoulos, Ion Stoica, Graham Neubig, Ameet Talwalkar, Chris Donahue
机构
*
Department of Electronic Engineering, City University of Hong Kong(DongGuan)(香港城市大学(东莞)电子工程系)
;
Department of Electronic Engineering, City University of Hong Kong(香港城市大学电子工程系)
专题命中
评测与基准
:LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI
Read Between the Lines: A Benchmark for Uncovering Political Bias in Bangla News Articles
Nusrat Jahan Lia, Shubhashis Roy Dipta, Abdullah Khan Zehady, Naymul Islam, Madhusodan Chakraborty, Abdullah Al Wasif
机构
*
University of Dhaka(达卡大学)
;
University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校)
;
Cisco Systems(思科系统)
;
BanglaLLM
;
Maharishi International University(玛哈里希国际大学)
;
Unityflow AI
专题命中
评测与基准
:LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL
专题命中
评测与基准
:LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL
CommentsPeer-reviewed and published version is in ICKG-2025 (The 16th IEEE International Conference on Knowledge Graphs, November 13-14, 2025, Limassol, Cyprus)
MARC: Multimodal and Multi-Task Agentic Retrieval-Augmented Generation for Cold-Start Recommender System
Seung Hwan Cho, Yujin Yang, Danik Baeck, Minjoo Kim, Young-Min Kim, Heejung Lee, Sangjin Park
机构
*
Department of Industrial Data Engineering, Hanyang University, Republic of Korea(工业数据工程系,翰阳大学)
;
School of Interdisciplinary Industrial Studies, Hanyang University, Republic of Korea(跨学科工业研究学院,翰阳大学)
专题命中
评测与基准
:LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI
Comments13 pages, 2 figures, Accepted at RDGENAI at CIKM 2025 workshop
Knots: A Large-Scale Multi-Agent Enhanced Expert-Annotated Dataset and LLM Prompt Optimization for NOTAM Semantic Parsing
Maoqi Liu, Quan Fang, Yang Yang, Can Zhao, Kaiquan Cai
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Beihang University(北航)
;
State Key Laboratory of CNS/ATM(国家空管流量管理技术实验室)
;
Aviation Data Communication Corporation(航空数据通信公司)
专题命中
评测与基准
:LLM(title);分类 cs.CL、cs.AI
CommentsAccepted to Advanced Engineering Informatics
MMD-Thinker: Adaptive Multi-Dimensional Thinking for Multimodal Misinformation Detection
Junjie Wu, Guohong Fu
机构
*
School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院)
;
Institute of Artificial Intelligence, Soochow University(苏州大学人工智能研究院)
专题命中
评测与基准
:large language model(abstract);language model(abstract);instruction tuning(abstract)
Language-Enhanced Generative Modeling for Amyloid PET Synthesis from MRI and Blood Biomarkers
Zhengjie Zhang, Xiaoxie Mao, Qihao Guo, Shaoting Zhang, Qi Huang, Mu Zhou, Fang Xie, Mianxin Liu
机构
*
organization= Shanghai Artificial Intelligence Laboratory , city= Shanghai , postcode= 200082 , country= China
;
organization= School of Medicine, Xiamen University , city= Xiamen , state= Fujian , country= China
;
organization= Department of Nuclear Medicine \& PET Center, Huashan Hospital, Fudan University , city= Shanghai , country= China
;
organization= Department of Gerontology, Shanghai Jiao Tong University Affiliated Sixth People’s Hospital , city= Shanghai , country= China
;
organization= Department of Computer Science, Rutgers University , city= New Brunswick , state= New Jersey , country= United States
;
organization= Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences , city= Shenzhen , country= China
专题命中
评测与基准
:LLM(abstract);large language model(abstract);language model(abstract)
机构
*
Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院)
;
College of AI, Tsinghua University(清华大学人工智能学院)
;
Shanghai Qi Zhi Institute(上海启智研究所)
;
State Key Laboratory of Virtual Reality Technology and Systems, Beihang University(北京航空航天大学虚拟现实技术与系统国家重点实验室)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG
Comments28 pages, 17 figures, accepted by NeruIPS 2025