ELLMPEG: An Edge-based Agentic LLM Video Processing Tool
ELLMPEG:一种基于边缘的代理LLM视频处理工具
Zoha Azimi, Reza Farahani, Radu Prodan, Christian Timmerer
机构
*
Christian Doppler Laboratory ATHENA, Department of Information Technology (ITEC) University of Klagenfurt(亚琛实验室ATHENA,信息科技系,克雷格弗尔特大学)
;
Department of Information Technology (ITEC) University of Klagenfurt(信息科技系,克雷格弗尔特大学)
;
Department of Computer Science University of Innsbruck(计算机科学系,因斯布鲁克大学)
机构
*
School of Informatics, Xiamen University(厦门大学信息学院)
;
Institute of Artificial Intelligence, Xiamen University(厦门大学人工智能研究院)
;
Zhejiang Expressway Co., Ltd.(浙江高速公路有限公司)
;
School of Science and Engineering, Chinese University of Hong Kong(香港中文大学科学与工程学院)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
Achieving Time Series Reasoning Requires Rethinking Model Design, Tasks Formulation, and Evaluation
实现时间序列推理需要重新思考模型设计、任务制定和评估
Yaxuan Kong, Yiyuan Yang, Shiyu Wang, Chenghao Liu, Yuxuan Liang, Ming Jin, Stefan Zohren, Dan Pei, Yan Liu, Qingsong Wen
机构
*
University of Oxford, UK(牛津大学)
;
Hong Kong University of Science(香港科学大学)
;
Griffith University, Australia(格里菲斯大学)
;
Tsinghua University, China(清华大学)
;
University of Southern California, USA(南加州大学)
EvalQReason: A Framework for Step-Level Reasoning Evaluation in Large Language Models
EvalQReason: 一种用于大型语言模型中分步推理评估的框架
Shaima Ahmad Freja, Ferhat Ozgur Catak, Betul Yurdem, Chunming Rong
机构
*
Department of Electrical Engineering and Computer Science, University of Stavanger(电子工程与计算机科学系,斯塔万格大学)
;
Department of Electrical and Electronics Engineering, Izmir Bakircay University(电子与电气工程系,伊兹密尔巴克伊大学)
机构
*
School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University, China(中山大学深圳校区计算机科学与技术学院)
;
Hong Kong University of Science(香港科技大学)
ClueTracer: Question-to-Vision Clue Tracing for Training-Free Hallucination Suppression in Multimodal Reasoning
ClueTracer: 问题到视觉线索追踪用于无训练 hallucination 抑制在多模态推理
Gongli Xi, Kun Wang, Zeming Gao, Huahui Yi, Haolang Lu, Ye Tian, Wendong Wang
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Nanyang Technological University(南洋理工大学)
;
West China Biomedical Big Data Center(西京生物大数据中心)
机构
*
Department of Computer Science and Technology, Beijing Jiaotong University, Beijing, China(北京交通大学计算机科学与技术学院)
;
Tianjin Tasly Digital Chinese Medicine Technology Co., Ltd.(天津塔斯丽数字中医科技有限公司)
;
Tasly Biopharmaceuticals Co., Ltd.(塔斯丽生物医药有限公司)
;
State Key Laboratory of Chinese Medicine Modernization, Tianjin, China(中药现代化国家工程实验室,天津,中国)
;
Institute of Liver Diseases, Hubei Key Laboratory of the theory and application research of liver and kidney in traditional Chinese medicine, Hubei Provincial Hospital of Traditional Chinese Medicine, Wuhan, China(肝病研究所,湖北省中医肝肾理论与应用研究重点实验室,湖北省中医药研究院,武汉,中国)
;
Affiliated Hospital of Hubei University of Chinese Medicine, Wuhan, China(湖北中医药大学附属医院,武汉,中国)
;
Hubei Province Academy of Traditional Chinese Medicine, Wuhan, China(湖北省中医药研究院,武汉,中国)
;
Department of Gastroenterology, Guang’anmen Hospital, China Academy of Chinese Medical Sciences, Beijing, China(消化内科,广安门医院,中国中医科学院,北京,中国)
;
China Academy of Chinese Medical Sciences, Beijing, China(中国中医科学院,北京,中国)
;
China Institute for History of Medicine and Medical Literature, China Academy of Chinese Medical Sciences, Beijing, China(中国中医科学院中国医学史与医学文献研究所,北京,中国)
;
Beijing Research Institute of Chinese Medicine, Beijing University of Chinese Medicine, Beijing, China(北京中医研究院,北京中医药大学,北京,中国)
;
Beijing University of Chinese Medicine Third Affiliated Hospital, Beijing University of Chinese Medicine, Beijing 100029, China(北京中医药大学第三附属医院,北京中医药大学,北京100029,中国)
Reasoning by Commented Code for Table Question Answering
通过注释代码进行表格问题回答
Seho Pyo, Jiheon Seok, Jaejin Lee
机构
*
Department of Data Science, Seoul National University, Seoul, Republic of Korea(数据科学系,首尔国立大学,首尔,大韩民国)
;
Department of Computer Science(计算机科学系)
;
Engineering, Seoul National University, Seoul, Republic of Korea(工程系,首尔国立大学,首尔,大韩民国)
机构
*
University of California, Los Angeles(加州大学洛杉矶分校)
;
Tianjin University(天津大学)
;
Peking University(北京大学)
;
Zhejiang University(浙江大学)
;
Beijing Institute of General Artificial Intelligence(北京一般人工智能研究院)
Ao Sun, Hongtao Zhang, Heng Zhou, Yixuan Ma, Yiran Qin, Tongrui Su, Yan Liu, Zhanyu Ma, Jun Xu, Jiuchong Gao, Jinghua Hao, Renqing He
机构
*
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
University of Chinese Academy of Sciences(中国科学院大学)
;
CAS Key Laboratory of AI Safety, Institute of Computing Technology, CAS(中国科学院人工智能安全重点实验室)
;
University of Science and Technology of China(中国科学技术大学)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
University of Oxford(牛津大学)
;
Tsinghua University(清华大学)
HalluHard: A Hard Multi-Turn Hallucination Benchmark
HalluHard: 一种具有挑战性的多轮 hallucination 评估基准
Dongyang Fan, Sebastien Delsad, Nicolas Flammarion, Maksym Andriushchenko
机构
*
EPFL(苏黎世联邦理工学院)
;
ELLIS Institute Tübingen(图宾根ELLIS研究所)
;
Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所)
;
Tübingen AI Center(图宾根人工智能中心)
机构
*
LARG, Research Center for Social Computing and Interactive Robotics, HIT(LARG,社会计算与交互机器人研究中心,哈尔滨工业大学)
;
School of Computer Science and Engineering, Central South University(计算机科学与工程学院,中南大学)
;
iFLYTEK
From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations
从分数到步骤:诊断和改进证据医学计算中LLM的性能
Benlu Wang, Iris Xia, Yifan Zhang, Junda Wang, Feiyun Ouyang, Shuo Han, Arman Cohan, Hong Yu, Zonghai Yao
机构
*
Department of Computer Science, Yale University, CT, USA(耶鲁大学计算机科学系)
;
Center for Healthcare Organization and Implementation Research, VA Bedford Health Care(VA贝福德医疗中心健康组织与实施研究中心)
;
Miner School of Computer and Information Sciences, UMass Lowell, MA, USA(UMass洛厄尔矿尔计算机与信息科学学院)
;
Manning College of Information and Computer Sciences, UMass Amherst, MA, USA(UMass阿默斯特马宁信息与计算机科学学院)
CommentsEqual contribution for the first two authors. To appear as an Oral presentation in the proceedings of the Main Conference on Empirical Methods in Natural Language Processing (EMNLP) 2025