LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks
LongWoF-Bench:评估用于可验证长工作流任务的EvoMap基因
机构 * EvoMap ; Tsinghua University(清华大学)
AI总结 该研究提出LongWoF-Bench基准,发现通过EvoMap将验证的执行轨迹转化为可复用的Gene,能显著提升多模型在长工作流任务上的表现,且优于Skill和参考蒸馏的Gene。
高校专区
LongWoF-Bench:评估用于可验证长工作流任务的EvoMap基因
机构 * EvoMap ; Tsinghua University(清华大学)
AI总结 该研究提出LongWoF-Bench基准,发现通过EvoMap将验证的执行轨迹转化为可复用的Gene,能显著提升多模型在长工作流任务上的表现,且优于Skill和参考蒸馏的Gene。
统一多模态模型中视觉生成何时能助力视觉理解?
机构 * Nanjing University(南京大学) ; Tsinghua University(清华大学) ; Fudan University(复旦大学)
AI总结 本研究提出VGAU-Diag评估框架,发现统一多模态模型的视觉生成仅在简单任务上助益理解,瓶颈多在理解侧,有效生成需瞄准理解瓶颈。
机构 * Tsinghua University(清华大学) ; Beijing Normal-Hong Kong Baptist University(北京师范大学-香港浸会大学联合国际学院) ; University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校) ; University of California San Diego(加利福尼亚大学圣迭戈分校)
Comments Paper source and compact evidence: this https URL (https://github.com/JulianZJN/GenCoord)
离散扩散模型:从词元化到生成的统一框架
机构 * McGill University(麦吉尔大学) ; Mila - Quebec AI Institute(米拉-魁北克人工智能研究所) ; University of Cambridge(剑桥大学) ; University of Toronto(多伦多大学) ; MBZUAI - Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) ; Tsinghua University(清华大学) ; Rochester Institute of Technology(罗彻斯特理工学院) ; Salesforce(Salesforce公司) ; University of Illinois Chicago(伊利诺伊大学芝加哥分校)
AI总结 研究离散扩散模型,引入统一框架从离散状态空间构建审视该模型,让现有公式成为共同设计空间实例,揭示训练、推理等方面权衡,为未来研究提供方向。
技能条件门控自蒸馏用于大语言模型推理
机构 * Tsinghua University(清华大学) ; Fudan University(复旦大学) ; City University of Hong Kong(香港城市大学) ; Huazhong University of Science and Technology(华中科技大学) ; University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
AI总结 提出技能条件门控自蒸馏(SGSD),通过从经验技能库中检索技能-错误对构建多教师池,并利用验证器验证教师极性,以鲁棒门控目标蒸馏信息性师生差异,在弱先验信息假设下提升数学推理性能。
Comments Accepted by EMNLP 2026 Findings. Code is available at this https URL (https://github.com/walawalagoose/SGSD)
具有鲁棒性的扩散求解器用于逆问题
机构 * School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院) ; Yau Mathematical Sciences Center, Tsinghua University(清华大学尤太数学科学中心)
AI总结 本文提出一种鲁棒扩散求解器,通过噪声估计和Huber损失构建迭代加权最小二乘目标,以解决逆问题中的异常值问题,并在多个图像数据集上验证了其有效性。
Comments Accepted by CVPR 2026
通过无训练KV缓存压缩实现高效的长周期GUI代理
机构 * Tsinghua University(清华大学) ; Zhejiang University(浙江大学) ; The Chinese University of Hong Kong(香港中文大学)
AI总结 ST-Lite通过无训练的KV缓存压缩方法,提升GUI代理的解码效率,实现2.45倍加速并保持性能优势。
Comments Accepted to Findings of EMNLP 2026. Camera-ready version. 37 pages
基于拓扑信息的AI基础模型实现全球河流预报
机构 * College of Water Sciences, Beijing Normal University, Beijing, China(北京师范大学水科学学院) ; School of Geography and the Environment, University of Oxford, Oxford, UK(牛津大学地理与环境学院) ; Department of Transdisciplinary Science and Engineering, Institute of Science Tokyo, Tokyo, Japan(东京科学研究所跨学科科学与工程系) ; School of Systems Science, Beijing Normal University, Beijing, China(北京师范大学系统科学学院) ; Institute of Industrial Science, University of Tokyo, Tokyo, Japan(东京大学工业科学研究所) ; China Institute of Water Resources and Hydropower Research, Beijing, China(中国水利水电科学研究院) ; State Key Laboratory of Hydro-Science and Engineering, Department of Hydraulic Engineering, Tsinghua University, Beijing, China(清华大学水利科学与工程国家重点实验室) ; School of Artificial Intelligence, Beijing Normal University, Beijing, China(北京师范大学人工智能学院) ; School of Water Resources and Hydropower Engineering, Wuhan University, Wuhan, China(武汉大学水利水电学院)
AI总结 GraphRiverCast通过拓扑信息和物理对齐的神经操作符架构,实现了全球河流系统的系统性水动力模拟,无需历史数据即可进行稳健预测。
Comments 21 pages, 4 figures. Main text only; the supplementary materials accompany the journal version
寻找黄金:利用通用知识图谱扩展领域特定知识图谱
机构 * National Key Laboratory of Big Data and Decision(大数据与决策国家重点实验室) ; National University of Defense Technology(国防科技大学) ; The Center for machine learning research(机器学习研究中心) ; Peking University(北京大学) ; Tsinghua University(清华大学) ; The Hong Kong University of Science and Technology(香港科学与技术大学)
AI总结 本文提出领域特定知识图谱融合任务,通过神经符号框架ExeFuse,利用通用知识图谱提升领域知识图谱的完整性和实用性。
Comments 13 pages, 5 figures
PatientHub: 一个统一的患者模拟框架
机构 * The CoAI Group, DCST Institute for Artificial Intelligence Tsinghua University(CoAI集团、DCST人工智能研究所清华大学)
AI总结 PatientHub提供了一个统一的患者模拟框架,通过标准化定义、组成和部署,促进跨方法和跨模型的基准测试,并降低新方法开发的门槛。
Comments EMNLP 2026 Demo Paper
TangramPuzzle: 通过组合空间推理评估多模态大语言模型
机构 * Tsinghua University(清华大学) ; Sun-Yat Sen University(孙逸人大学) ; Tencent Youtu Lab(腾讯优图实验室) ; University of Illinois Chicago(伊利诺伊大学香槟分校)
AI总结 TangramPuzzle通过几何基准测试评估多模态大语言模型的组合空间推理能力,发现模型在匹配轮廓时忽视几何约束,导致碎片变形。
Comments EMNLP 2026 Findings
大语言模型对顺序敏感性如何?OrderProbe用于确定性结构重建
机构 * Peking University(北京大学) ; The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) ; City University of Hong Kong(香港城市大学) ; Fudan University(复旦大学) ; The Hong Kong Polytechnic University(香港理工大学) ; Tsinghua University(清华大学) ; Zhejiang University(浙江大学) ; University of Illinois Urbana-Champaign(伊利诺伊大学香槟分校) ; Marquette University(马凯特大学) ; Juniata College(朱尼阿特学院)
AI总结 研究通过OrderProbe基准评估大语言模型对输入顺序的敏感性,发现即使在前沿模型上,结构重建仍面临挑战,且语义能力与结构鲁棒性存在脱节。
Comments EMNLP 2026 Findings