arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期刊&会议

NeurIPS

Conference on Neural Information Processing Systems · 会议 · Machine Learning

2026-03-31 至 2026-03-31 共收录 7
2509.26351 2026-03-31 cs.LG

LLM-Assisted Emergency Triage Benchmark: Bridging Hospital-Rich and MCI-Like Field Simulation

基于大语言模型的急救分级基准:连接医院丰富与MCI类现场模拟

Joshua Sebastian, Karma Tobden, KMA Solaiman

机构 * University of Maryland Baltimore County(马里兰大学巴尔的摩县分校)

AI总结 本文提出一个开放的急救分级基准,通过大语言模型辅助构建,解决现有基准不可用的问题,涵盖医院环境和MCI模拟两种场景,提升临床AI数据集的可重复性和可访问性。

Comments Submitted to GenAI4Health@NeurIPS 2025. This was the first version of the LLM-assisted emergency triage benchmark dataset and baseline models. A related but separate benchmark-focused study on emergency triage under constrained sensing has been accepted at the IEEE International Conference on Healthcare Informatics (ICHI) 2026 (see arXiv:2602.20168)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27528 2026-03-31 cs.SD cs.IR

Advancing Multi-Instrument Music Transcription: Results from the 2025 AMT Challenge

推进多乐器音乐转录:2025年AMT挑战赛的结果

Ojas Chaturvedi, Kayshav Bhardwaj, Tanay Gondil, Benjamin Shiue-Hal Chou, Kristen Yeon-Ji Yun, Yung-Hsiang Lu, Yujia Yan, Sungkyun Chang

AI总结 本文报告了2025年自动音乐转录挑战赛的结果,展示了多乐器转录的进展与挑战,强调了转录准确性的提升及多声部和音色变化的难点。

Comments 7 pages, 3 figures. Accepted to the AI for Music Workshop at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21636 2026-03-31 cs.AI cs.CL

Silicon Bureaucracy and AI Test-Oriented Education: Contamination Sensitivity and Score Confidence in LLM Benchmarks

硅 bureaucracy 与 AI 考试导向教育:LLM 测试基准中的污染敏感性与分数可信度

Yiliang Song, Hongjun An, Jiangan Chen, Xuanchen Yan, Huan Song, Jiawei Shao, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom(中国电信人工智能研究院(TeleAI)) Guangxi Normal University(广西师范大学) Northwestern Polytechnical University(西北工业大学)

AI总结 本文探讨了基于测试基准的LLM评估体系,指出其依赖于基准分数直接反映真实泛化能力的假设,并提出审计框架以分析污染敏感性和分数可信度。

Comments Remove the NeurIPS 2026 template

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11417 2026-03-31 cs.CV

Robust Ego-Exo Correspondence with Long-Term Memory

具有长期记忆的鲁棒自体-外体对应

Yijun Hu, Bing Fan, Xin Gu, Haiqing Ren, Dongfang Liu, Heng Fan, Libo Zhang

机构 * University of Chinese Academy of Sciences(中国科学院大学) University of North Texas(北德克萨斯大学) Institute of Software Chinese Academy of Sciences(中国科学院软件研究所) Rochester Institute of Technology(罗切斯特理工学院)

AI总结 本文提出基于SAM 2的新型EEC框架,通过双记忆架构和自适应特征路由模块提升长期记忆能力,实现更鲁棒的自体-外体对应。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02028 2026-03-31 cs.CV cs.CR

See No Evil: Adversarial Attacks Against Linguistic-Visual Association in Referring Multi-Object Tracking Systems

见无所见:对抗性攻击针对参考多目标跟踪系统中的语言-视觉关联

Halima Bouzidi, Haoyu Liu, Mohammad Abdullah Al Faruque

机构 * University of California, Irvine(加利福尼亚大学尔湾分校)

AI总结 本文研究了参考多目标跟踪系统在设计逻辑上的安全问题,揭示了语言-视觉关联和目标匹配组件的对抗性漏洞,并提出VEIL框架以破坏其统一的参考匹配机制。

Comments Accepted to the NeurIPS 2025 Workshop on Reliable ML from Unreliable Data

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21356 2026-03-31 cs.CV

ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models

ShotBench:视觉语言模型中的专家级电影叙事理解

Hongbo Liu, Jingwen He, Yi Jin, Dian Zheng, Yuhao Dong, Fan Zhang, Ziqi Huang, Yinan He, Yangguang Li, Weichao Chen, Yu Qiao, Wanli Ouyang, Shengjie Zhao, Ziwei Liu

机构 * Tongji University(同济大学) The Chinese University of Hong Kong(香港中文大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) S-Lab, Nanyang Technological University(南洋理工大学S-Lab)

AI总结 本文提出ShotBench基准测试,用于评估视觉语言模型对电影叙事语言的理解能力,通过专家标注的3500多对问答数据揭示现有模型在细粒度视觉线索和复杂空间推理上的不足,并开发ShotVL模型在该基准上取得新突破。

Journal ref Advances in Neural Information Processing Systems 38 (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00200 2026-03-31 cs.LG cs.CR math.OC

Scalable Neural Network Verification with Branch-and-bound Inferred Cutting Planes

可扩展的神经网络验证与分支界限推导切割平面

Duo Zhou, Christopher Brix, Grani A Hanasusanto, Huan Zhang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) RWTH Aachen University(亚琛工业大学)

AI总结 本文提出BICCOS方法,通过神经网络验证问题的结构生成高效可扩展的切割平面,利用分支界限搜索树中的逻辑关系生成约束,提升验证效率和可扩展性。

Comments Accepted by NeurIPS 2024. BICCOS is part of the alpha-beta-CROWN verifier, the VNN-COMP 2024 winner; fixed Theorem 3.2 and clarified experimental results

详情

展开后加载摘要…

URL PDF HTML 收藏