arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高校专区

Carnegie Mellon University(卡内基梅隆大学)

2026-02-17 至 2026-02-17 共收录 17
2602.14989 2026-02-17 cs.CV cs.AI cs.LG

ThermEval: A Structured Benchmark for Evaluation of Vision-Language Models on Thermal Imagery

ThermEval: 一种用于评估视觉语言模型在热成像上的性能的结构化基准

Ayush Shrivastava, Kirtan Gangani, Laksh Jain, Mayank Goel, Nipun Batra

机构 * Indian Institute of Technology, Gandhinagar, India(印度理工学院加尔各答分校) Carnegie Mellon University, Pittsburgh, USA(卡内基梅隆大学)

AI总结 ThermEval通过结构化基准评估视觉语言模型在热成像上的性能,揭示其在温度推理和颜色映射转换上的不足,推动热视觉语言模型的发展。

Comments 8 Pages with 2 figures of main content. 2 pages of References. 10 pages of appendix with 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14874 2026-02-17 cs.RO

Affordance Transfer Across Object Instances via Semantically Anchored Functional Map

通过语义锚定的功能映射实现跨物体实例的 affordance 转移

Xiaoxiang Dong, Weiming Zhi

机构 * Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所) School of Computer Science, The University of Sydney(悉尼大学计算机科学学院) Australian Centre for Robotics, The University of Sydney(悉尼大学机器人中心) College of Connected Computing, Vanderbilt University(范德比大学连接计算学院)

AI总结 本文提出语义锚定的功能映射方法,通过单个视觉示范实现跨不同几何物体的 affordance 转移,提升机器人感知与行动的实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14432 2026-02-17 cs.LG cs.AI stat.ML

S2D: Selective Spectral Decay for Quantization-Friendly Conditioning of Neural Activations

S2D:选择性谱衰减用于神经激活的量化友好条件化

Arnav Chavan, Nahush Lele, Udbhav Bamba, Sankalp Dayal, Aditi Raghunathan, Deepak Gupta

机构 * Amazon(亚马逊公司) Carnegie Mellon University(卡内基梅隆大学)

AI总结 S2D通过选择性谱衰减减少神经激活异常值,提升模型在量化过程中的准确性和部署效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14224 2026-02-17 cs.SD cs.CL cs.MM

The Interspeech 2026 Audio Reasoning Challenge: Evaluating Reasoning Process Quality for Audio Reasoning Models and Agents

Interspeech 2026音频推理挑战:评估音频推理模型和代理的推理过程质量

Ziyang Ma, Ruiyang Xu, Yinghao Ma, Chao-Han Huck Yang, Bohan Li, Jaeyeon Kim, Jin Xu, Jinyu Li, Carlos Busso, Kai Yu, Eng Siong Chng, Xie Chen

机构 * Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学) Queen Mary University of London(伦敦大学Queen Mary) NVIDIA(NVIDIA公司) Carnegie Mellon University(卡内基梅隆大学) Qwen Team, Alibaba Group(通义实验室,阿里巴巴集团) Microsoft Corporation(微软公司)

AI总结 Interspeech 2026音频推理挑战通过评估推理过程质量,探讨了音频推理模型和代理在事实性和逻辑性方面的表现及改进方向。

Comments The official website of the Audio Reasoning Challenge: https://audio-reasoning-challenge.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13928 2026-02-17 cs.SD cs.LG

voice2mode: Phonation Mode Classification in Singing using Self-Supervised Speech Models

voice2mode:基于自监督语音模型的歌唱发声模式分类

Aju Ani Justus, Ruchit Agrawal, Sudarsana Reddy Kadiri, Shrikanth Narayanan

机构 * University of Birmingham, School of Computer Science, Birmingham, UK(伯明翰大学计算机科学学院) Carnegie Mellon University, Information Systems, Doha, Qatar(卡内基梅隆大学信息系统) University of Southern California, Department of Electrical(南加州大学电气工程系)

AI总结 voice2mode利用自监督语音模型提取的嵌入对四种歌唱发声模式进行分类,实验表明基础模型特征在准确率上显著优于传统方法。

Comments Accepted to the Speech, Music and Mind (SMM26) workshop at the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026). This is the preprint version of the paper to appear in the proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25260 2026-02-17 cs.AI cs.CL cs.LG

Internal Planning in Language Models: Characterizing Horizon and Branch Awareness

语言模型中的内部规划:刻画视野与分支意识

Muhammed Ustaomeroglu, Baris Askin, Gauri Joshi, Carlee Joe-Wong, Guannan Qu

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) Carnegie Mellon University(卡内基梅隆大学)

AI总结 研究通过分析语言模型内部计算结构,揭示规划视野与分支意识的特性,为理解模型内部动态提供通用工具。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18053 2026-02-17 cs.RO

V2V-GoT: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models and Graph-of-Thoughts

V2V-GoT: 基于多模态大语言模型和思维图的车对车协同自动驾驶

Hsu-kuang Chiu, Ryo Hachiuma, Chien-Yi Wang, Yu-Chiang Frank Wang, Min-Hung Chen, Stephen F. Smith

机构 * NVIDIA Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文提出V2V-GoT框架,结合多模态大语言模型和图-思维方法,提升车对车协同自动驾驶的感知、预测和规划能力。

Comments Accepted by ICRA 2026 (IEEE International Conference on Robotics and Automation). Project: https://eddyhkchiu.github.io/v2vgot.github.io/ Code: https://github.com/eddyhkchiu/V2V-GoT Dataset: https://huggingface.co/datasets/eddyhkchiu/V2V-GoT-QA

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15990 2026-02-17 cs.RO cs.CV

GelSLAM: A Real-time, High-Fidelity, and Robust 3D Tactile SLAM System

GelSLAM:一种实时、高保真度和鲁棒的3D触觉SLAM系统

Hung-Jui Huang, Mohammad Amin Mirzaee, Michael Kaess, Wenzhen Yuan

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 GelSLAM通过触觉传感实现实时高保真3D SLAM,以高精度重建物体形状并提升手部操作任务的鲁棒性。

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02873 2026-02-17 cs.AI

It's the Thought that Counts: Evaluating the Attempts of Frontier LLMs to Persuade on Harmful Topics

想法才是关键:评估前沿大语言模型在有害话题上的说服尝试

Matthew Kowal, Jasper Timm, Jean-Francois Godbout, Thomas Costello, Antonio A. Arechar, Gordon Pennycook, David Rand, Adam Gleave, Kellin Pelrine

机构 * Université de Montréal, MILA(蒙特利尔大学,MILA) Carnegie Mellon University(卡内基梅隆大学) MIT, Center for Research and Teaching in Economics(麻省理工学院,经济研究与教学中心) Cornell University, University of Regina(康奈尔大学, Regina大学) Cornell University, MIT(康奈尔大学,麻省理工学院)

AI总结 本文提出APE基准测试,评估前沿大语言模型在有害话题上的说服意愿,揭示模型在有害情境下尝试说服的倾向及风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07861 2026-02-17 cs.CL cs.AI cs.LG

Scalable LLM Reasoning Acceleration with Low-rank Distillation

可扩展的大语言模型推理加速与低秩蒸馏

Harry Dong, Bilge Acun, Beidi Chen, Yuejie Chi

机构 * CMU Department of Electrical and Computer Engineering(卡内基梅隆大学电气与计算机工程系) Carnegie Mellon University(卡内基梅隆大学) Meta FAIR at Meta(Meta FAIR) CMU(卡内基梅隆大学)

AI总结 Caprese通过低秩蒸馏方法恢复高效推理方法中丢失的数学推理能力,同时减少参数数量和延迟,提升响应效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09980 2026-02-17 cs.CV cs.RO

V2V-LLM: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models

V2V-LLM:基于多模态大语言模型的车与车协同自动驾驶

Hsu-kuang Chiu, Ryo Hachiuma, Chien-Yi Wang, Stephen F. Smith, Yu-Chiang Frank Wang, Min-Hung Chen

机构 * NVIDIA Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文提出基于多模态大语言模型的V2V-LLM,通过车与车协同感知提升自动驾驶安全性和性能。

Comments Accepted by ICRA 2026 (IEEE International Conference on Robotics and Automation). Project: https://eddyhkchiu.github.io/v2vllm.github.io/ Code: https://github.com/eddyhkchiu/V2V-LLM Dataset: https://huggingface.co/datasets/eddyhkchiu/V2V-GoT-QA

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13576 2026-02-17 cs.CR cs.AI cs.CL

Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges

评分标准作为攻击面:LLM裁判中的隐蔽偏好漂移

Ruomeng Ding, Yifei Pang, He Sun, Yizhong Wang, Zhiwei Steven Wu, Zhun Deng

机构 * University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) Carnegie Mellon University(卡内基梅隆大学) Yale University(耶鲁大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 研究揭示了基于评分标准的LLM裁判中隐蔽的偏好漂移问题,通过评分标准攻击可系统性降低目标领域准确性,影响模型对齐流程。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13418 2026-02-17 cs.LG

Text Has Curvature

文本具有曲率

Karish Grover, Hanqing Zeng, Yinglong Xia, Christos Faloutsos, Geoffrey J. Gordon

机构 * Carnegie Mellon University(卡内基梅隆大学) Meta

AI总结 本文提出Texture,一种文本原生的词级曲率信号,通过定义和应用曲率检测与利用,建立文本曲率范式,提升长上下文推断和生成性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13398 2026-02-17 cs.LG q-bio.QM

Accelerated Discovery of Cryoprotectant Cocktails via Multi-Objective Bayesian Optimization

通过多目标贝叶斯优化加速冻保护剂混合物的发现

Daniel Emerson, Nora Gaby-Biegel, Purva Joshi, Yoed Rabin, Rebecca D. Sandlin, Levent Burak Kara

机构 * Mechanical Engineering Department, Carnegie Mellon University(卡内基梅隆大学机械工程系) Center for Engineering in Medicine & Surgery, Department of Surgery, Massachusetts General Hospital, Harvard Medical School, and Shriners Children’s(医学与手术工程中心,外科部,麻省总医院,哈佛医学院,以及谢尔曼儿童医院)

AI总结 通过多目标贝叶斯优化结合高通量筛选,高效发现同时具有高CPA浓度和高细胞存活率的冻保护剂混合物。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13346 2026-02-17 q-bio.GN cs.AI cs.CV

CellMaster: Collaborative Cell Type Annotation in Single-Cell Analysis

CellMaster: 单细胞分析中的协作细胞类型注释

Zhen Wang, Yiming Gao, Jieyuan Liu, Enze Ma, Jefferson Chen, Mark Antkowiak, Mengzhou Hu, JungHo Kong, Dexter Pratt, Zhiting Hu, Wei Wang, Trey Ideker, Eric P. Xing

机构 * Halicioglu Data Science Institute, University of California, San Diego, CA, USA(哈利奇奥格卢数据科学研究所,加州大学圣地亚哥分校) Department of Electrical & Computer Engineering, Texas A&M University, College Station, TX, USA(电气与计算机工程系,德克萨斯A&M大学) Department of Medicine, University of California, San Diego, CA, USA(医学系,加州大学圣地亚哥分校) Department of Chemistry and Biochemistry, University of California San Diego, La Jolla, CA, USA(化学与生物化学系,加州大学圣地亚哥分校) Moores Cancer Center, University of California, San Diego, La Jolla, CA, USA(摩尔癌症中心,加州大学圣地亚哥分校) Mohamed bin Zayed University of AI, Abu Dhabi, UAE(穆罕默德·本·扎耶德人工智能大学) School of Computer Science, Carnegie Mellon University, Pittsburgh, PA, USA(计算机科学学院,卡内基梅隆大学)

AI总结 CellMaster通过利用LLM编码知识实现零样本细胞类型注释,提升单细胞分析的准确性和可解释性。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13324 2026-02-17 cs.CV cs.AI cs.RO

Synthesizing the Kill Chain: A Zero-Shot Framework for Target Verification and Tactical Reasoning on the Edge

构建杀伤链:一种零样本框架,用于边缘节点上的目标验证和战术推理

Jesse Barkley, Abraham George, Amir Barati Farimani

机构 * With the Department of Mechanical Engineering, Carnegie Mellon University(卡内基梅隆大学机械工程系)

AI总结 本文提出了一种零样本框架,通过边缘设备实现目标验证和战术推理,验证了分层架构在动态军事环境中的有效性。

Comments 8 Pages, 3 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13212 2026-02-17 cs.RO cs.MA cs.SY eess.SY

UAVGENT: A Language-Guided Distributed Control Framework

UAVGENT: 一种语言引导的分布式控制框架

Ziyi Zhang, Xiyu Deng, Guannan Qu, Yorie Nakahira

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) Carnegie Mellon University(卡内基梅隆大学)

AI总结 UAVGENT通过结合语言引导的任务推理与分布式反馈控制,实现了多无人机系统在复杂任务中的鲁棒性和稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏