arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Texas at Austin(得克萨斯大学奥斯汀分校)

共收录 1921
2609.03894 2026-09-04 cs.CL cs.SE 新提交

CROCODIL: Cross-Model Code Editing with LLMs

CROCODIL:基于大语言模型的跨模型代码编辑

Linghan Zhong, Aditya Thimmaiah, Jayanth Srinivasa, Milos Gligoric, Junyi Jessy Li

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Cisco Research(思科研究院)

AI总结 本研究针对多LLM协作时的外来代码过度编辑问题,提出CROCODIL后训练框架,通过相似度奖励与执行奖励的乘积优化,在保证功能正确的同时减少了不必要的代码编辑。

Comments EMNLP 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00038 2026-09-04 stat.ML cs.CE cs.LG cs.NA math.NA 版本更新

Active learning for data-driven reduced models of parametric differential systems with Bayesian operator inference

基于贝叶斯算子推断的参数微分系统的数据驱动降阶模型主动学习

Shane A. McQuarrie, Mengwu Guo, Anirban Chaudhuri

机构 * Department of Mathematics, Brigham Young University(数学系, Brigham Young University) Centre for Mathematical Sciences, Lund University(数学科学中心, Lund University) Oden Institute for Computational Engineering and Sciences, The University of Texas at Austin(计算工程与科学研究院, The University of Texas at Austin)

AI总结 本文提出基于贝叶斯算子推断的主动学习方法,用于提升参数微分系统数据驱动降阶模型的稳定性和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.02089 2026-09-03 cs.CL cs.LG 新提交

IDEEA: training-free Input-Dependent stEEring via Activation cluster matching

IDEEA:基于激活簇匹配的无训练输入依赖引导

Zheng Wang, Muchen Li, Renjie Liao, Yan Leng

机构 * University of Waterloo(滑铁卢大学) University of British Columbia(不列颠哥伦比亚大学) Vector Institute for AI(人工智能矢量研究所) University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本研究针对现有无训练引导方法输入无关的局限,提出IDEEA框架,通过聚类注意力头的激活并匹配输入激活选择对应方向,在保留输入表示的同时对齐LLMs,使TruthfulQA的truth×info指标平均提升9.9%。

Comments Accepted to EMNLP 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.01794 2026-09-03 cs.CL 新提交

Disentangling Statistical Preemption from Entrenchment in Language Models' Avoidance of Overgeneralization

在语言模型避免过度概括时,区分统计优先与固化效应

Yixuan Wang, Freda Shi, Kanishka Misra

机构 * University of Waterloo(滑铁卢大学) Vector Institute(向量研究所) University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 该研究通过对语言模型开展受控实验,区分了语言模型避免过度概括时的优先效应与固化效应,发现其在动词特定层面无优先效应,仅存在微弱抽象优先效应,还揭示了模型对竞争结构的证据处理方式,为相关人类实验提供了新方向。

Comments Accepted to EMNLP 2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.01788 2026-09-03 cs.CL cs.AI 新提交

VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages

VakyArth:评估跨印度语言的大语言模型语用能力

Usneek Singh, Poorvaja Veera Balaji Kumar, Parth Nanda, Anand Madhusoodanan, Geyang Guo, Wei Xu, Junyi Jessy L

机构 * Georgia Institute of Technology(佐治亚理工学院) University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本研究推出首个印度语言语用基准VakyArth,评估多语言LLM在印度语言文化相关语用现象上的表现,发现其存在系统性语用推理失败及语言间差异。

Comments Findings of EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.01876 2026-09-03 cs.CV cond-mat.mtrl-sci 新提交

RAFT-DVC: Resolution-Aware Machine Learning-Based Digital Volume Correlation

RAFT-DVC:分辨率感知的基于机器学习的数字体积相关

Zixiang Tong, Lehu Bu, Jin Yang

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Texas Materials Institute(德克萨斯材料研究所)

AI总结 本文提出分辨率感知的RAFT-DVC密集DVC框架,明确其精度与工作模式,在不同纹理、位移条件下性能优异,可实现跨纹理迁移与大体积估计。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22642 2026-09-03 cs.LG cs.AI 版本更新

Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for Molecules

Mol-JEPA:用于分子的多模态联合嵌入预测架构

Florian Rottach, Sebastian Schieferdecker, William Rudman, Randall Balestriero, Carsten Eickhoff

机构 * University of Tübingen(蒂宾根大学) Boehringer Ingelheim(勃林格殷格翰) The University of Texas at Austin(德克萨斯大学奥斯汀分校) Brown University(布朗大学)

AI总结 针对分子基础模型的化学无效增强等局限,提出Mol-JEPA多模态框架,利用模态掩码融入生化上下文,在基准测试中展现出优异性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08389 2026-09-03 cs.AI cs.IR cs.MA 版本更新

Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents

不值得再增加一个Token:面向高效深度研究智能体的边际价值估计

Harshitha Kolukuluru, Reshma Ashok, Kirat Arora, Evan William Ciccarelli, Nischal Ashok Kumar, Lunyiu Nie, Franck Dernoncourt, Samyadeep Basu, Ryan A. Rossi, Nedim Lipka

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) The University of Texas at Austin(德克萨斯大学奥斯汀分校) Adobe Research(奥多比研究院)

AI总结 该研究针对长周期研究智能体的上下文冗余问题,系统比较不同阶段的剪枝策略,发现早期剪枝可大幅降本,轻量级启发式方法能减73%Token用量且质量损失小,为高效智能体设计提供指导。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07678 2026-09-03 cs.LG cs.AI 版本更新

DOG-DPO:Dynamic Optimization in Geometry for Safety Alignment

DOG-DPO:几何中的动态优化用于安全对齐

Yi Nian, Tiankai Yang, Yudi Zhang, Qi Pan, Zelong Xu, Shenzhe Zhu, Qingqing Luan, Yue Huang, Xiangliang Zhang, Yue Zhao

机构 * University of Southern California(南加州大学) Iowa State University(爱荷华州立大学) University of Wisconsin–Madison(威斯康星大学麦迪逊分校) UT Austin(德克萨斯大学奥斯汀分校) Independent Researcher(独立研究员) University of Notre Dame(圣母大学)

AI总结 提出DOG-DPO框架,将偏好对表示为模型表示空间中的方向,通过几何分解和多样性覆盖选择子集,仅用11%数据即可恢复大部分安全增益。

Comments Accepted by 2026 EMNLP (Conference on Empirical Methods in Natural Language Processing)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05616 2026-09-03 cs.CL 版本更新

What's in a Name? Morphological Shortcuts by LLMs in Pharmacology

名字里有什么?LLM在药理学中的形态捷径

Kaijie Mo, Thomas Yang, Chantal Shaib, Qing Yao, William Rudman, Ramez Kouzy, Kanishka Misra, Byron C. Wallace, Junyi Jessy Li

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Northeastern University(东北大学) MD Anderson Cancer Center(MD安德森癌症中心)

AI总结 研究LLM在药理学中依赖词缀线索进行推理的形态捷径行为,通过虚构药物名称实验和归因框架揭示其机制及安全风险。

Comments Accepted to the Main Conference of EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02824 2026-09-03 cs.CL 版本更新

CALIBURN: Self-Calibrated LLM Unlearning Alignment

CATNIP:通过校准和令牌化的负偏好对齐实现LLM去学习

Zhengbang Yang, Yisheng Zhong, Junyuan Hong, Zhuangdi Zhu

机构 * George Mason University(乔治·马歇尔大学) University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 CATNIP通过校准和令牌化的负偏好对齐方法,实现有效LLM去学习,无需保留数据或对比对,提升知识遗忘与保留的平衡。

Comments EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00290 2026-09-03 cs.CL cs.LG stat.ML 版本更新

DLM-One: Diffusion Language Models for One-Step Sequence Generation

DLM-One:用于单步序列生成的扩散语言模型

Tianqi Chen, Shujian Zhang, Mingyuan Zhou

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 该研究提出DLM-One框架,通过分数蒸馏实现连续扩散语言模型的单步序列生成,大幅提升采样速度并保持基准任务性能,还提出对抗正则化两阶段训练方案防止模型退化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00759 2026-09-03 cs.CV cs.AI cs.CL 版本更新

Multimodal Language Models as Text-to-Image Model Evaluators

多模态语言模型作为文本到图像模型的评估器

Jiahui Chen, Candace Ross, Reyhane Askari-Hemmat, Koustuv Sinha, Melissa Hall, Amy Zhang, Michal Drozdzal, Adriana Romero-Soriano

机构 * FAIR at Meta - Montreal, New York(Meta 蒙特利尔 FAIR) University of Texas at Austin(德克萨斯大学奥斯汀分校) Mila McGill University(麦吉尔大学) Canada CIFAR AI chair(加拿大 CIFAR 人工智能主席)

AI总结 该研究提出MT2IE框架,以MLLM为评估智能体,生成少量提示即可高效准确评估T2I模型,其排名比现有指标更忠实一致,还可定制基准,助力动态评估框架发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00966 2026-09-03 cs.LG cs.CC 版本更新

Smoothed Analysis for Learning Concepts with Low Intrinsic Dimension

针对具有低固有维度的概念学习的平滑分析

Gautam Chandrasekaran, Adam Klivans, Vasilis Kontonis, Raghu Meka, Konstantinos Stavropoulos

机构 * UT Austin(德克萨斯大学奥斯汀分校) UCLA(加州大学洛杉矶分校)

AI总结 本文提出平滑分析框架,实现低固有维度概念的学习,得到不可知学习k个半空间交集的首个多项式时间算法,规避传统学习的强困难性结果。

Comments 50 pages. This is the TheoretiCS journal version

Journal ref TheoretiCS, Volume 5 (2026), Article 13, 1-50

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.01491 2026-09-02 cs.CL cs.AI cs.MA 新提交

GlossoGen: Emergent Language in Complex Multi-Agent LLM Interactions

GlossoGen:复杂多智能体大语言模型交互中的涌现语言

Elias Stengel-Eskin, Newton Sander, Carlos Bonetti, Sasha Boguraev, James Bowler, Hale Sirin, Simon Kirby

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校) AE Studio(AE工作室) Schmidt Sciences(施密特科学机构) University of Edinburgh(爱丁堡大学)

AI总结 研究人员推出GlossoGen平台,发现多LLM智能体交互中会涌现出人类无法理解的组合性新语言,明确了语言演化的关键条件及不同模型在语言传递中的作用,证实LLM具备人类特有的累积文化演化潜力。

Comments GlossoGen code: https://github.com/agencyenterprise/GlossoGen Paper code: https://github.com/esteng/emergent_communication

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.00579 2026-09-02 cs.PL cs.AI cs.CL cs.SE 新提交

Predicting Program Exit Code with LLMs and Programming Language Semantics

用大语言模型(LLM)和编程语言语义预测程序退出码

Lara Marinov, Aditya Thimmaiah, Jayanth Srinivasa, Junyi Jessy Li, Milos Gligoric

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Cisco Research(思科研究院)

AI总结 本研究提出程序可执行性预测(PrEx)任务,构建含系统生成非法变换的数据集,评估开源编码LLM,发现其依赖预训练先验而非给定语义,在修改语义及复杂程序上表现差。

Comments Accepted at LMPL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.00443 2026-09-02 cs.CL cs.AI 新提交

(V)LMs generalize beyond surface co-occurrence: Evidence from cross-modal number agreement

(视觉)语言模型((V)LMs)可泛化至表面共现之外:来自跨模态数一致的证据

Zach Studdiford, Kanishka Misra

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 该研究以(V)LMs为对象,通过跨模态泛化实验发现其可超越表面共现实现泛化,表现出与抽象兼容的行为,为语言模型具备抽象能力提供了证据。

Comments 9 pages main text

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26872 2026-09-02 cs.LG cs.AI cs.CL 版本更新

When the Strongest Teacher Is Not the Best Teacher: Student-Centric Answer Selection

最强的教师并不总是最好的教师:以学生为中心的答案选择

Zhengyu Hu, Zheyuan Xiao, Linxin Song, Fengqing Jiang, Yuetai Li, Zhihan Xiong, Yue Liu, Junhao Lin, Yao Su, Lijie Hu, Kaize Ding, Teng Xiao, Radha Poovendran

机构 * University of Washington(华盛顿大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Southern California(南加州大学) Independent Researcher(独立研究者) National University of Singapore(新加坡国立大学) Microsoft(微软) Google(谷歌) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Northwestern University(西北大学) Allen Institute for AI (AI2)(人工智能研究院(AI2))

AI总结 提出以学生为中心的答案采样(SCAS)框架,通过估计学生中心的学习成本选择教师生成的答案,从而提升学生模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29928 2026-09-01 cs.CV 新提交

Biomechanical 3D Body: Self-Supervised Distillation of Biomechanical Pose from a 3D Body Foundation Model

生物力学三维人体:从三维人体基础模型自监督蒸馏生物力学姿态

R. James Cotton, J. D. Peiffer, Lucinda Williamson, John Leske, Georgios Pavlakos

机构 * Shirley Ryan AbilityLab(雪莉·瑞安能力实验室) Northwestern University(西北大学) University of Illinois Chicago(伊利诺伊大学芝加哥分校) University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 该研究扩展SAM-3D-Body模型新增生物力学预测头,用JAX结合Equinox实现,在公开数据集训练后,在多数据集验证中优于现有生物力学回归模型,仅略逊于需高成本轨迹优化的同类方法。

Comments Accepted to the ECCV 2026 Workshop MoCha

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.28934 2026-09-01 cs.LG 新提交

Revisiting the Provable-Auditable Privacy Gap of DP-SGD

重新审视DP-SGD的可证明可审计隐私差距

Saloni Modi, Srivi Balaji, Yusong Zhu, Gautam Kamath, Kevin Tian

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Waterloo(滑铁卢大学) Vector Institute(矢量研究院)

AI总结 本研究针对DP-SGD的可证明可审计隐私差距,提出轻量级防御框架提升其经验隐私且不损失理论隐私,经多场景评估验证其灵活性。

Comments The code for this paper can be found here: https://github.com/pineappleEnthusiast/empirical-privacy-defense

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.28924 2026-09-01 cs.CL 新提交

Causal Interventions Reveal Typologically Organized Syntactic Mechanisms in Multilingual Language Models

因果干预揭示多语言语言模型中类型学组织的句法机制

Sasha Boguraev, Toshiki Nakai, Kyle Mahowald, Julius Steuer

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Saarland University(萨尔大学) Leipzig University(莱比锡大学) Heidelberg Institute for Theoretical Studies(海德堡理论研究所)

AI总结 本研究利用机制可解释性技术,在四种多语言语言模型的三种句法结构上证实跨语言机制迁移,且迁移程度与语言类型学相似度正相关,为语言学理论提供新假说。

Comments 20 Pages, 7 Figures, 11 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.28718 2026-09-01 cs.RO cs.AI cs.CV cs.ET cs.SY eess.SY 新提交

RoboPhys-3D: A Comprehensive Embodied World Model Evaluation via 3D Reconstruction

RoboPhys-3D:通过三维重建进行全面的具身世界模型评估

Tianyi Wang, Jiazhou Chen, Yiming Xu, Xiangyu Li, Tianyi Zeng, Chih-Hsien Chou, Ning Lu, Liang Peng, Junfeng Jiao, Christian Claudel

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Purdue University(普渡大学) Futurewei Technologies, Inc(星纪时代科技公司)

AI总结 本研究推出基于RoboTwin 2.0的三维具身世界模型基准RoboPhys-3D,含50个操作任务等数据,提出两个评估分数,发现Cosmos 3的RoboPhyscore最高,该分数与人类评估一致性强。

Comments 66 pages, 12 figures, 55 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06714 2026-09-01 cs.AI 版本更新

The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows

优化器即智能体:跨提示、程序与机器学习工作流的推理驱动搜索

Junbo Li, Boyi Liu, Canwen Xu, Yite Wang, Yuxiong He, Zhangyang Wang, Qiang Liu, Zhewei Yao

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Snowflake(思诺飞克公司)

AI总结 该研究提出ReASearch统一框架,让智能体自主优化提示、程序与ML工作流,在14项任务中优于专用系统,部分发现超人类最佳结果的方案。

Journal ref COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29793 2026-09-01 cs.CL q-fin.GN 版本更新

Fund2Persona: A Framework for Building and Refining Financial Advisor Personas from Fund Disclosure Data

Fund2Persona:基于基金披露数据构建与精炼金融顾问角色的框架

Suhwan Park, Hoyoung Lee, Zhangyang Wang, Alejandro Lopez-Lira, Young Cha, Chanyeol Choi, Jaewon Choi, Yongjae Lee

机构 * UNIST(蔚山科学技术院) LinqAlpha University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Florida(佛罗里达大学) Blackstone(黑石集团) Hanwha Life(韩华生命)

AI总结 提出Fund2Persona框架,利用基金披露、持仓变动、市场背景和管理人评论构建金融顾问角色,并通过智能体循环精炼,在持仓重建和管理人评论对齐任务上优于通用基线,生成更具体有用的建议。

Comments 18 pages, 4 figures, 18 tables. EMNLP 2026 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17188 2026-09-01 cs.CV cs.CL 版本更新

Not Truly Multilingual: Script Consistency as a Missing Dimension in VLM Evaluation

并非真正的多语言:脚本一致性作为VLM评估中缺失的维度

Prabhjot Singh, Bhushan Pawar, Madhu Reddiboina, Rajvee Sheth

机构 * RediMinds Inc.(RediMinds公司) The University of Texas at Austin(德克萨斯大学奥斯汀分校) Independent Researcher(独立研究员)

AI总结 提出PuMVR基准,评估10个VLM在旁遮普语三种文字上的表现,发现显著的脚本差距,并提出脚本一致性率(SCR)作为必要评估指标。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27790 2026-09-01 cs.LG 版本更新

SYNAPSE: Neuro-Symbolic Visual Thought-to-Text Decoding via Topological Semantic Denoising

SYNAPSE: 通过拓扑语义去噪的神经符号视觉思维到文本解码

Akshaj Murhekar, Abhijit Mishra

机构 * School of Information University of Texas at Austin(信息学院德克萨斯大学奥斯汀分校)

AI总结 提出SYNAPSE框架,利用常识图结构和潜在样本进行推理时符号正则化,稳定脑电到文本解码中的语义生成,无需微调大语言模型。

Comments Accepted to Findings of EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06433 2026-09-01 physics.comp-ph cs.LG physics.flu-dyn 版本更新

Operator Learning for Predicting Bulk Wave Parameters of Spectral Wave Models

运算学习用于海面波浪诱导力的替代建模

Shukai Cai, Sourav Dutta, Mark Loveland, Eirik Valseth, Peter Rivera-Casillas, Corey Trahan, Clint Dawson

机构 * Oden Institute for Computational Engineering and Sciences, The University of Texas at Austin(德克萨斯大学奥斯汀分校奥登计算工程与科学研究所) Information Technology Laboratory, U.S. Army Engineer Research and Development Center(美国陆军工程兵研究与发展中心信息技术实验室) Department of Mechanical Engineering and Technology Management, Norwegian University of Life Sciences(挪威生命科学大学机械工程与技术管理系) Department of Numerical Analysis and Scientific Computing, Simula Research Laboratory(Simula研究实验室数值分析与科学计算系)

AI总结 本文利用深度运算网络替代传统波浪模型,通过不同边界条件和风场的数值实验,验证了其在预测辐射应力梯度和显著波高方面的高精度。

Comments 46 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28900 2026-09-01 cs.RO cs.AI cs.LG cs.SY eess.SY 版本更新

Robust Multi-Agent Reinforcement Learning for Small UAS Separation Assurance under GPS Degradation and Spoofing

针对GPS退化和欺骗的鲁棒多智能体强化学习用于小型无人机分离保障

Alex Zongo, Filippos Fotiadis, Ufuk Topcu, Peng Wei

机构 * George Washington University(乔治华盛顿大学) University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本文通过多智能体强化学习解决小型无人机在GPS退化和欺骗下的鲁棒分离保障问题,提出闭式表达式对抗扰动,实现线性时间评估,并在高密度无人机模拟中取得近零碰撞率。

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01605 2026-09-01 cs.LG stat.ML 版本更新

Universal Redundancies in Time Series Foundation Models

时间序列基础模型中的通用冗余性

Anthony Bao, Venkata Hasith Vattikuti, Jeffrey Lai, William Gilpin

机构 * ECE Department, UT Austin(UT奥斯汀电子工程系) Department of Physics, UT Austin(UT奥斯汀物理系) Oden Institute, UT Austin(UT奥斯汀奥登研究所)

AI总结 本研究揭示了时间序列基础模型中普遍存在的冗余特性,并通过理论框架和消融实验揭示了模型退化现象的根源。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14129 2026-08-31 cs.CV 版本更新

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models

别让视频说话:面向音频-视觉语言模型的音频对比偏好优化

Ami Baid, Zihui Xue, Kristen Grauman

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本文提出ACPO方法,通过引入输出对比和输入对比目标,解决音频-视觉语言模型中视频驱动的音频幻觉问题,提升音频真实性和多模态能力。

Comments ECCV 2026 camera-ready version. Project page: https://vision.cs.utexas.edu/projects/acpo/

详情

展开后加载摘要…

URL PDF HTML 收藏