arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AudioJev:具有顺序校准概率的直接音频决策

AudioJev: Direct Audio Decisions with Order-Calibrated Probabilities

Sihan Lv, Zhen Li, Zhiqi Cao, Jinshan Zhang, Ying Li, Meng Xi, Jianwei Yin

arXiv 2610.01293首次发表:更新:

发表机构

Zhejiang University; Binjiang Institute of Zhejiang University; China Academy of Space Technology(浙江大学; 浙江大学滨江研究院; 中国空间技术研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AudioJev通过共享全参数模型直接映射音频到候选分布,定义顺序校准并采用随机错排SKL训练,在MMAU/MMAR上显著提升准确率和概率稳定性。

AI 中文摘要

音频决策通常依赖于转录文本无法保留的证据,而其概率估计可能取决于答案选项的排序方式。AudioJev通过一个共享的全参数模型,将波形、问题和提供的备选方案直接映射到候选分布。我们将顺序校准定义为在保持语义的选项排列下,保留答案的概率。随机错排SKL训练将每个问题与一个重新排序的视图配对,在该视图中每个备选方案都改变位置,监督两个答案,并在应用对称KL惩罚之前对齐两个分布。推理保留单一候选评分前向传播,无需校准头或顺序集成。在三个训练种子中,AudioJev在完整的MMAU/MMAR上达到68.88%/55.33%的平均准确率,并将随机顺序SKL相对于单视图移除消融降低了43.8%/60.5%。同一模型处理意图、环境声音、音符属性、语音活动和对话转换。配对移除消融和多顺序评估共同衡量预测准确性和概率稳定性,建立了一个直接音频接口,其校准目标作用于候选含义而非呈现位置。推理代码和模型权重可在以下https URL和https URL获取。

英文摘要

Audio decisions often depend on evidence that a transcript does not preserve, while their probability estimates can depend on how answer options are ordered. AudioJev maps a waveform, question and supplied alternatives directly to a candidate distribution through one shared full-parameter model. We define order calibration as preserving an answer's probability under meaning-preserving option permutations. Random-derangement SKL training pairs each question with a reordered view in which every alternative changes position, supervises both answers, and aligns the two distributions before applying a symmetric KL penalty. Inference retains a single candidate-scoring forward, with no calibration head or order ensemble. Across three training seeds, AudioJev reaches 68.88%/55.33% mean accuracy on complete MMAU/MMAR and reduces random-order SKL by 43.8%/60.5% relative to the single-view removal ablation. The same model handles intent, environmental sound, note properties, speech activity and conversational transitions. Paired removal ablations and multi-order evaluation measure predictive accuracy and probability stability together, establishing a direct audio interface whose calibration objective acts on candidate meaning rather than presentation position. Inference code and model weights are available at https://github.com/SihanLv/AudioJev-Inference and https://huggingface.co/shlv/AudioJev.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑