arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19751cs.CLcs.AIcs.CYcs.LGcs.NE

将阿拉伯方言连续体学习为连续空间:一种用于说话者起源预测的回归方法

Learning the Arabic Dialect Continuum as a Continuous Space: A Regression Approach to Speaker Origin Prediction

  • Higher School of Communication of Tunis (SUP’COM)(突尼斯高等传播学院(SUP’COM))
  • Prince Sultan University(苏丹王子大学)

机构由 AI 辅助整理,请以论文原文为准。

Mohamed Aziz Khadraoui, Adel Ammar, Bilel Benjdira, Zahid Khan, Skander Turki, Wadii Boulila

中文总结 AI 辅助

该研究提出基于回归的方法预测阿拉伯语说话者起源,通过分层神经架构融合多种表示,用球面测地线损失优化距离,在特定协议下取得一定定位误差和预测准确率,为阿拉伯方言地理定位提供框架并量化其优缺点。

中文摘要 AI 辅助

我们提出了一种基于回归的阿拉伯方言地理定位方法,将方言变异建模为连续地理空间而非离散类别。使用分层神经架构预测说话者起源为连续经纬度坐标,该架构通过Transformer编码器和可学习注意力池化查询融合帧级XLS-R-300M和Whisper-large-v3编码器表示与音位描述符。球面测地线损失直接优化地球表面的大圆距离。在按源记录分组的无泄漏5折GroupKFold协议下,模型获得481.2公里的合并中位数定位误差。辅助国家和城市预测准确率分别达64.5%和45.2%。排列曼特尔检验为阿拉伯方言连续体假设提供定量支持。通过城市掩码协议进一步探究真实泛化能力,零样本情况下平均误差升至1173.3公里。我们的发现将连续地理建模确立为阿拉伯方言地理定位的原则框架,并量化了其优势和仍存在的巨大空间。

英文摘要

We present a regression-based approach to Arabic dialect geolocation that models dialectal variation as a continuous geographic space rather than discrete categories. Speaker origin is predicted as continuous latitude-longitude coordinates using a hierarchical neural architecture that fuses frame-level XLS-R-300M and Whisper-large-v3 encoder representations with phonotactic descriptors through a Transformer encoder and a learnable attention-pooled query. A spherical geodesic loss directly optimizes great-circle distance on Earth's surface, avoiding distortions inherent to planar coordinate regression. Under a leakage-free 5-fold GroupKFold protocol grouped by source recording, our model attains a pooled median localization error of 481.2 km. Auxiliary country and city heads reach 64.5% and 45.2% accuracy, respectively. A permutation Mantel test on the learned latent space provides quantitative support for the Arabic dialect continuum hypothesis. To probe true generalization, we further introduce a city-masking protocol in which two cities per fold are removed from training but retained in validation. Under this zero-shot regime, the mean error rises to 1173.3 km, a 1.32x degradation relative to seen cities. Our findings establish continuous geographic modeling as a principled framework for Arabic dialect geolocation and quantify both its strengths and the substantial headroom that remains.

补充信息

↑