arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AWARE-FX:一种可审计的知识引导型AI系统,用于衡量企业外汇套期保值披露情况

AWARE-FX: An Auditable Knowledge-Guided AI System for Measuring Corporate Foreign-Exchange Hedging Disclosure

Qi Wang

arXiv 2607.27611首次发表:更新:

发表机构

University of Nottingham(诺丁汉大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究开发可审计的AWARE-FX系统,将企业年报文本转换为可追溯的套期保值披露指标,通过多维度实验验证其可靠性,为外汇风险管理提供可单独审计的决策支持架构。

AI 中文摘要

企业年度报告包含关于外汇风险管理、衍生品使用、自然套期保值及明确未使用相关的弱结构化证据。本研究开发了AWARE-FX,这是一种可审计的AI/NLP决策支持系统,可将报告文本转换为可追溯的公司年度套期保值披露衡量指标。该系统结合了专业来源词典、否定与会计状态逻辑、渠道特定财务编码器、精确证据门、保守聚合及审计账本。在2008-2025年的24909个香港公司年度样本中,它检索并评分了543527个片段。通过消融实验、300个片段的分层人工审计、三种子的FinBERT-ModernBERT比较、严格的2023-2025时间测试、概率校准、选择性预测及固定提示生成式模型基准对可靠性进行评估。在8项编码器任务拆分比较中,FinBERT有7项的平均F1值更高;其时间F1范围为0.702至0.872。对20%置信度最低的时间观测值弃权(不执行),可使保留样本的F1值提升0.050-0.077。确定性Qwen3-8B在商品和否定证据上表现强劲,但在外国债务和会计上下文标签上表现不佳,表明通用大语言模型无法统一替代领域约束。严格的外汇评分与关联基线及压力时期外汇暴露呈负相关,而通用宽泛评分则不相关。这些关联提供了外部构念验证,而非套期保值有效性的因果估计。AWARE-FX贡献了一种经过测试的决策支持架构,其中检索、状态逻辑、分类、不确定性处理、聚合及外部验证均保持可单独审计。

英文摘要

Corporate annual reports contain weakly structured evidence about foreign-exchange risk management, derivative use, natural hedging, and explicit non-use. This study develops AWARE-FX, an auditable AI/NLP decision-support system that converts report text into traceable firm-year hedging-disclosure measures. The system combines a professional-source lexicon, negation and accounting-status logic, channel-specific financial encoders, exact evidence gates, conservative aggregation, and an audit ledger. Across 24,909 Hong Kong firm-years from 2008-2025, it retrieves and scores 543,527 snippets. Reliability is evaluated through ablations, a stratified 300-snippet human audit, three-seed FinBERT-ModernBERT comparisons, strict 2023-2025 temporal tests, probability calibration, selective prediction, and fixed-prompt generative-model benchmarks. FinBERT has the higher mean F1 in seven of eight encoder task-split comparisons; its temporal F1 ranges from 0.702 to 0.872. Abstaining on the 20% least-confident temporal observations raises retained-sample F1 by 0.050-0.077. Deterministic Qwen3-8B performs strongly on commodity and negation evidence but poorly on foreign-debt and accounting-context labels, showing that a general-purpose LLM does not uniformly replace domain constraints. The strict FX score is negatively associated with linked baseline and stress-period FX exposure, whereas the generic broad score is not. These associations provide external construct validation, not causal estimates of hedging effectiveness. AWARE-FX contributes a tested decision-support architecture in which retrieval, status logic, classification, uncertainty handling, aggregation, and external validation remain separately auditable.

Comments40 pages, 4 figures, 12 tables. Preprint; not peer reviewed

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑