arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于多动作会计日志的中小企业财务指导观测策略排序

Observational Policy Ranking for SMB Financial Guidance from Multi-Action Accounting Logs

Shrutendra Harsola, Vignesh Subrahmaniam, Vikas Raturi, Kamalika Das, Xiang Gao, Kratika Gupta, Ruocheng Guo, Padmaja Jonnalagedda, Ananya Pramod, Sricharan Kumar

arXiv 2608.10050首次发表:更新:

发表机构

Foresight-AI; Intuit(预见人工智能; 英图易公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对中小企业财务指导的观测策略排序问题,提出CAR-PL方法,通过85078个公司-月度数据验证,其在毛利润等指标上表现优异,可实现多动作会计日志下的财务指导类别排序。

AI 中文摘要

中小企业需要及时的财务指导,但历史会计日志记录的是企业自行选择且常同时发生的业务变更,而非随机推荐。我们将该场景建模为观测策略排序:基于决策前财务信息,策略为目标财务关键绩效指标(KPI)从34个分类账衍生的业务变更类别中选择一个。使用来自7505家企业的85078个公司-月度观测数据,我们提出协变量调整残差策略学习(CAR-PL),这是一种逐动作R学习器,可直接处理多热日志并通过观测支持正则化选择。我们在公司不相交的保留企业上,采用共享模型辅助评分规则,将CAR-PL与提升T学习器、保守上下文价值模型、零样本大语言模型(LLM)及非个性化参考进行比较。CAR-PL的毛利润点估计值最高(0.084),T学习器的营收点估计值最高(0.085),上下文价值模型的速动比率点估计值最高(0.062)。在匹配的公司聚类比较中,CAR-PL与T学习器在任一增长KPI上均无统计差异,而CAR-PL会选择33-34个类别,且在目录中的选择集中度更低。仅结果模型评分保留相同的KPI级别点估计领导者或顶级对,当全零处理参考被最常见的训练协同动作模式取代时,类别排名仍相似。这些发现支持从多动作会计日志中进行针对特定目标的中小企业财务指导排序。

英文摘要

Small and medium-sized businesses need timely financial guidance, yet historical accounting logs record self-selected and often co-occurring business changes rather than randomized recommendations. We formulate this setting as observational policy ranking: from pre-decision financial information, a policy selects one of 34 ledger-derived business-change categories for a target financial KPI. Using 85,078 company-month observations from 7,505 firms, we introduce Covariate-Adjusted Residual Policy Learning (CAR-PL), an action-wise R-learner that operates directly on multi-hot logs and regularizes selection by observational support. We compare CAR-PL with an uplift T-Learner, a conservative contextual value model, a zero-shot LLM, and non-personalized references on company-disjoint held-out firms under a shared model-assisted scoring rule. CAR-PL has the highest Gross Profit point estimate (0.084), the T-Learner has the highest Revenue point estimate (0.085), and the contextual value model has the highest Quick Ratio point estimate (0.062). CAR-PL and the T-Learner are not statistically separated on either growth KPI in matched company-clustered comparisons, while CAR-PL selects 33-34 categories and produces less concentrated selections across the catalog. Outcome-model-only scoring retains the same KPI-level point-estimate leader or top pair, and category rankings remain similar when the all-zero treatment reference is replaced by the most common training co-action pattern. These findings support objective-specific ranking of SMB financial guidance from multi-action accounting logs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑