arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11620cs.CL

一种免训练、免对齐的企业情报方法:应用于SEC文件

A Training-Free, Alignment-Free Approach to Corporate Intelligence: Application to SEC Filings

Jean-François Delpech

首次发表
浏览论文内容

中文总结 AI 辅助

提出免训练免对齐的稀疏种子向量框架,用于SEC文件分析,实现快速文档比较和事件追踪,无需LLM推理。

中文摘要 AI 辅助

高维稠密文本嵌入和大语言模型在财务披露分析中面临实际障碍:上下文窗口限制、幻觉风险、高计算成本,以及独立训练模型间向量空间的任意旋转。我们提出了一种基于确定性稀疏种子向量的免训练、免对齐的企业情报框架。通过将词字符串哈希到固定的高维基中,所有文档和所有时间段在构造上处于同一坐标系,从而无需训练或对齐。在句子上下文中累积这些种子向量,可生成可线性组合的语料特定语义签名,支持亚秒级文档比较、发行人指纹识别、跟踪发行人在文件之间词汇变化,以及主题句子提取,所有这些均在普通CPU硬件上完成。我们在多年SEC文件语料库(10-K、10-Q、8-K)上展示了该方法,表明重大企业事件(包括波音737 MAX危机、英特尔供应链中断和邦吉收购维特拉)如何以独特、可解释的语义特征出现,每个特征都可追溯到产生它的确切源句子,且无需特定领域训练和LLM推理。

英文摘要

High-dimensional dense text embeddings and large language models face real obstacles in financial-disclosure analysis: context-window limits, hallucination risk, high computational cost, and the arbitrary rotation of vector spaces across independently trained models. We present a training-free, alignment-free framework for corporate intelligence built on deterministic sparse seed vectors. Hashing word strings into a fixed high-dimensional basis places all documents and all temporal epochs in a common coordinate system by construction, removing any need for training or alignment. Accumulating these seed vectors across sentence contexts yields corpus-specific semantic signatures that compose linearly, supporting sub-second document comparison, issuer fingerprinting, tracking of how an issuer's vocabulary shifts between filings, and thematic sentence extraction, all on ordinary CPU hardware. Demonstrating the approach on a multi-year corpus of SEC filings (10-K, 10-Q, 8-K), we show how material corporate events, among them Boeing's 737 MAX crisis, Intel's supply-chain disruptions, and Bunge's acquisition of Viterra, emerge as distinct, interpretable semantic profiles, each traceable to the exact source sentences that produced it, with no domain-specific training and no LLM inference.

发表机构

  • WebGlyphs, Inc.(WebGlyphs公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑