The Stanford EDGAR Filings Dataset: Reconstructing U.S. Corporate and Financial Disclosures into Layout-Faithful and Token-Efficient Pretraining Data
Stanford researchers have released the Stanford EDGAR Filings Dataset (SEFD), a major open-source resource that cleans and organizes over 18.5 million U.S. SEC filings for training large language models in finance. SEFD preserves complex layouts using MultiMarkdown, covering 550 billion tokens with minimal overlap with web data. The dataset introduces two benchmarks for financial prediction and table transcription. Notably, a few long filings hold most value, and structured data outperforms sheer volume in model training.
What it examines
This paper introduces the Stanford EDGAR Filings Dataset (SEFD), a large, open-source collection of U.S. SEC filings reconstructed into a layout-faithful, token-efficient format. SEFD aims to provide high-quality, structured financial documents for training and evaluating large language models, supporting tasks like financial reasoning and document understanding.
What it concludes
SEFD enables advanced financial language modeling, long-context pretraining, and evaluation tasks such as table transcription and forecasting. Its structure preserves important financial details often lost in other datasets. Applications include financial analysis, compliance, and document parsing. Limitations involve image parsing and coverage of rare filing types; future work may address these gaps.
Evidence objects
Stanford EDGAR Filings Dataset (SEFD) revolutionizes SEC filings by offering a clean, layout-faithful, and token-efficient resource for training large language models in finance, covering 18.5 million filings and 550 billion tokens.
key_findings bullet 1 · key_findings · validation V0
SEFDs innovative parsing pipeline preserves complex document structurestables, indentation, visual hierarchiesusing MultiMarkdown, enabling advanced financial reasoning and outperforming plain-text datasets, with less than 0.1% overlap with common web corpora.
key_findings bullet 2 · key_findings · validation V0
Surprisingly, a few long filings contribute most dataset value, showing high-quality, well-structured data beats sheer quantity in LLM training; SEFD introduces new benchmarks but lacks image coverage and may lag SEC schema changes.
key_findings bullet 3 · key_findings · validation V0
The Stanford EDGAR Filings Dataset (SEFD) offers a novel, open, layout-faithful, and token-efficient reconstruction of SEC filings for LLM pretraining. Preserving semantic and layout cues, SEFD uniquely enables accurate financial reasoning. Two new benchmarks, EDGAR-Forecast and EDGAR-OCR, further highlight its originality and substantial impact on financial AI research.
key_findings bullet 4 · key_findings · validation V0
Raw abstract and provenance
As high-quality public web corpora become increasingly exhausted, clean long-context documents have become a scarce and expensive source of training data for large language models (LLMs). Existing long-context corpora are often proprietary and costly to acquire, synthetically generated, or concentrated in narrow domains such as programming. We introduce the Stanford EDGAR Filings Dataset (SEFD), an open reconstruction of SEC filings into layout-faithful MultiMarkdown for financial language modeling and evaluation. SEFD makes audited financial statements, risk disclosures, ownership reports, accounting notes, and market-moving event filings usable as long-context pretraining data and as a basis for financial reasoning, forecasting, compliance, and document understanding. The resulting corpus is token-efficient, model-ready, and has less than 0.1% overlap with Common Crawl-derived corpora. We release SEFD-v1, a 152B-token initial public snapshot, and provide corpus-level analyses of a larger 18.5M-filing archive estimated at 550B tokens. We further introduce two SEFD-derived benchmarks: EDGAR-Forecast, which evaluates filing-grounded numerical forecasting after model knowledge cutoffs, and EDGAR-OCR, which evaluates transcription of complex financial tables.
Source row: 2011 · abstract type: unknown