Finding 8043Emerging EvidenceValidation V0
Surprisingly, a few long filings contribute most dataset value, showing high-quality, well-structured data beats sheer quantity in LLM training; SEFD introduces new benchmarks but lacks image coverage and may lag SEC schema changes.
86%Confidence
1Evidence objects
v1Version
DraftStatus
Evidence trail
Supporting86% linkage confidence
Surprisingly, a few long filings contribute most dataset value, showing high-quality, well-structured data beats sheer quantity in LLM training; SEFD introduces new benchmarks but lacks image coverage and may lag SEC schema changes.
key_findings bullet 3 · key_findings
Inspect source: The Stanford EDGAR Filings Dataset: Reconstructing U.S. Corporate and Financial Disclosures into Layout-Faithful and Token-Efficient Pretraining Data →This Finding was extracted from the configured corpus. It is versioned, traceable, and may evolve through editorial review or new corpus evidence.