SEC-BERT is a domain-specific language model pre-trained and fine-tuned exclusively on SEC filings, including 10-K annual reports, 10-Q quarterly reports, and 8-K current event filings. Developed by researchers at the University of Athens, it represents the most specialized open-source model for SEC document analysis available today.
The model was trained on a corpus of over 260,000 SEC filings sourced from the EDGAR database, covering approximately 10 years of corporate disclosures. This extensive training on regulatory text gives SEC-BERT an unmatched understanding of the language, structure, and conventions used in SEC filings β from risk factor disclosure language to management discussion and analysis (MD&A) sections.
For financial institutions, auditors, and compliance teams that process large volumes of SEC filings, SEC-BERT provides a powerful foundation for automating document classification, risk factor extraction, and entity recognition tasks. Its Apache 2.0 license makes it freely available for commercial use, and it integrates directly into EDGAR document processing pipelines without licensing costs or API dependencies.