πŸ¦™Stalecollected in 66m

590GB SEC EDGAR Dataset Open-Sourced

590GB SEC EDGAR Dataset Open-Sourced
PostLinkedIn
πŸ¦™Read original on Reddit r/LocalLLaMA

πŸ’‘Free 43B-token SEC dataset unlocks finance LLMs without API costs (590GB on HF)

⚑ 30-Second TL;DR

What Changed

590GB dataset: 8M samples, 43B tokens from major SEC filings

Why It Matters

Provides free, open access to vast financial data for AI training, reducing reliance on paid services and enabling finance-focused LLMs. Democratizes corporate filing analysis in the closed AI ecosystem.

What To Do Next

Load the SEC-EDGAR dataset from Hugging Face and experiment with fine-tuning on 10-K filings.

Who should care:Developers & AI Engineers

Key Points

  • β€’590GB dataset: 8M samples, 43B tokens from major SEC filings
  • β€’Includes raw contents, parsed HTML/XML plaintext, and metadata
  • β€’Collected over 10 days using datamule-python respecting EDGAR rate limits
  • β€’Parsed with selectolax, modified doc2dict, and secsgml libraries
πŸ“°

Weekly AI Recap

Read this week's curated digest of top AI events β†’

πŸ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA β†—

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.