LLMgram · AI News · 2026-07-24

Hugging Face releases The Stack v3, largest open code dataset

Hugging Face releases The Stack v3, largest open code dataset

Hugging Face has released The Stack v3, described as the largest open code dataset to date. It offers a pre-processed train set for direct model ingestion and a raw full 114 TB bucket for finer-grained control.

Key takeaway

Open code-model builders gain a new default corpus where the bottleneck shifts from scraping scarce data to filtering and quality control on a shared baseline.

Context

The Stack v3 is positioned as a standardized open code training resource aimed at reducing the need for each lab to rebuild large-scale cleaning and deduplication pipelines from scratch.

By shipping both a ready-to-train split and a raw full dump, Hugging Face is targeting teams that want either quick ingestion or deeper custom filtering over the same underlying collection.

Numbers to know

  • 114 TBSize of the raw full Stack v3 bucket available for granular control
LLMgram editorial selection and synthesis · @llmgram. LLMgram is not the original publisher of this information.
Continue on LLMgram: Open in AI Signal →
Original reporting: Hugging Face releases The Stack v3 – largest open code dataset yet