Data collection at petabyte scale: the good, the bad, and the ugly
Where pretraining data really comes from: the open sets, the gaps you scrape yourself, and the shadow libraries that keep showing up in training corpora.
Collecting the data to train a model. The story, in its good, bad, and ugly pieces.
First, the scale of what is out there:
The good
There is far more high-quality open data now than people realize. The Stack v2 has GitHub code, snapshotted in late 2023. FineWeb-Edu and DCLM cover the web. OpenWebMath, FineMath, and Proof-Pile-2 cover math. peS2o has the usable papers. Wikipedia, StackExchange, and Project Gutenberg fill the rest. Almost all of it sits on HuggingFace, though the terms vary by set, and the Stack takes two steps: metadata from HuggingFace, the file contents from S3.
HuggingFace is not the only source, and the best finds are not indexed anywhere. X is full of datasets people share and never list: reasoning traces, Claude and GPT coding trajectories, SWE-bench runs. I found more than twenty that way that you cannot find by searching. ModelScope, the Chinese HuggingFace, holds a lot that never gets mirrored west. I found over a hundred datasets there, mostly coding traces and reasoning data. When HuggingFace returns nothing, ModelScope usually has a version.
The code data needs balancing by language. The public code sets skew hard to Python, 65 to 70 percent by volume. For the training mix Python is capped at 20 percent, with 21+ languages represented. The separate GitHub scrape, measured in raw bytes before dedup, came out very different. C dominates, not because C is popular but because C projects are huge: kernel forks, drivers, embedded SDKs, vendored code. Dedup later shrinks C the hardest for the same reason:
The bad
The open sets have gaps, and filling them is manual. The Stack's snapshot ends in late 2023. Everything after that I scraped myself, repo by repo. That was about 600 GB of newer code, plus the pieces around it:
Then there is licensing, muddier than people admit. Some of the best book sets are released under noncommercial terms. The Harvard and Google Books collection is 983,000 books, 242 billion tokens, and its terms restrict commercial use and redistribution. Whether a given model release complies depends on those terms and the law, and releasing the weights as open or non-commercial does not settle it by itself.
The scale is its own problem. The catalogue runs to about 1.2 PB, more than you can store and more than you need. Most of the job is choosing which slice to keep. None of it is complicated. Running it reliably for months is the work.
The ugly
Books and papers have no clean source at scale. There is no licensed, ready-to-download corpus of every book ever written. So a lot of the field has ended up in the same place: the shadow libraries.
Anna's Archive aggregates LibGen, Sci-Hub, Z-Library, and DuXiu. An early 2026 snapshot showed around 61.6 million books and 95.7 million papers, over a petabyte, most of it available as bulk torrents:
Books from here show up in the training data of many popular models, and disclosure is usually thin. One of the largest labs was sued over these books; the training claim was dismissed on that record, and claims over how the files were acquired and shared continued. Books3, a set of pirated books, was part of The Pile, which several well-known model projects trained on. It is the ugly part of the field.
Then you clean it
Everything you pull lands as JSONL, one document per line, a few thousand shards across code, web, math, books, and papers. That uniform format is the whole point. Once every source looks the same, one pipeline can run across all of it.
Deduplication goes first, because the raw data is full of copies. The same book shows up in nine editions. The same article is mirrored across hundreds of sites. You catch them in layers: exact copies by hash, near-copies by MinHash, repeated passages inside otherwise-different documents by suffix array.
The tool for this is datatrove, HuggingFace's data processing library. It is what they used to build FineWeb, and it runs MinHash dedup as a four-stage pipeline: compute a signature per document, group signatures into buckets, cluster the buckets into duplicate sets, then filter keeping one document per cluster. My config was 14 bands of 8 hashes over 5-gram shingles. On the books it ran 25 hours across the stages. It is good software and it still is not turnkey at this scale. Stage three crashed on a bug in the release version, documents that never joined any cluster hit a lookup that assumed they had, and I patched the library to finish the run. Budget for that kind of thing. Nothing in this pipeline runs clean the first time at full size.
The run itself taught me something about tuning. I set stage one to 4 workers because an earlier run had crashed when 16 workers ran alongside extraction jobs fighting for the same RAM. Wrong lesson: on their own, dedup workers only use about 2 GB each, and the caution cost most of a day. Stage four, after I measured, ran at 24 workers:
The result on the books alone: 40 percent were near-duplicates of documents already in the set. Not junk, mostly the same works under different filenames. Skip this and you pay compute to train on the same text twice, and the model over-memorizes whatever repeats.
Then decontamination, which matters more than it sounds. If a benchmark's test questions are in your training data, your score on that benchmark goes up and you can no longer trust it. So you take the test splits of every benchmark you plan to report, cut them into overlapping n-grams, and scan every document for a match. I dropped 261,823 documents this way. The worst single file was one Python source with over a thousand hits, a tutorial repo that had pasted HumanEval and MBPP problems straight into its code. Miss that scan and the affected coding numbers are inflated, and you cannot measure by how much.
What you end up with
The final number is a choice, not what survived. The open web sets alone hold trillions of tokens, and nobody tokenizes all of that for a small model. Out of everything collected, I selected and tokenized about 700 billion tokens, sized to what I was training. The cleaning above made that keep trustworthy, it did not shrink it much: dedup removed 40 percent of the books, decontamination removed a fraction of a percent. A training run then uses around 200 billion of it. Text and code in near equal parts, math the rest, books and papers folded into the text bucket:
Months can go into this before any training happens. Finding the sources, scraping the gaps, deduplicating, decontaminating. That is where the time goes.
References
The open sources, for anyone building their own corpus:
- The Stack v2, GitHub code, late-2023 snapshot, with detected-license metadata
- FineWeb-Edu, education-filtered web
- DCLM-Baseline, curated Common Crawl
- OpenWebMath, math from the web
- FineMath, curated math
- Proof-Pile-2, math and formal proofs
- peS2o, cleaned academic papers
- Wikipedia
- SmolLM corpus, Cosmopedia, deduplicated FineWeb-Edu, and Python-edu
- Dolma, AllenAI's pretraining mix; the Reddit, arXiv, and peS2o slices came through here
- Dolmino and OLMo-mix, the OLMo training mixes
- MathPile, more math
- Cosmopedia, synthetic textbooks
- OpenCodeReasoning, execution-verified code solutions
- LoC PD Books, Library of Congress public-domain books
- StackExchange dump, the official data dump
- Institutional Books, the Harvard and Google Books set, noncommercial terms
- Project Gutenberg, public-domain books
- datatrove, the dedup pipeline
- ModelScope, the Chinese hub, for what HuggingFace does not have