[untitled]
1 points·by hynky··0 comments
[untitled]
1 points·by hynky··0 comments
FinePDFs: 3T token dataset made from internet PDFs
3 points·by hynky··1 comments
FineWeb2: Adapting Pre-Training Data Processing to Every Language
arxiv.org7 points·by hynky··0 comments
FineWeb2 dataset: A sparkling update with 1000s of languages
huggingface.co2 points·by hynky··0 comments
Alongside that they also shared a graph how much PDFs they were able to fetch by date, and the old internet seems to be mostly dead.