Datasets:
The dataset viewer is not available because its heuristics could not detect any supported data files. You can try uploading some data files, or configuring the data files location manually.
Synthetic Pre-pretraining Datasets
This repository contains the pre-processed datasets used in the paper Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior.
The datasets cover synthetic pre-pretraining (PPT) tasks (e.g., k-Shuffle Dyck, Set, MP-Struct Core, NCA) and corresponding control/pre-training data mixtures (e.g., C4, SmolLM3, Olmo3, Marin, FineWeb-Edu). They are tokenized and intended for language model pre-pretraining and pre-training experiments across multiple model scales.
- Project page: https://huggingface.co/verify-ppt
- Code: https://github.com/gucci-j/verify-ppt-at-scale
For the full list of available dataset repositories, preprocessing scripts, and training/evaluation details, please refer to the GitHub repository and the project page.
- Downloads last month
- 1,098