Dataset Viewer

The dataset viewer is not available because its heuristics could not detect any supported data files. You can try uploading some data files, or configuring the data files location manually.

Synthetic Pre-pretraining Datasets

This repository contains the pre-processed datasets used in the paper Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior.

The datasets cover synthetic pre-pretraining (PPT) tasks (e.g., k-Shuffle Dyck, Set, MP-Struct Core, NCA) and corresponding control/pre-training data mixtures (e.g., C4, SmolLM3, Olmo3, Marin, FineWeb-Edu). They are tokenized and intended for language model pre-pretraining and pre-training experiments across multiple model scales.

For the full list of available dataset repositories, preprocessing scripts, and training/evaluation details, please refer to the GitHub repository and the project page.

Downloads last month
1,098

Paper for verify-ppt/ppt-mpstructcore_v2