You need to agree to share your contact information to access this dataset
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
This dataset is a compilation of ClinicalTrials.gov data (a U.S. Government
database) by the dataset author. The compilation is licensed under
CC-BY-NC-4.0. As a condition of accessing these files, you agree to:
- Cite this dataset, and attribute the underlying data to ClinicalTrials.gov,
in any publication or derivative work. - Not use these files, or copies of them, for commercial purposes without a
separate commercial license from the author. Examples: ML training data
for commercial models, redistribution in a commercial product or service,
inclusion in databases offered for sale, or quantitative trading /
commercial market analysis.
These terms apply to this dataset only, not to data you obtain directly from
ClinicalTrials.gov. For commercial licensing, contact licensing@anamnesisdata.com
or visit https://anamnesisdata.com.
Log in or Sign Up to review the conditions and access this dataset content.
Clinical Trials Version History
The Clinical Trials Version History dataset by Anamnesis Data: structured medical and scientific data for biotech, pharma and healthcare research. Public trial registries show only a trial's latest state. This dataset keeps every version, so you can see what a trial said on any past date, or when a sponsor moved a completion date.
606,620 trials (4,456,564 versions), one folder of Parquet files per table (79 tables, 541,236,196 rows). Access is gated; request it above. A free sample of 20 complete trials has the same tables and columns.
Example queries
sample_queries.sql
in the sample holds ready-to-run DuckDB queries, from a trial's current record to serious adverse
event rates. They run unchanged against this dataset's
setup.sql,
which creates one view per table; on the full data, add a LIMIT to the row-level ones.
Quick start
Download the repository, then in DuckDB, from that folder:
.read setup.sql -- one view per table, named after its folder
-- the current public record of every trial
SELECT nct_id, nct_version, overall_status, lead_sponsor_name
FROM core
WHERE last_update_post_date IS NOT NULL -- leaves out versions that failed review
QUALIFY row_number() OVER (PARTITION BY nct_id ORDER BY nct_version DESC) = 1;
Or read a table in place with a token from an account that has been granted access:
CREATE SECRET hf_token (TYPE huggingface, PROVIDER credential_chain); -- uses `hf auth login`
SELECT * FROM 'hf://datasets/anamnesis-data/clinical_trials_history/core/*.parquet' LIMIT 5;
from datasets import load_dataset
core = load_dataset("anamnesis-data/clinical_trials_history", "core", split="train")
What's inside
Each folder is one table. core has one row per trial version; interventions, conditions,
phases, arm_groups, locations and the results_* tables (participant flow, baseline,
outcomes, adverse events) hang off it. Tables with nct_version join to core on
(nct_id, nct_version); per-trial tables join on nct_id.
Things to know:
- Versions.
nct_versionis the source's own 0-based version number for a trial. - Public dates. A version is public from
last_update_post_date. Versions that failed the registry's review were never public: theirlast_update_post_dateis empty, andreset_notice_post_dateholds the date of the review's reset notice. - No "current" flag. The current record is a trial's highest public
nct_version, as in the query above. It is the latest version held here, which can trail the live record. - Versions as captured. Each version is kept as captured; the registry later re-derives some
fields in about 1 trial in 10.
version_historyfollows the registry's current change list. - Obsolete IDs. Every table ends with
is_obsoleteandobsolete_by, which flag registrations since merged into another (22 trials). Filter onNOT is_obsoleteto drop them. - List entries. List tables (for example
interventions,arm_groups,secondary_ids,overall_officials,conditionsandkeywords) keep every source entry, in order:ordis its position in the source list, so entries that repeat a name are separate rows. - Text. No field holds HTML: entities are decoded (
A & BbecomesA & B), tags are removed, and<word>used as brackets becomes(word). Rich-text fields (descriptions, eligibility criteria, results comments) are Markdown, like the public ClinicalTrials.gov API.
Releases
Each release is tagged vYYYY.MM.DD; pass it as revision= (or @vYYYY.MM.DD in an hf://
path) to pin one. A tag moves if the same day is republished; pin a commit hash for a fixed
snapshot. This release's newest change: 2026-10-09.
Breaking changes since v2026.09.28
Code written against v2026.09.28 needs these changes; that release stays available under its tag.
interventionsis keyed by(nct_id, nct_version, ord): entries with the same name and type are separate rows, and entries without a name or type are kept.- New
ordcolumn, ordered by it:arm_groups,collaborators,overall_officials,secondary_ids,conditions,keywords,phases,who_masked,std_ages,ipd_info_types,id_aliases.intervention_other_namesgainsintervention_ordandother_name_ord. intervention_arm_linkshas one row per source element (source,parent_ord,item_ord,raw_label,intervention_ord,arm_group_ord);intervention_typeuses the enum values.core.last_update_post_dateis empty for versions that failed review; the new last columnreset_notice_post_dateholds their reset-notice date.- Plain-text fields are decoded (see Text);
version_history.module_labelsis{}rather than NULL for changes that list no module. - No HTML in text: rich-text fields are Markdown, and tags and entities are removed (see Text).
Repository move
Releases up to v2026.09.26 were published at brbk/clinical_trials_history; new releases are
published here only. Access approvals do not carry over, so request access again here.
Source and license
Trial records come from ClinicalTrials.gov, a U.S. Government
database operated by the National Library of Medicine at NIH; no proprietary rights are claimed in
them, and the site's terms apply to them.
We restructured the records into tables and a subset of fields; dates given only as a month or year
are set to its first day (a *_precision column, where present, keeps the original), and HTML is
removed from text and rich-text fields are converted to Markdown. Other values are as returned.
The compilation (structure, schema, version history and documentation) is licensed CC BY-NC 4.0. Access is gated: you agree to cite this dataset and attribute the data to ClinicalTrials.gov, and not to use these files commercially without a separate license. These terms do not apply to data you obtain directly from ClinicalTrials.gov. For commercial licensing, contact licensing@anamnesisdata.com or visit anamnesisdata.com.
Citation
@misc{clinical_trials_history_2026,
title = {Clinical Trials Version History},
author = {{Anamnesis Data}},
year = {2026},
url = {https://huggingface.co/datasets/anamnesis-data/clinical_trials_history}
}
- Downloads last month
- 72