You need to agree to share your contact information to access this dataset

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

This dataset is a compilation of ClinicalTrials.gov data (a U.S. Government
database) by the dataset author. The compilation is licensed under
CC-BY-NC-4.0. As a condition of accessing these files, you agree to:

  1. Cite this dataset, and attribute the underlying data to ClinicalTrials.gov,
    in any publication or derivative work.
  2. Not use these files, or copies of them, for commercial purposes without a
    separate commercial license from the author. Examples: ML training data
    for commercial models, redistribution in a commercial product or service,
    inclusion in databases offered for sale, or quantitative trading /
    commercial market analysis.

These terms apply to this dataset only, not to data you obtain directly from
ClinicalTrials.gov. For commercial licensing, contact licensing@anamnesisdata.com
or visit https://anamnesisdata.com.

Log in or Sign Up to review the conditions and access this dataset content.

Clinical Trials Version History

The Clinical Trials Version History dataset by Anamnesis Data: structured medical and scientific data for biotech, pharma and healthcare research. Public trial registries show only a trial's latest state. This dataset keeps every version, so you can see what a trial said on any past date, or when a sponsor moved a completion date.

606,620 trials (4,456,564 versions), one folder of Parquet files per table (79 tables, 541,236,196 rows). Access is gated; request it above. A free sample of 20 complete trials has the same tables and columns.

Example queries

sample_queries.sql in the sample holds ready-to-run DuckDB queries, from a trial's current record to serious adverse event rates. They run unchanged against this dataset's setup.sql, which creates one view per table; on the full data, add a LIMIT to the row-level ones.

Quick start

Download the repository, then in DuckDB, from that folder:

.read setup.sql          -- one view per table, named after its folder

-- the current public record of every trial
SELECT nct_id, nct_version, overall_status, lead_sponsor_name
FROM core
WHERE last_update_post_date IS NOT NULL      -- leaves out versions that failed review
QUALIFY row_number() OVER (PARTITION BY nct_id ORDER BY nct_version DESC) = 1;

Or read a table in place with a token from an account that has been granted access:

CREATE SECRET hf_token (TYPE huggingface, PROVIDER credential_chain);  -- uses `hf auth login`
SELECT * FROM 'hf://datasets/anamnesis-data/clinical_trials_history/core/*.parquet' LIMIT 5;
from datasets import load_dataset
core = load_dataset("anamnesis-data/clinical_trials_history", "core", split="train")

What's inside

Each folder is one table. core has one row per trial version; interventions, conditions, phases, arm_groups, locations and the results_* tables (participant flow, baseline, outcomes, adverse events) hang off it. Tables with nct_version join to core on (nct_id, nct_version); per-trial tables join on nct_id.

Things to know:

  • Versions. nct_version is the source's own 0-based version number for a trial.
  • Public dates. A version is public from last_update_post_date. Versions that failed the registry's review were never public: their last_update_post_date is empty, and reset_notice_post_date holds the date of the review's reset notice.
  • No "current" flag. The current record is a trial's highest public nct_version, as in the query above. It is the latest version held here, which can trail the live record.
  • Versions as captured. Each version is kept as captured; the registry later re-derives some fields in about 1 trial in 10. version_history follows the registry's current change list.
  • Obsolete IDs. Every table ends with is_obsolete and obsolete_by, which flag registrations since merged into another (22 trials). Filter on NOT is_obsolete to drop them.
  • List entries. List tables (for example interventions, arm_groups, secondary_ids, overall_officials, conditions and keywords) keep every source entry, in order: ord is its position in the source list, so entries that repeat a name are separate rows.
  • Text. No field holds HTML: entities are decoded (A &amp; B becomes A & B), tags are removed, and <word> used as brackets becomes (word). Rich-text fields (descriptions, eligibility criteria, results comments) are Markdown, like the public ClinicalTrials.gov API.

Releases

Each release is tagged vYYYY.MM.DD; pass it as revision= (or @vYYYY.MM.DD in an hf:// path) to pin one. A tag moves if the same day is republished; pin a commit hash for a fixed snapshot. This release's newest change: 2026-10-09.

Breaking changes since v2026.09.28

Code written against v2026.09.28 needs these changes; that release stays available under its tag.

  • interventions is keyed by (nct_id, nct_version, ord): entries with the same name and type are separate rows, and entries without a name or type are kept.
  • New ord column, ordered by it: arm_groups, collaborators, overall_officials, secondary_ids, conditions, keywords, phases, who_masked, std_ages, ipd_info_types, id_aliases. intervention_other_names gains intervention_ord and other_name_ord.
  • intervention_arm_links has one row per source element (source, parent_ord, item_ord, raw_label, intervention_ord, arm_group_ord); intervention_type uses the enum values.
  • core.last_update_post_date is empty for versions that failed review; the new last column reset_notice_post_date holds their reset-notice date.
  • Plain-text fields are decoded (see Text); version_history.module_labels is {} rather than NULL for changes that list no module.
  • No HTML in text: rich-text fields are Markdown, and tags and entities are removed (see Text).

Repository move

Releases up to v2026.09.26 were published at brbk/clinical_trials_history; new releases are published here only. Access approvals do not carry over, so request access again here.

Source and license

Trial records come from ClinicalTrials.gov, a U.S. Government database operated by the National Library of Medicine at NIH; no proprietary rights are claimed in them, and the site's terms apply to them. We restructured the records into tables and a subset of fields; dates given only as a month or year are set to its first day (a *_precision column, where present, keeps the original), and HTML is removed from text and rich-text fields are converted to Markdown. Other values are as returned.

The compilation (structure, schema, version history and documentation) is licensed CC BY-NC 4.0. Access is gated: you agree to cite this dataset and attribute the data to ClinicalTrials.gov, and not to use these files commercially without a separate license. These terms do not apply to data you obtain directly from ClinicalTrials.gov. For commercial licensing, contact licensing@anamnesisdata.com or visit anamnesisdata.com.

Citation

@misc{clinical_trials_history_2026,
  title  = {Clinical Trials Version History},
  author = {{Anamnesis Data}},
  year   = {2026},
  url    = {https://huggingface.co/datasets/anamnesis-data/clinical_trials_history}
}
Downloads last month
72