Tool · Methodology

How the Site Overlap Explorer works

Where the data comes from, how often it changes, what “similar study” and “same site” mean here, and what the numbers cannot tell you.

Data sources and coverage

ClinicalTrials.gov is the only trial source. The explorer calls the public ClinicalTrials.gov REST API v2 directly from your browser for each search; nothing is stored on a ClinBolt server and no account or key is involved. Each result shows the exact API requests it came from, the retrieval time, and each record's “last update posted” date from the registry.

Coverage is therefore exactly the registry's coverage: studies registered on ClinicalTrials.gov, with the site lists their sponsors chose to post. Studies registered only elsewhere (EU CTR, ISRCTN, jRCT, CTRI …) are not included. The explorer does not currently query other registries.

ChEMBL (EMBL-EBI) supplies compound synonyms: the international non-proprietary name, brand names, development codes and other names, each typed by ChEMBL. The exact requests are linked in the “Matched terms” panel. ChEMBL is a curated public database; a very new compound or a code not yet curated may be missing, in which case only your own term is used. ClinicalTrials.gov additionally applies its own synonym expansion to intervention and condition searches, so a trial may match through a term the panel does not show; those trials are labelled “registry synonym match”.

Limits per request. To keep the browser responsive and the registry API unburdened, a search loads a fixed number of studies per request (100 to 1,000, most recently updated first) and lets you load more. The summary cards always say how many studies matched in the registry and how many are loaded; site counts, maps and overlap are computed from loaded studies only.

Refresh behaviour

There is no cached dataset. Every search is a live query, so results reflect the registry at the moment shown as “retrieved”. ClinicalTrials.gov itself updates its public data daily; the API's data timestamp is shown under the page title. Within one browser session identical requests are answered from memory to avoid repeating calls when you change filters; reloading the page clears that.

The demonstration snapshot is different: it is a copy of real API responses captured on a stated date, offered only if the registry cannot be reached from your network. It is labelled on screen and in every export, and only the searches listed in its banner work.

Compound role: studied, comparator, combination or background

For a compound search, each trial's role is read from the registry's arm-group data, never guessed from the title:

Matching a term to an intervention uses the intervention name and its “other names”, after lower-casing, accent stripping and punctuation removal; development codes also match without their separator (MK-3475 ≡ MK3475). Terms shorter than three characters are ignored.

Study similarity

The overlap view compares a selected study with candidate studies retrieved from the registry in two ways, both shown with their request links and counts:

  1. Indication search: the selected study's listed conditions and derived MeSH condition terms (editable chips), OR-ed in the registry's condition search, restricted to a comparable phase by default (the selected phase and its neighbours: Phase 3 retrieves Phase 2, Phase 2/3, Phase 3 and Phase 4; “N/A” retrieves “N/A”). You can switch to the exact phase or any phase.
  2. Intervention search (optional, on by default when the study has drug or biological interventions): the selected study's therapy names OR-ed in the intervention search. Candidates found this way are the “exact compound” candidates.

Each candidate is then scored. The default weights and rules are:

ComponentDefault pointsRule
Indication40Full points for an identical condition string or MeSH condition term on both records; half points when the study was only found through the registry's synonym expansion (no identical term); 0 otherwise.
Phase25Full points for the same phase (or both N/A); half points for an adjacent phase (Phase 2 vs Phase 2/3, Phase 2/3 vs Phase 3); 0 otherwise or when either phase is missing.
Compound / intervention20Full points when any drug/biological intervention name, “other name” or MeSH intervention term is shared (generic entries such as “placebo” or “standard of care” are ignored); a quarter of the points for the same intervention type only. Studies with full points carry the exact compound label.
Population870% of the points scaled by the overlap (Jaccard) of standard age groups (child / adult / older adult) and 30% when sex eligibility matches.
Geography7Scaled by the overlap (Jaccard) of the countries in each study's reported sites.

A candidate counts as “similar” when its score reaches the minimum, 40 points by default: a matching indication alone qualifies, as does a comparable phase together with the same compound. Every weight and the threshold are editable in the explorer, and each row's “why” link lists the points and reason for every component.

Mechanism of action is not a scoring factor. ClinicalTrials.gov does not record it, and looking it up for every candidate's interventions would need many extra calls to a second database with uneven coverage. It is listed here so that nobody assumes it is included.

Overlap metrics and timing

Overlap % = number of sites shared with the comparison study ÷ number of reported sites in the selected study × 100
Jaccard similarity % = shared sites ÷ distinct sites across both studies × 100

“Reported sites” are distinct site groups (a study listing the same facility twice counts once), including placeholder-named sites, which can never be shared. A related study with no posted site list shows “no site list” and no percentage: its overlap is unknown, not zero.

Timing. Each related study is labelled from study-level data only:

None of these proves that two studies recruited at the same site at the same time. The registry does not record per-site dates. The only site-level evidence is the site status ClinicalTrials.gov publishes for recruiting studies; where both studies report “recruiting” at a shared site, the site panel and the shared-site table say so explicitly. Anything else is study-level context.

Site matching

Registry site names are free text. The explorer groups location records into sites with these rules, in this order, and reports a confidence level for every group:

  1. Normalise: lower-case; strip accents and punctuation; “&” → “and”; drop stop words (the, of, and, de, …) and legal forms (Inc, LLC, GmbH, …); expand common abbreviations (Univ → University, Hosp → Hospital, Ctr/Centre → Center, Inst → Institute, St → Saint, Mem → Memorial, Ped/Paediatric → Pediatric, …); join initials (M.D. → MD); remove site numbers (“( Site 0123 )”, “- 0042”, “#12”). Cities and countries are normalised the same way, with a short alias list (USA → United States, South Korea → Korea, Republic of, …).
  2. Placeholder names such as “Local Institution”, “Research Site”, “Novartis Investigative Site”, “Site 45” are kept as reported sites but never merged with anything, not even with each other. They are flagged in every table.
  3. Exact groups (high confidence): identical normalised name, city and country.
  4. Near-identical names (medium confidence): within the same city and country only, two groups merge when (a) their word sets overlap by at least 75% (Jaccard), or (b) the shorter name starts the longer one, sharing at least two words (“Samsung Medical Center” and “Samsung Medical Center, Sungkyunkwan University School of Medicine”), or (c) one name contains the other and every extra word is a generic descriptor such as “university”, “medical” or “school” (“Hospital Clinico San Carlos” and “Hospital Clinico Universitario San Carlos”). Rule (a) applies only when both names have words of their own; a name that is a subset of another goes through (b) or (c), so “Tongji Hospital, Tongji Medical College” is not merged into “Union Hospital, Tongji Medical College”, and “Beijing Hospital” is not merged into “Beijing Cancer Hospital”. In every case, if each name carries a distinctive word the other lacks, they are treated as different places (“Asan Medical Center” vs “Samsung Medical Center”, “Mount Sinai West” vs “Mount Sinai Morningside”), and extra words that name a campus or department (north, children's, cardiology, …) block the merge.
  5. Coordinates are not used. ClinicalTrials.gov's location coordinates are city centroids, so every site in a city shares one point; they place sites on the map but say nothing about identity.

A site group's confidence is the weakest rule used to build it. The Site matching control switches between Exact only (rule 3 only), Standard (the thresholds above) and Loose, which additionally accepts 60% word overlap at low confidence and one shared word for the prefix rule. Every group keeps every original registry name, visible in the site panel and the CSV exports. Two facilities that share a health-system name in different cities are never merged, because city and country must agree first.

Sponsors

“Same sponsor” and “different sponsor” are decided on the lead sponsor only. Collaborators are shown but never used for grouping. Lead-sponsor names are grouped after removing trailing legal forms and punctuation (Inc, LLC, Ltd, GmbH, AG, S.A., “and Company”, …), so “Merck Sharp & Dohme LLC” and “Merck Sharp & Dohme Corp.” group together while “Merck KGaA” stays separate. No other merging (subsidiaries, renamed companies, acquisitions) is attempted; the grouping key is shown on the study card when it differs from the name.

Known limitations

Exports

Every CSV includes a header block with the source, retrieval time and generation time, plus source identifiers (NCT IDs, registry URLs), the site group IDs used in the matrix, and the match confidence of every site group. In demonstration mode the file name and header say so.

ClinBolt is an experimental project. This tool is for informational and educational use by clinical operations and feasibility teams; it is not a substitute for verified site intelligence.