AI use case

AI for clinical trial site selection and feasibility

AI that scores and ranks candidate investigator sites and countries for a planned clinical trial by predicted enrollment speed, access to the eligible patient population and historical performance, so clinical operations teams choose and activate a shortlist of sites with a higher chance of meeting enrollment targets on time, instead of relying on which sites a study team happens to know.

By Len Debets · Last verified 29 September 2026 · 2 public deployments

USD 3 million to USD 90 million
Indicative value per year
A biopharma sponsor starting 10 trials a year. Worked example, see how it is calculated.

What problem does it solve?

Choosing where to run a clinical trial has traditionally been a manual, relationship driven process: a study team asks the sites and investigators it already knows, supplemented by feasibility surveys that ask sites to self report how many eligible patients they can enroll. Past performance data sits in different systems for different trials, and a site with real potential but no history with the sponsor rarely makes the list because nobody thought to ask it.

The cost of getting this wrong is high. Enrollment takes almost half the time a medicine spends in clinical trials, and a large share of trials never reach their enrollment target on the original timeline, which delays the medicine and adds cost for every month the trial runs open. Underneath that average sits a second, quieter problem: sites that are fast to enroll are not necessarily the ones whose patients look like the population the medicine will eventually treat, so trials can end up both slow and unrepresentative unless someone deliberately corrects for it.

  • About 80% of clinical trials fail to enroll enough participants to move forward, and with almost half the time to bring a medicine through clinical trials spent on enrollment, this causes serious delays in getting potential new drugs to patients who need them now.Follow the Data (2025)

How does it work?

  1. Pull together site and investigator history. The model ingests every site's and investigator's past performance across the sponsor's own trials, screen failure rates, enrollment rates and protocol deviations, rather than the one or two trials a single study team happened to work on.
  2. Add real world data on where patients actually are. Anonymized or aggregated data, such as insurance claims or electronic health records, shows where people who match the trial's eligibility criteria are treated today, including countries and sites the sponsor has never worked with.
  3. Score and rank candidates. The model learns from trials with a similar indication and design to predict each candidate site's likely enrollment rate, then returns a ranked list with the reasoning behind each rank, not just a single number.
  4. Feed the ranking into feasibility, not around it. The shortlist goes into the standard feasibility survey and site selection committee process, so sites still confirm their own interest, capacity and any competing trials before anyone commits to them.
  5. Check predictions against what actually happens. As trials close out, actual enrollment is compared back against the prediction, and the model is retrained so it keeps up with new therapeutic areas, geographies and trial designs.
Audience
Employee facing
Autonomy
Assist
Adoption
Emerging
Channels
Internal tools

What is it worth?

Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.

No public deployment has disclosed a measurable outcome yet.

Value drivers: Speed and cycle time, Risk and loss reduction, Inclusion and access.

Indicative value

A biopharma sponsor starting 10 trials a year

USD 3 million to USD 90 million

Value of enrollment delay avoided per year per year

How this is calculated

Formula: trialsPerYear * enrollmentDaysSaved * costPerDayOfDelay. The low scenario uses every low input, the high scenario every high input.

InputLowHighBasis
Trials started per year trialsPerYear, trials per year515Editorial assumption for a mid sized biopharma sponsor. Replace with your own pipeline.
Enrollment days saved per trial from better site selection enrollmentDaysSaved, days per trial2060Conservative against the evidence on this page: Amgen reports that sites at the top of its model's ranked lists enrolled participants up to three times faster, on average, than lower ranked sites, and a Novartis pilot reported that investigators tagged fast starting and high performing recruited at 3.4 times the median rate. This range covers only part of the enrollment period, since site selection is one of several factors that determine how long enrollment takes.
Cost of a day of delayed enrollment costPerDayOfDelay, USD per day30,000100,000Editorial assumption, replace with your own trial's fully loaded daily cost, which varies widely by phase and therapeutic area.

What it leaves out: Gross value of faster enrollment only. It leaves out the cost of building and maintaining the site selection model, the value of a more representative patient population, and trials where a factor other than site selection is what actually sets the critical path.

Who already uses it?

2 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.

Amgen

United States · Pharma and life sciences · 2025

ProductionGrade B

Amgen's own data scientists, engineers and analysts built ATOMIC (Analytical Trial Optimization Module), a machine learning model that analyzes hundreds of factors, including eligibility criteria, treatment duration, historical site enrollment and anonymized real world data such as electronic health records and medical claims, to generate ranked lists of candidate clinical trial sites by predicted enrollment rate, alongside relevant country and investigator data. ATOMIC was first piloted for an ulcerative colitis trial and has since been used for site selection across cardiometabolic disease, atopic dermatitis and multiple cancer trials, alongside a separate scorecard, built with Amgen's Representation in Clinical Research (RISE) team, that flags sites with higher concentrations of historically underrepresented populations.

No outcome disclosed.

Novartis

Switzerland · Pharma and life sciences · 2024

PilotGrade C

Novartis leaders described at the DPHARM 2024 conference how the company built its own AI algorithms and a "Unified Ontology" data platform, drawing on data from 460,000 clinical trials, more than 700,000 clinical sites and 600,000 principal investigators across the industry, to analyze large datasets in minutes and identify suitable trial sites and investigators, work that traditionally could take weeks or months of manual review. In a pilot on a roughly 1,700 patient trial in the United States, principal investigators the model tagged as fast starting and high performing recruited at 3.4 times the median rate, and a separately tagged group of highly representative investigators recruited 2.7 times more Black or African American patients than their peers.

No outcome disclosed.

How do you implement it?

A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.

Data you need

  • Historical site and investigator performance data (enrollment rates, screen failure rates, protocol deviations) from the sponsor's own past trials
  • Anonymized or aggregated real world data showing where the eligible patient population is treated today
  • The draft protocol's eligibility criteria, treatment duration and visit schedule

Systems to integrate

  • Clinical trial management system (CTMS) for site and investigator history
  • Electronic data capture (EDC) platform for enrollment milestones
  • A real world data or claims data provider
  • The sponsor's feasibility survey or site portal, so sites still confirm interest and capacity

Complexity: Medium

The ranking model itself is contained, but it needs access to sensitive real world data (claims or health record extracts) alongside the sponsor's own trial systems, a review process for the study team to actually act on a ranked list, and retraining as trial designs and regions change.

  1. 1

    Centralize site and investigator history

    Pull enrollment, screen failure and protocol deviation history for every site and investigator the sponsor has worked with out of the CTMS and EDC, instead of relying on what one study team remembers from its last trial.

  2. 2

    Add real world data for the target population

    Layer in anonymized claims or health record data to see where patients who match the trial's eligibility criteria are actually treated, including sites and countries the sponsor has no history with.

  3. 3

    Train the model on comparable trials

    Score candidate sites against trials with a similar indication, phase and design, so the ranking predicts enrollment rate rather than simply repeating whichever sites enrolled the most patients historically, regardless of trial type.

  4. 4

    Route the ranked list into feasibility, not around it

    Send the model's shortlist and its reasoning into the standard feasibility survey and site selection committee; do not let a rank substitute for a site confirming its own interest and capacity.

  5. 5

    Track predicted against actual enrollment

    Compare predicted versus actual enrollment per site every quarter, and retrain or flag low confidence when a therapeutic area, region or trial design drifts from what the model was trained on.

Guardrails

  • Every recommendation carries the reasoning behind it (the data points behind the rank) for the site selection committee to review, never a bare score
  • Sites confirm their own interest, capacity and any competing trials; the model narrows the list, it does not commit a site
  • Real world data used for scoring is anonymized or aggregated before it reaches the model

KPIs to instrument

  • Predicted versus actual enrollment rate per site, reviewed by quarter
  • Share of activated sites that were in the model's top tier
  • Number of new sites identified that the sponsor had no prior history with
  • Time from protocol finalization to first site activated

Human in the loop

The clinical operations team and the site selection committee make the final call on every site. The model's job is to surface sites the team would not otherwise have found and to flag when a familiar site's recent performance has slipped; it does not activate a site by itself.

Common failure modes

Stale or narrow training data
A model trained mostly on one therapeutic area or region ranks poorly for a new one. Retrain per indication family and flag low confidence predictions instead of a bare rank.
Representativeness dropped under time pressure
Teams under enrollment pressure default back to the fastest known sites, which are rarely the most representative. Give the committee a representativeness view alongside the speed ranking, not as an afterthought.
A ranking treated as a commitment
A high ranked site still needs its own confirmed capacity and interest. Keep the feasibility survey mandatory before any site is activated, whatever its rank.

What are the risks and rules?

EU AI Act

Depends on design

Ranking candidate sites as institutions, by data such as facility capacity and historical trial throughput, is an internal research operations decision and stays minimal risk: it does not by itself decide a natural person's access to healthcare, credit, employment or another Annex III listed area. But the same models often also score and tag named principal investigators on performance, speed and representativeness, as in the Novartis evidence on this page, to help decide which investigator's site gets a trial contract. Annex III point 4 covers AI used to evaluate the performance and behaviour of natural persons in a work related context, so a deployment that profiles named investigators, not only institutions, may fall inside that point depending on how the sponsor uses the score. Sponsors should classify each deployment on whether it scores named investigators or only sites, and treat investigator scoring as high risk until a legal review of the specific design says otherwise.

Rules that apply

Controls to put in place

  • Inventory entry for the model with an accountable owner in clinical operations
  • A data use agreement and anonymization review for every real world data source before it reaches the model
  • Documented validation of the ranking against at least one completed trial before it is used to choose sites for a new one
  • A GDPR review of investigator profiling, since scoring named principal investigators processes their personal performance data separately from any patient data

Frequently asked questions

Does AI site selection replace feasibility surveys?
No. It narrows a long list of possible sites to a short list worth surveying, based on predicted enrollment and where the eligible patient population is actually treated. The site still has to confirm its own interest and capacity before it is activated.
How much faster does AI ranked site selection enroll a trial?
Amgen reports that sites at the top of its ATOMIC model's ranked lists enrolled participants up to three times faster, on average, than sites lower on the list, across 13 studies sponsored by Amgen and analyzed in 2025. Novartis has reported a US pilot in which investigators the model rated as high performing and fast starting recruited at 3.4 times the median rate. Both figures compare sites within one sponsor's own trials, not against an industry average.
Can this improve diversity in clinical trial enrollment?
It can help find sites with access to underrepresented populations that a sponsor has no prior history with, which is why Amgen built a separate representativeness scorecard into ATOMIC alongside its enrollment speed ranking. It does not by itself fix a trial design that excludes a population; that is a protocol decision.
What data does a site selection model need?
Its own past trial and site performance data (enrollment rates, screen failures, protocol deviations), plus anonymized or aggregate real world data such as insurance claims or electronic health records, to see where the eligible patient population is actually treated.

How to cite this page

Blits.ai AI Use Case Library, "AI for clinical trial site selection and feasibility", last verified 29 September 2026, https://www.blits.ai/ai-use-cases/clinical-trial-site-selection-and-feasibility. Licensed under CC BY 4.0. Method: how we verify use cases.

Changelog
  • 29 September 2026: First published

Related use cases

HealthcarePharma and life sciences

AI clinical trial patient matching and prescreening

AI that reads structured data and clinical notes in the health record, compares each patient with the inclusion and exclusion criteria of open clinical trials, and gives research staff and treating clinicians a ranked list of likely eligible patients with the evidence for each criterion, so that people confirm eligibility and invite the patient.

Deployments
3 public, best grade B
Reported accuracy
100%
Cleveland Clinic, organization claim
Pharma and life sciences

AI native platform for drug target discovery and molecule design

An AI native research platform that prioritizes disease targets from biological data, generates and optimizes candidate drug molecules computationally, and predicts their properties before a chemist synthesizes and tests them, so a pharmaceutical or biotech company reaches a validated preclinical candidate with far fewer molecules made and tested than a conventional medicinal chemistry program.

Deployments
2 public, best grade B
Reported cycle time reduction
about 60%
Insilico Medicine, organization claim
Energy and utilities

AI analytics for smart meter and AMI data

AI that turns the flood of readings from smart electricity, gas and water meters into usable information: it monitors meter and network health at scale, estimates which appliances drive a household's usage from the meter signal alone, flags unusual consumption, and targets efficiency and electrification programmes at the customers who will benefit most, instead of a utility treating every meter and every customer the same way.

Deployments
2 public, best grade C
Autonomy
Assist
Healthcare

AI command center for hospital bed and staff capacity planning

An AI powered operations center that predicts patient admissions, discharges and transfers across a hospital or health system, and helps a team of coordinators sitting in one room sequence real time bed assignments, staffing levels and patient moves, so patients get into the right bed faster and existing capacity is used fully without adding beds.

Deployments
2 public, best grade B
Reported cycle time reduction
38%
Johns Hopkins Medicine, organization claim