What problem does it solve?
Many clinical trials struggle to enroll enough patients, and slow enrollment extends trial timelines. Eligibility criteria are long and specific (biomarkers, prior treatments, lab values, stage), and the facts needed to check them are scattered across notes, pathology reports and lab results. Research coordinators screen charts by hand, one trial and one patient at a time, so they see only a fraction of the patients who might qualify, and patients treated outside the flagship hospital are considered less often.
The result is lost opportunity on both sides: patients are not offered trials that could help them, sponsors wait longer for results, and trial participation stays concentrated at flagship academic sites. Language models can read clinical notes at scale, but eligibility errors in either direction matter, so the design has to keep people in charge of the final decision.
- Researchers at Yale Cancer Center write that cancer clinical trial enrollment remains critically low at 5% to 7% of adult patients, despite exponential growth in the number of available trials.Clinical Trial Patient Matching: A Real-Time, Common Data Model and Artificial Intelligence-Driven System for Semiautomated Patient Prescreening in Cancer Clinical Trials (2026)
How does it work?
- Encode the criteria. Each trial's inclusion and exclusion criteria are turned into checks, some structured (age, diagnosis codes, lab values) and some that need the notes (prior lines of therapy, performance status, biomarkers).
- Find the population. Rules on structured data narrow the whole patient population to candidates, for example patients with a relevant diagnosis and an upcoming visit.
- Read the record. A language model or NLP pipeline reads notes and reports for each candidate and marks each criterion as met, not met or unknown, with the passage that supports it.
- Present a ranked list. Research staff and treating clinicians see likely eligible patients and the evidence, ideally before the patient's next visit.
- Confirm and invite. Staff verify eligibility in the record, discuss the trial with the treating clinician and approach the patient; outcomes feed back to improve the criteria and the model.
- Audience
- Employee facing
- Autonomy
- Assist
- Adoption
- Early adopters
- Channels
- Internal tools
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
| KPI | Median | Reported range | Data points | Claimed by |
|---|---|---|---|---|
| Interactions handled | Not pooled | 904 to 98,348 | 2 | 2 organization |
| Accuracy | Too few to pool | 100% | 1 | 1 organization |
| Handling time reduction | Too few to pool | 41% | 1 | 1 organization |
Value drivers: Inclusion and access, Speed and cycle time, Employee productivity.
Indicative value
A cancer center whose research staff prescreen 15,000 charts a year
USD 12,000 to USD 210,000
Research staff screening time released per year
How this is calculated
Formula: charts * minutesPerChart / 60 * reviewAvoided * costPerHour. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Charts prescreened by hand per year charts, charts per year | 15,000 | 15,000 | Editorial assumption for the reference cancer center, replace with your own. |
| Minutes of manual review per chart minutesPerChart, minutes per chart | 3 | 15 | Yale Cancer Center measured 3.1 minutes per chart for its prescreening workflow. Cleveland Clinic writes that a manual chart review can take more than 30 minutes per record, depending on the complexity of the criteria and the volume of history; the high value is an editorial assumption well below that. Replace with your own time data. |
| Share of manual chart review avoided reviewAvoided, fraction of review minutes | 0.4 | 0.8 | Conservative against Yale Cancer Center's report of a tenfold reduction in chart review workload and 41% less screening time per chart that was still reviewed. |
| Fully loaded cost per research coordinator hour costPerHour, USD per hour | 40 | 70 | Editorial assumption. Replace with your own. |
What it leaves out: Screening effort only, and likely small next to the main value: more patients offered trials and faster enrollment, neither of which is included. It also leaves out the cost of encoding criteria, integrating the record and validating the tool.
Who already uses it?
3 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
Mount Sinai Health System
United States · Healthcare · 2026
The Mount Sinai Tisch Cancer Center deployed PRISM, an oncology specific trial matching platform from Triomics built on its OncoLLM language model pipeline, across the Mount Sinai Health System in January 2026. The platform reviews patient records against trial protocols so that patients seen at other hospitals in the system, such as Mount Sinai Queens and Mount Sinai Brooklyn, have the same access to trials as those treated at The Mount Sinai Hospital. Mount Sinai says the aim is also to let clinicians focus on conversations with patients rather than manual chart review; the January 2026 release reports no outcome figures. Mount Sinai said it would evaluate outcomes and publish them later.
No outcome disclosed.
Cleveland Clinic
United States · Healthcare · 2025
Cleveland Clinic researchers used Dyania Health's Synapsis platform, a medically trained language model system embedded in the electronic medical record, to prescreen patients for a phase 3 polycythemia vera trial. From 4.7 million active records it identified 28,200 patients with an oncology diagnosis in the past three years, narrowed them to 904 patients with polycythemia vera, assessed each against the trial's seven eligibility and 20 exclusion criteria within one week and found 22 eligible patients, all confirmed by research staff (100% positive predictive value). The usual workflow had prescreened nine patients and enrolled four over twelve months. Looking ahead, Cleveland Clinic and Dyania Health have announced a collaboration to integrate the platform across the health system's clinical research enterprise; Cleveland Clinic has also invested in Dyania Health.
- Interactions handled: 904, patients assessed against the trial criteria in one week
"“The AI tool completed full eligibility assessments on these 904 patients within one week, against the trial’s criteria and identified 22 eligible patients,” Dr. Gerds and his colleagues reported."
Claimed by: organization - Accuracy: 100%, of the 22 patients identified as eligible, confirmed by research staff (positive predictive value)
"In a study presented at the 2025 American Society of Hematology (ASH) Annual Meeting, investigators reported that the Dyania Health’s Synapsis™ AI platform, an artificial intelligence tool, identified seven times more eligible patients for a polycythemia vera trial than standard workflows, while achieving 100% positive predictive value following research-staff verification."
Claimed by: organization
Yale Cancer Center
United States · Healthcare · 2022
Yale Cancer Center built a clinical trial patient matching (CTPM) tool that combines rules with natural language processing over structured and unstructured record data standardized to the OMOP common data model. Validated first on one metastatic colorectal cancer trial, it was then implemented across 29 trials in several cancer specialties. Since September 2022 it has screened 98,348 patients, identified 825 eligible candidates and contributed to 117 enrollments, and it cut screening time per reviewed chart by 41%. The prescreening is semiautomated: research teams review the candidates the tool puts forward instead of reading every chart.
- Handling time reduction: 41%, screening time per chart for patients who underwent review
"Implementation reduced chart review workload 10-fold and screening time by 41% (3.1 to 1.8 minutes per chart) for those patients who did undergo review."
Claimed by: organization - Interactions handled: 98,348, patients screened since September 2022, across 29 trials
"Since September 2022, the system has screened 98,348 patients across 29 trials, identifying 825 eligible candidates and facilitating 117 patient enrollments with 9%-37% consent rates."
Claimed by: organization
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- Access to structured record data and clinical notes, pathology and lab reports
- Trial protocols with inclusion and exclusion criteria, and a list of open trials
- Visit schedules to time outreach before appointments
- Past screening decisions to measure accuracy
Systems to integrate
- Electronic health record or clinical data warehouse (for example on the OMOP model)
- Clinical trial management system for trials and enrollment status
- Research staff worklists and secure messaging to treating clinicians
Complexity: Medium
Reading the record is feasible with current models; the effort goes into access to structured and unstructured data (often through a common data model such as OMOP), turning free text criteria into checks for each new trial, and fitting the output into how research teams already work.
- 1
Start with trials that struggle to enroll
Pick a few open trials with clear criteria and slow accrual, and measure how many eligible patients the current process finds.
- 2
Split criteria into structured and text checks
Use structured data to narrow the population cheaply and reserve language model reading for criteria that only the notes contain.
- 3
Show evidence per criterion
For every criterion show met, not met or unknown and the passage behind it, so staff can verify in seconds rather than rereading the chart.
- 4
Measure against manual screening
Compare accuracy, missed patients and time per chart with the manual process on the same trials, as Yale Cancer Center and Cleveland Clinic did, before scaling.
- 5
Scale across trials and sites
Add trials and community sites, and check that patients at every site and in every group are identified at similar rates.
Guardrails
- Eligibility is always confirmed by research staff or the investigator before a patient is approached
- The treating clinician is involved before any patient contact
- Evidence shown for every criterion, with unknowns marked rather than guessed
- Access to records for prescreening limited to what research rules and local approvals allow
- Identification rates monitored by site, sex, age and ethnicity to catch unequal access
KPIs to instrument
- Eligible patients identified per trial per month, compared with manual screening
- Screening minutes per chart and charts reviewed per enrollment
- Accuracy of eligibility suggestions on a verified sample, including missed eligible patients
- Enrollment and time to first patient per trial
- Identification and enrollment rates by site and demographic group
Human in the loop
Research coordinators and investigators decide who is eligible and who is approached, together with the treating clinician. The tool prioritizes and explains; it never enrolls or contacts a patient itself.
Common failure modes
- Missed eligible patients
- Criteria encoded too strictly, or facts hidden in scanned documents, exclude patients who qualify. Measure sensitivity against manual screening, not only precision.
- Confident but wrong eligibility
- The model marks a criterion as met from an outdated or negated note. Show the source passage and date and require verification.
- More candidates, no more enrollments
- Lists grow but staff and clinicians have no time to act. Fit the output to visit schedules and worklists and track enrollments, not matches.
- Conflicts of interest in evaluation
- Health systems that invest in the vendor they evaluate may overstate results. Look for independent or prospective evaluations.
What are the risks and rules?
EU AI Act
Depends on design
Prescreening for research that staff verify is not listed in Annex III and is usually minimal risk. The Article 2(6) exclusion covers only systems developed and put into service for the sole purpose of scientific research and development, so an operational recruitment tool used across a health system usually falls inside the Act. If the software recommends trials to a clinician as a treatment option for an individual patient, it may qualify as medical device software under the Medical Device Regulation; where that needs a notified body assessment, it is high risk under Article 6(1). Processing health records for research falls under GDPR Article 9 and national research rules.
Rules that apply
Guidance
- Reflection paper on the use of artificial intelligence (AI) in the medicinal product lifecycle (European Medicines Agency, Europe). Says AI used in clinical trials should meet applicable ICH E6 good clinical practice requirements, and that where a use could have high regulatory impact or high patient risk and the method has not been previously qualified by the EMA for that context of use, the model documentation may be treated as clinical trial data and requested at marketing authorization, clinical trial application or GCP inspection. A reflection paper, not binding, and general to AI in clinical trials rather than specific to recruitment.
- Regulation (EU) No 536/2014 on clinical trials on medicinal products for human use (European Union, Europe). Sets the EU rules on trial conduct, informed consent and subject protection that recruitment processes supported by AI must respect.
Controls to put in place
- Documented approval of prescreening under the institution's research governance and privacy rules
- Versioned criteria per trial with an owner and review against protocol amendments
- Audit trail of suggestions, staff decisions and patient contacts
- Periodic accuracy and fairness review per trial and site
Frequently asked questions
- Does AI trial matching increase enrollment?
- The deployments on this page do not report enrollment with and without the tool. Yale Cancer Center's tool screened 98,348 patients across 29 trials since September 2022 and facilitated 117 enrollments, and Cleveland Clinic found 22 eligible patients for a rare disease trial in one week, where the usual process had prescreened nine in a year. Mount Sinai deployed matching systemwide in 2026 and has promised published results; until comparisons are published, treat an enrollment increase as something to measure, not to assume.
- How accurate is it?
- Good enough to prioritize, not to decide. On one trial, Cleveland Clinic's research staff confirmed all 22 patients the system identified as eligible (100% positive predictive value), but missed eligible patients are not reported on the Cleveland Clinic page. On its validation trial Yale Cancer Center reports 94% retrospective and 88% prospective accuracy with 100% sensitivity. Both keep staff verification in the loop, so measure missed eligible patients as well as false matches.
- What should we watch out for?
- Unequal identification across sites and groups, criteria that fall out of date after protocol amendments, and evaluations run by organizations with a financial interest in the vendor, as Cleveland Clinic discloses for its investment in Dyania Health.
How to cite this page
Blits.ai AI Use Case Library, "AI clinical trial patient matching and prescreening", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/clinical-trial-patient-matching. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 27 September 2026: First published