AI use case

AI drafting of clinical study reports and regulatory documents

Generative AI that drafts clinical study reports and other regulated documents, such as protocols, patient materials and submission modules, from the trial's statistical tables, listings and figures and from approved template text, for medical writers to verify, edit and approve before anything is submitted to a regulator.

By Len Debets · Last verified 27 September 2026 · 2 public deployments

50%
Reported error reduction
Merck & Co., organization claim.
90%
Reported handling time reduction
Novo Nordisk, organization claim.
USD 57,600 to USD 297,000
Indicative value per year
A drug developer that writes 20 clinical study reports a year. Worked example, see how it is calculated.

What problem does it solve?

A clinical study report describes the design, conduct and results of a trial in the structure regulators expect, and can run to 300 pages. Medical writers assemble it from thousands of pages of tables, listings and figures, reconcile conflicting numbers, and write narrative in precise regulatory language, followed by rounds of review. The same pattern repeats for protocols, investigator brochures, safety narratives, patient materials and the modules of a marketing application.

The work sits on the critical path to approval, so weeks spent writing are weeks a medicine is not with patients, and each report ties up experienced writers for a long time. Manual drafting is also prone to errors, and every error has to be found and corrected in review. Generative AI fits the task because much of the text restates structured results in standard language, but the output must be exactly right: a transposed number or an invented statement in a regulatory document is a serious quality failure.

How does it work?

  1. Ingest the study outputs. Statistical tables, listings and figures, the protocol and the statistical analysis plan are loaded and preprocessed, so each table can be read reliably.
  2. Map content to the template. Each section of the report template is linked to the tables and approved boilerplate it needs, following the house structure and the ICH format.
  3. Draft section by section. A language model writes each section from its mapped sources, with every number traced back to the table it came from and approved text reused where possible.
  4. Check automatically. Rules compare numbers in the text with the source tables, flag missing sections and check terminology and style before a person sees the draft.
  5. Review and approve. Medical writers and study experts review, edit and interpret the results; the document follows the normal quality control and approval process before submission.
Audience
Employee facing
Autonomy
Copilot
Adoption
Early adopters
Channels
Internal tools

What is it worth?

Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.

Value benchmarks for AI drafting of clinical study reports and regulatory documents
KPIMedianReported rangeData pointsClaimed by
Error reductionToo few to pool
50%
11 organization
Handling time reductionToo few to pool
90%
11 organization

Value drivers: Speed and cycle time, Employee productivity, Compliance quality, Lower cost to serve.

Indicative value

A drug developer that writes 20 clinical study reports a year

USD 57,600 to USD 297,000

Medical writing effort released on first drafts per year

How this is calculated

Formula: reports * hoursPerDraft * timeSaved * costPerHour. The low scenario uses every low input, the high scenario every high input.

InputLowHighBasis
Clinical study reports per year reports, reports per year2020The reference developer.
Hours to a fully human reviewed first draft today hoursPerDraft, hours per report120180The high end is Merck's reported average of 180 hours before its platform. The low end is an editorial assumption, replace both with your own time data.
Share of drafting time saved timeSaved, fraction of drafting hours0.30.55The high end is Merck's achieved result: Merck went from an average of 180 to 80 hours for a fully human reviewed first draft (about 55% less). The low end is an editorial assumption for a less mature platform. Novo Nordisk reports 90% less writing time, which excludes review and approval, so it is not used as the high end.
Fully loaded cost per medical writer hour costPerHour, USD per hour80150Editorial assumption covering internal writers and agency rates. Replace with your own.

What it leaves out: Drafting effort only. It leaves out the value of earlier submissions, which for a marketed medicine can far exceed the writing cost, review and quality control time that remains, the cost of building and validating the platform, and other document types.

Who already uses it?

2 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.

Merck & Co.

United States · Pharma and life sciences · 2025

ProductionGrade B

Merck built an internal generative AI platform that combines table preprocessing with large language model authoring to produce first drafts of clinical study reports, under the oversight of qualified medical writers. Merck reports that first drafts now take three to four days instead of two to three weeks, that the time to a fully human reviewed first draft fell from an average of 180 hours to 80 hours, and that draft errors halved. The first live reports built on the platform were submitted in 2025, and Merck said it was scaling the platform across its late phase pipeline.

  • Error reduction: 50%, CSR first drafts, across multiple studies
    "Increased the quality of CSR drafts – as measured by reducing the number of errors by 50% – in categories such as data, messaging, citations, terminology and typography."
    Claimed by: organization

Novo Nordisk

Denmark · Pharma and life sciences · 2025

ProductionGrade C

Novo Nordisk built NovoScribe, a documentation platform that combines retrieval augmented generation over expert approved text with case specific variables to draft clinical study reports. It runs on Amazon Bedrock and MongoDB Atlas with Claude models, and has been extended to device verification protocols and patient materials. Anthropic's case study quotes Novo Nordisk saying writing times on clinical study reports fell by 90%, with drafts going to people for review and approval; the company aims to extend it to full Common Technical Documents.

  • Handling time reduction: 90%, writing time per clinical study report
    "“Claude has helped us cut writing times on CSRs by 90% so we can get documentation directly into human hands for review and approval,” said Waheed Jowiya, Digitalization Strategy Director at Novo Nordisk."
    Claimed by: organization

How do you implement it?

A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.

Data you need

  • Statistical tables, listings and figures in a consistent, machine readable format
  • Approved templates and boilerplate text per document type, with owners
  • Past reports and reviewer comments to test against
  • A style guide and terminology list

Systems to integrate

  • Statistical computing environment that produces the study outputs
  • Document management and regulatory information management systems
  • Quality management system for review, approval and electronic signatures
  • Submission publishing tools

Complexity: High

The hard parts are reliable extraction from complex statistical tables, traceability from every sentence to its source, validation of the system under quality management rules, and changing the writing process. Merck reports a team of more than 80 people across data science, AI and medical expertise; Novo Nordisk spent months integrating legacy systems for device protocols.

  1. 1

    Start with one document type and its most formulaic sections

    Clinical study report results sections that restate tables are the usual starting point; leave interpretation, discussion and conclusions to writers at first.

  2. 2

    Make every number traceable

    Store the link between each sentence and the table cell it uses, and check numbers automatically, so reviewers verify instead of recalculating.

  3. 3

    Reuse approved text

    Put expert approved boilerplate and templates in a governed library with versions, and generate only what changes per study.

  4. 4

    Validate under your quality system

    Treat the platform as a computerized system used in a regulated process: define its intended use, validate it on real studies, and control changes to prompts, models and templates.

  5. 5

    Redesign the review, then scale

    Train writers to review generated drafts, measure hours, error rates and review cycles per report, and extend to further document types once quality is stable.

Guardrails

  • Every generated number checked automatically against its source table before human review
  • No document submitted without review and approval by qualified medical writers and study experts
  • Interpretation of results and conclusions written or approved by named experts
  • Unpublished trial data processed only in approved, access controlled environments
  • Change control and revalidation for prompts, models and templates

KPIs to instrument

  • Writer hours to a human reviewed first draft, per document type
  • Errors found in quality control per draft, by category (data, terminology, citations)
  • Review cycles and elapsed days from database lock to final report
  • Share of generated text kept after review
  • Regulator questions attributable to document quality

Human in the loop

Medical writers own every document: they review each section against its sources, edit the narrative and interpretation, and take the draft through the normal quality control and approval workflow. Study clinicians and statisticians approve the interpretation of results.

Common failure modes

Numbers that do not match the tables
A model transposes, rounds or invents a figure. Check every number against its source automatically and block drafts that fail.
Plausible but wrong interpretation
The draft states a conclusion the data do not support. Keep interpretation and discussion with experts and mark generated interpretive text clearly.
Faster drafts, same timeline
Drafting gets quicker but review and approval stay slow. Train reviewers to work with generated drafts, and measure review cycles and elapsed days, not only drafting hours. Merck reports that it revamped its operations and trained teams in the skills needed to oversee its platform.
Unvalidated changes
A model or prompt update changes output quality unnoticed. Regression test on reference studies before every change.

What are the risks and rules?

EU AI Act

Depends on design

Drafting regulated documents for expert review is not listed in Annex III and is not a practice prohibited by Article 5, so the tier turns on the sponsor's role under Article 50. A sponsor that deploys a third party drafting tool has no specific AI Act obligations beyond AI literacy: the Article 50(4) disclosure duty covers AI generated text published to inform the public on matters of public interest, which clinical study reports and regulatory submissions are not. For that sponsor the tier is minimal. A sponsor that builds its own generating system, as Merck (a proprietary platform) and Novo Nordisk (NovoScribe) did, is its provider under Article 50(2) and must mark the synthetic text in a machine readable format, unless the exemption for an assistive function for standard editing applies; drafting whole report sections goes beyond that exemption, so for that sponsor the tier is limited. Quality expectations come from medicines regulation and EMA guidance: the EMA reflection paper expects close human supervision and quality review when AI drafts medicinal product information documents, and makes the clinical trial sponsor, marketing authorisation applicant or holder, or manufacturer responsible for ensuring that models and data pipelines are fit for purpose and meet GxP standards and EMA guidelines.

Guidance

Controls to put in place

  • Documented intended use, validation report and change control for the drafting platform
  • Traceability from generated text to source tables and approved templates
  • Automated numeric consistency checks with results kept in the document history
  • Access controls and logging for unpublished trial data
  • Quality control sampling of generated sections by experienced writers

Frequently asked questions

How much faster does AI make clinical study reports?
Merck reports that first drafts now take three to four days instead of two to three weeks, and that the time to a fully human reviewed first draft fell from an average of 180 to 80 hours. Novo Nordisk says writing times on clinical study reports fell by 90%. Review and approval still take time, so measure the whole cycle, not only drafting.
What do regulators expect for AI drafted submissions?
Responsibility stays with the company that uses the AI: the EMA reflection paper makes the clinical trial sponsor, marketing authorisation applicant or holder, or manufacturer responsible for ensuring that models and data pipelines are fit for purpose and meet GxP standards and EMA guidelines. For medicinal product information documents, the EMA asks for close human supervision and quality review so that all generated text is factually and syntactically correct. The FDA's January 2025 draft guidance on AI for regulatory decision making does not address AI used to draft a submission when it does not affect patient safety, drug quality or the reliability of study results.
Does AI reduce errors or add them?
Both are possible. Merck reports 50% fewer errors in drafts across data, messaging, citations, terminology and typography, but language models can also produce plausible wrong numbers, which is why automatic checks against the source tables matter.

How to cite this page

Blits.ai AI Use Case Library, "AI drafting of clinical study reports and regulatory documents", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/clinical-and-regulatory-document-drafting. Licensed under CC BY 4.0. Method: how we verify use cases.

Changelog
  • 27 September 2026: First published

Related use cases

Pharma and life sciencesGovernment and public sector

AI for pharmacovigilance adverse event case intake

AI that takes in adverse event reports about medicines, vaccines and devices from calls, emails, forms, literature and partner files, decides whether each is a valid case, flags seriousness, extracts and codes the case data into the safety database format, and routes it to drug safety professionals, who review medical content and regulatory reporting.

Deployments
4 public, best grade B
Autonomy
Supervised agent
HealthcarePharma and life sciences

AI clinical trial patient matching and prescreening

AI that reads structured data and clinical notes in the health record, compares each patient with the inclusion and exclusion criteria of open clinical trials, and gives research staff and treating clinicians a ranked list of likely eligible patients with the evidence for each criterion, so that people confirm eligibility and invite the patient.

Deployments
3 public, best grade B
Reported accuracy
100%
Cleveland Clinic, organization claim
Cross industryBanking

AI for drafting customer letters and outbound notices

AI that drafts the letters and notices operations must send at scale, such as arrears notices, decline letters, complaint responses, servicing confirmations and product change notices, from case data and approved templates and clauses, in the customer's language, for a person to approve where the notice is regulated.

Deployments
5 public, best grade B
Reported cycle time reduction
25%
SS&C GIDS and RS, vendor claim
BankingWealth and asset management

AI assistant for Shariah compliance screening and review

An AI assistant that screens Islamic financing contracts, deal structures and investments for Shariah compliance risks such as riba, gharar and exposure to prohibited activities, retrieves the relevant standards and fatwas, drafts the Shariah review documentation and flags issues for the Shariah board, which keeps sole authority over any ruling.

Deployments
1 public, best grade B
Autonomy
Copilot