What problem does it solve?
A clinical study report describes the design, conduct and results of a trial in the structure regulators expect, and can run to 300 pages. Medical writers assemble it from thousands of pages of tables, listings and figures, reconcile conflicting numbers, and write narrative in precise regulatory language, followed by rounds of review. The same pattern repeats for protocols, investigator brochures, safety narratives, patient materials and the modules of a marketing application.
The work sits on the critical path to approval, so weeks spent writing are weeks a medicine is not with patients, and each report ties up experienced writers for a long time. Manual drafting is also prone to errors, and every error has to be found and corrected in review. Generative AI fits the task because much of the text restates structured results in standard language, but the output must be exactly right: a transposed number or an invented statement in a regulatory document is a serious quality failure.
- Anthropic's case study reports that Novo Nordisk's staff writers averaged only 2.3 clinical study reports per year.Novo Nordisk accelerates clinical documentation and drug development with Claude (2025)
How does it work?
- Ingest the study outputs. Statistical tables, listings and figures, the protocol and the statistical analysis plan are loaded and preprocessed, so each table can be read reliably.
- Map content to the template. Each section of the report template is linked to the tables and approved boilerplate it needs, following the house structure and the ICH format.
- Draft section by section. A language model writes each section from its mapped sources, with every number traced back to the table it came from and approved text reused where possible.
- Check automatically. Rules compare numbers in the text with the source tables, flag missing sections and check terminology and style before a person sees the draft.
- Review and approve. Medical writers and study experts review, edit and interpret the results; the document follows the normal quality control and approval process before submission.
- Audience
- Employee facing
- Autonomy
- Copilot
- Adoption
- Early adopters
- Channels
- Internal tools
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
| KPI | Median | Reported range | Data points | Claimed by |
|---|---|---|---|---|
| Error reduction | Too few to pool | 50% | 1 | 1 organization |
| Handling time reduction | Too few to pool | 90% | 1 | 1 organization |
Value drivers: Speed and cycle time, Employee productivity, Compliance quality, Lower cost to serve.
Indicative value
A drug developer that writes 20 clinical study reports a year
USD 57,600 to USD 297,000
Medical writing effort released on first drafts per year
How this is calculated
Formula: reports * hoursPerDraft * timeSaved * costPerHour. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Clinical study reports per year reports, reports per year | 20 | 20 | The reference developer. |
| Hours to a fully human reviewed first draft today hoursPerDraft, hours per report | 120 | 180 | The high end is Merck's reported average of 180 hours before its platform. The low end is an editorial assumption, replace both with your own time data. |
| Share of drafting time saved timeSaved, fraction of drafting hours | 0.3 | 0.55 | The high end is Merck's achieved result: Merck went from an average of 180 to 80 hours for a fully human reviewed first draft (about 55% less). The low end is an editorial assumption for a less mature platform. Novo Nordisk reports 90% less writing time, which excludes review and approval, so it is not used as the high end. |
| Fully loaded cost per medical writer hour costPerHour, USD per hour | 80 | 150 | Editorial assumption covering internal writers and agency rates. Replace with your own. |
What it leaves out: Drafting effort only. It leaves out the value of earlier submissions, which for a marketed medicine can far exceed the writing cost, review and quality control time that remains, the cost of building and validating the platform, and other document types.
Who already uses it?
2 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
Merck & Co.
United States · Pharma and life sciences · 2025
Merck built an internal generative AI platform that combines table preprocessing with large language model authoring to produce first drafts of clinical study reports, under the oversight of qualified medical writers. Merck reports that first drafts now take three to four days instead of two to three weeks, that the time to a fully human reviewed first draft fell from an average of 180 hours to 80 hours, and that draft errors halved. The first live reports built on the platform were submitted in 2025, and Merck said it was scaling the platform across its late phase pipeline.
- Error reduction: 50%, CSR first drafts, across multiple studies
"Increased the quality of CSR drafts – as measured by reducing the number of errors by 50% – in categories such as data, messaging, citations, terminology and typography."
Claimed by: organization
Novo Nordisk
Denmark · Pharma and life sciences · 2025
Novo Nordisk built NovoScribe, a documentation platform that combines retrieval augmented generation over expert approved text with case specific variables to draft clinical study reports. It runs on Amazon Bedrock and MongoDB Atlas with Claude models, and has been extended to device verification protocols and patient materials. Anthropic's case study quotes Novo Nordisk saying writing times on clinical study reports fell by 90%, with drafts going to people for review and approval; the company aims to extend it to full Common Technical Documents.
- Handling time reduction: 90%, writing time per clinical study report
"“Claude has helped us cut writing times on CSRs by 90% so we can get documentation directly into human hands for review and approval,” said Waheed Jowiya, Digitalization Strategy Director at Novo Nordisk."
Claimed by: organization
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- Statistical tables, listings and figures in a consistent, machine readable format
- Approved templates and boilerplate text per document type, with owners
- Past reports and reviewer comments to test against
- A style guide and terminology list
Systems to integrate
- Statistical computing environment that produces the study outputs
- Document management and regulatory information management systems
- Quality management system for review, approval and electronic signatures
- Submission publishing tools
Complexity: High
The hard parts are reliable extraction from complex statistical tables, traceability from every sentence to its source, validation of the system under quality management rules, and changing the writing process. Merck reports a team of more than 80 people across data science, AI and medical expertise; Novo Nordisk spent months integrating legacy systems for device protocols.
- 1
Start with one document type and its most formulaic sections
Clinical study report results sections that restate tables are the usual starting point; leave interpretation, discussion and conclusions to writers at first.
- 2
Make every number traceable
Store the link between each sentence and the table cell it uses, and check numbers automatically, so reviewers verify instead of recalculating.
- 3
Reuse approved text
Put expert approved boilerplate and templates in a governed library with versions, and generate only what changes per study.
- 4
Validate under your quality system
Treat the platform as a computerized system used in a regulated process: define its intended use, validate it on real studies, and control changes to prompts, models and templates.
- 5
Redesign the review, then scale
Train writers to review generated drafts, measure hours, error rates and review cycles per report, and extend to further document types once quality is stable.
Guardrails
- Every generated number checked automatically against its source table before human review
- No document submitted without review and approval by qualified medical writers and study experts
- Interpretation of results and conclusions written or approved by named experts
- Unpublished trial data processed only in approved, access controlled environments
- Change control and revalidation for prompts, models and templates
KPIs to instrument
- Writer hours to a human reviewed first draft, per document type
- Errors found in quality control per draft, by category (data, terminology, citations)
- Review cycles and elapsed days from database lock to final report
- Share of generated text kept after review
- Regulator questions attributable to document quality
Human in the loop
Medical writers own every document: they review each section against its sources, edit the narrative and interpretation, and take the draft through the normal quality control and approval workflow. Study clinicians and statisticians approve the interpretation of results.
Common failure modes
- Numbers that do not match the tables
- A model transposes, rounds or invents a figure. Check every number against its source automatically and block drafts that fail.
- Plausible but wrong interpretation
- The draft states a conclusion the data do not support. Keep interpretation and discussion with experts and mark generated interpretive text clearly.
- Faster drafts, same timeline
- Drafting gets quicker but review and approval stay slow. Train reviewers to work with generated drafts, and measure review cycles and elapsed days, not only drafting hours. Merck reports that it revamped its operations and trained teams in the skills needed to oversee its platform.
- Unvalidated changes
- A model or prompt update changes output quality unnoticed. Regression test on reference studies before every change.
What are the risks and rules?
EU AI Act
Depends on design
Drafting regulated documents for expert review is not listed in Annex III and is not a practice prohibited by Article 5, so the tier turns on the sponsor's role under Article 50. A sponsor that deploys a third party drafting tool has no specific AI Act obligations beyond AI literacy: the Article 50(4) disclosure duty covers AI generated text published to inform the public on matters of public interest, which clinical study reports and regulatory submissions are not. For that sponsor the tier is minimal. A sponsor that builds its own generating system, as Merck (a proprietary platform) and Novo Nordisk (NovoScribe) did, is its provider under Article 50(2) and must mark the synthetic text in a machine readable format, unless the exemption for an assistive function for standard editing applies; drafting whole report sections goes beyond that exemption, so for that sponsor the tier is limited. Quality expectations come from medicines regulation and EMA guidance: the EMA reflection paper expects close human supervision and quality review when AI drafts medicinal product information documents, and makes the clinical trial sponsor, marketing authorisation applicant or holder, or manufacturer responsible for ensuring that models and data pipelines are fit for purpose and meet GxP standards and EMA guidelines.
Rules that apply
Guidance
- ICH E3, Structure and Content of Clinical Study Reports (International Council for Harmonisation, Global). Sets the structure and content that a clinical study report follows for submission to regulators, the format the drafting tool's templates map against.
- Reflection paper on the use of artificial intelligence (AI) in the medicinal product lifecycle (European Medicines Agency, Europe). States that AI used for drafting, compiling or reviewing medicinal product information documents should be used under close human supervision, with quality review so that all generated text is factually and syntactically correct before submission.
- Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products (draft guidance) (US Food and Drug Administration, North America). The January 2025 draft, still marked as draft guidance on the FDA site in September 2026, does not address AI used for operational efficiencies such as drafting or writing a regulatory submission when it does not affect patient safety, drug quality or the reliability of study results; its credibility framework applies when AI produces data that support regulatory decisions.
- Part 11, Electronic Records; Electronic Signatures, Scope and Application (US Food and Drug Administration, North America). States that FDA interprets the scope of Part 11 narrowly, covering electronic records and signatures required by predicate rules or submitted to FDA, and announces enforcement discretion for the validation, audit trail, record retention and record copying requirements, and for all Part 11 requirements on systems operational before 20 August 1997 (legacy systems). Predicate rules stay fully enforced regardless.
Controls to put in place
- Documented intended use, validation report and change control for the drafting platform
- Traceability from generated text to source tables and approved templates
- Automated numeric consistency checks with results kept in the document history
- Access controls and logging for unpublished trial data
- Quality control sampling of generated sections by experienced writers
Frequently asked questions
- How much faster does AI make clinical study reports?
- Merck reports that first drafts now take three to four days instead of two to three weeks, and that the time to a fully human reviewed first draft fell from an average of 180 to 80 hours. Novo Nordisk says writing times on clinical study reports fell by 90%. Review and approval still take time, so measure the whole cycle, not only drafting.
- What do regulators expect for AI drafted submissions?
- Responsibility stays with the company that uses the AI: the EMA reflection paper makes the clinical trial sponsor, marketing authorisation applicant or holder, or manufacturer responsible for ensuring that models and data pipelines are fit for purpose and meet GxP standards and EMA guidelines. For medicinal product information documents, the EMA asks for close human supervision and quality review so that all generated text is factually and syntactically correct. The FDA's January 2025 draft guidance on AI for regulatory decision making does not address AI used to draft a submission when it does not affect patient safety, drug quality or the reliability of study results.
- Does AI reduce errors or add them?
- Both are possible. Merck reports 50% fewer errors in drafts across data, messaging, citations, terminology and typography, but language models can also produce plausible wrong numbers, which is why automatic checks against the source tables matter.
How to cite this page
Blits.ai AI Use Case Library, "AI drafting of clinical study reports and regulatory documents", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/clinical-and-regulatory-document-drafting. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 27 September 2026: First published