What problem does it solve?
Every new system, release and model needs data that looks like the real thing: customers with plausible histories, transactions with realistic patterns, claims, statements and conversations. The easy answer is to copy production into test environments, which puts personal and confidential data in places with weaker controls, in the hands of contractors and offshore teams, and in the training sets of models. Some organizations have sealed production off from development altogether, as Kin Insurance did, and data protection law makes such copies hard to justify.
The alternatives each have a cost. Hand written mock data is safe but thin, so tests pass that would fail on real data. Masking production data keeps realism but can still leave people identifiable, and preparing a usable copy by hand is slow: Patterson Dental needed 2.5 hours per dataset before it automated the work. Rare but important cases, such as fraud patterns or unusual products, are often missing entirely. Synthetic data generation aims to give teams realistic, referentially consistent data on demand, with a measured privacy risk.
How does it work?
- Profile the source. A generator learns the schema, relationships and statistical distributions of a production dataset inside the secure environment, or works from a schema and business rules when no production data may be used.
- Generate. It produces new records (customers, accounts, transactions, claims, documents or conversation transcripts) that follow the same patterns and keep keys consistent across tables, with the option to add rare cases such as fraud or edge conditions on purpose.
- Measure privacy, utility and fidelity. Automated checks confirm that no generated record copies a real one, test for reidentification risk and compare the synthetic statistics with the real ones.
- Approve and publish. A data owner or privacy officer signs off each new dataset configuration; routine refreshes then run on a schedule or through an API into test, analytics or sandbox environments.
- Use and monitor. Teams test, train or demo on the synthetic data, and results that matter (a model going live, a performance benchmark) are confirmed on real data under proper controls.
- Audience
- Back office
- Autonomy
- Supervised agent
- Adoption
- Early adopters
- Channels
- Internal tools, API and system to system
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
| KPI | Median | Reported range | Data points | Claimed by |
|---|---|---|---|---|
| Users served | Not pooled | 28 to 2000 | 2 | 1 organization, 1 vendor |
| Cycle time | Not pooled | 1 days | 1 | 1 vendor |
| Cycle time reduction | Too few to pool | 75% | 1 | 1 vendor |
Value drivers: Speed and cycle time, Risk and loss reduction, Compliance quality, Employee productivity.
Indicative value
A bank with 30 delivery teams that request test data for their environments
USD 24,000 to USD 360,000
Engineering time released from test data preparation per year
How this is calculated
Formula: requests * hoursPerRequest * timeSaved * hourlyCost. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Test data requests per year requests, requests per year | 200 | 400 | Editorial assumption for about 30 teams refreshing environments every few weeks. Replace with your own ticket volume. |
| Engineering hours to prepare one masked or hand built dataset hoursPerRequest, hours per request | 4 | 12 | Editorial assumption covering extraction, masking, fixing broken references and checks. Patterson Dental reports 2.5 hours per dataset before automation. |
| Share of preparation time removed timeSaved, fraction of hours | 0.5 | 0.75 | Capped at the benchmark on this page (Patterson Dental reports a 75% reduction in test data generation time). |
| Fully loaded cost of an engineer hour hourlyCost, USD per hour | 60 | 100 | Editorial assumption. Replace with your own blended rate. |
What it leaves out: Counts only the preparation effort. It leaves out licence and compute cost, the risk reduction of removing personal data from lower environments, faster releases and fewer defects found late.
Who already uses it?
7 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
Financial Conduct Authority
United Kingdom · Government and public sector · 2023
The UK Financial Conduct Authority has built synthetic datasets since a 2020 DataSprint, gave participants in two Digital Sandbox pilots access to synthetic data, opened a permanent Digital Sandbox in August 2023 and released an authorised push payment fraud synthetic dataset in September 2023, so firms can build and test solutions without real customer data. Its Synthetic Data Expert Group published a March 2024 report with use cases on system testing, model validation and data sharing, including the trade offs between privacy, utility and fidelity.
- Users served: 28, First Digital Sandbox pilot, organisations given access to synthetic data
"The first pilot, involving 28 organisations, underscored the value of synthetic data, emphasising the need for more referentially linked datasets and finer granularity."
Claimed by: organization
Internal Revenue Service
United States · Government and public sector · 2022
The IRS runs an AI based synthetic data generator that builds a large population of synthetic people, households and businesses, ages them over time and correlates them with the socioeconomic patterns of US taxpayers. It outputs synthetic individual and business tax returns for several tax years, plus the reference files that seed test systems, so tax processing systems can be tested, including simulated fraud cases, without exposing taxpayer information. Automated checks run on every schema version to catch anomalies in the generated returns. No outcome figures are published.
No outcome disclosed.
JPMorgan Chase
United States · Banking · 2020
J.P. Morgan AI Research develops generators for realistic synthetic datasets in financial services and makes them available to researchers on request. The published sets cover anti money laundering customer traces, retail customer journeys, payments data for fraud detection, market order books, synthetic documents for layout recognition and simulated equity market data. Its documented method computes metrics on the real data, builds and optionally calibrates a generator (statistical or agent based simulation) and then compares the metrics of the synthetic and the real data. No business outcome figures are published.
No outcome disclosed.
Boomi
United States · Technology and software · 2026
Boomi, an integration platform company, gave employees in its citizen developer programme a synthetic copy of its Snowflake data warehouse, generated with Tonic Structural, so they can build and test AI agents and workflows without touching customer data. When a prototype is approved, the enterprise AI team points it at the production warehouse. Refreshes run automatically each weekend, and five agents were being rolled out when the story was published.
- Users served: about 2000, Employees with access to the synthetic environment
"With de-identified data from Tonic Structural, Boomi has opened its citizen developer program to approximately 2,000 people."
Claimed by: vendor
Patterson Dental
United States · Healthcare · 2025
Patterson Dental, a division of Patterson Companies, uses Tonic Structural to generate deidentified, production like data for performance and functional testing of its dental practice platforms, so protected health information stays out of developer workflows. The vendor reports that test data preparation fell from 2.5 hours to 35 minutes per dataset and that performance testing grew from one practice to between 15 and 25 practices a day, with up to seven development teams using the data.
- Cycle time reduction: 75%, Test data generation time per dataset
"With Tonic Structural, the company achieved immediate and measurable improvements in their testing workflows, reducing test data generation time by 75% and cutting it down from 2.5 hours to just 35 minutes."
Claimed by: vendor
Merkur Versicherung AG
Austria · Insurance · 2023
The Merkur Innovation Lab, the innovation arm of the Austrian insurer Merkur Versicherung, runs an automated pipeline that extracts its active customer data (about 600,000 rows and 55 columns) from an Oracle database, has MOSTLY AI generate a synthetic version through a REST call and writes the result to a PostgreSQL database with Apache Airflow every day. The synthetic health data feeds internal analysis dashboards and is used to explore data sharing with third parties. The vendor reports that time to data fell from one month to one day.
- Cycle time: 1 days, Time from data request to usable synthetic data
"The end-to-end automated workflow has cut Merkur’s time-to-data from 1-month, to 1-day."
Claimed by: vendor
Kin Insurance
United States · Insurance · 2022
Kin Insurance, a digital home insurer, sealed its production database off from development and now gives engineers and QA only subsetted, masked copies generated with Tonic, including differential privacy to protect people who stand out in the data. Developers use the subsets to fix bugs and build features, and QA uses the same data to check that releases behave as they did in the sandbox. The vendor reports faster data access and fewer security concerns but no quantified outcome beyond a database that can be pulled down in an hour or less.
No outcome disclosed.
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- Access to the source data inside a controlled environment, or a complete schema with business rules
- A data classification that marks personal, special category and confidential fields
- Agreed privacy, utility and fidelity thresholds per use (testing, model training, external sharing)
- Examples of the rare cases and edge conditions tests must cover
Systems to integrate
- Source databases and data warehouse for profiling
- Target test, sandbox and analytics environments
- CI pipelines and environment provisioning, so data refreshes run with deployments
- Data catalogue for lineage and approval records
Complexity: Medium
Generating a single table is easy. The work is in multi table data with referential integrity, business rules that must hold (a closed account has no new transactions), privacy measurement a privacy officer will accept, and pipelines that keep the synthetic data in step with schema changes.
- 1
Start where production copies hurt most
List the environments and teams that still receive production or lightly masked data, and pick one with a clear pain, such as a performance test that needs volume or a supplier that should never see real customers.
- 2
Decide the privacy standard before generating anything
Agree with the privacy officer which checks a dataset must pass (no copied records, distance to closest real record, attribute inference tests) and who signs off. Treat the generator itself as processing of personal data, because it learns from it.
- 3
Model the relationships, not just the columns
Map keys and business rules across tables and add them as constraints, so tests exercise real journeys. A customer without accounts or a claim without a policy produces false failures.
- 4
Add the cases production lacks
Deliberately oversample fraud patterns, edge values, long names, rare products and multiple languages. This is where synthetic data beats a production copy.
- 5
Automate refresh and measure use
Run generation from the pipeline on a schedule or per environment, and track requests, time to data and the defects found with synthetic data versus escaped defects.
- 6
Keep real data for the final proof
For model training and validation, compare models trained on synthetic, mixed and real data before relying on synthetic data alone, and confirm go live decisions on real data under controls.
Guardrails
- Generation runs inside the secured data environment; only the approved synthetic output leaves it
- Every dataset passes automated privacy checks for copied records and reidentification risk before release
- Stricter thresholds, or rule based generation without production data, for datasets shared outside the organization
- Labels on every synthetic dataset so it is never mistaken for, or merged with, real data
- Model decisions that affect customers are validated on real data, not on synthetic data alone
KPIs to instrument
- Time from data request to usable test data
- Share of test environments that hold no production personal data
- Privacy test results per release (copied records, closest record distance)
- Fidelity scores against the source statistics per dataset
- Defects found in test versus defects that escape to production
Human in the loop
A data owner and the privacy officer approve each new dataset configuration and each external release, and review the privacy report. Test and data science leads confirm that the data is fit for purpose, and a person decides when a result on synthetic data is strong enough to act on.
Common failure modes
- Synthetic data that leaks real people
- Overfitted generators can reproduce real records or rare outliers. Measure copies and closest records on every release and use differential privacy or stricter settings for sensitive data.
- Realistic columns, broken journeys
- Distributions match but relationships and business rules do not, so tests fail for the wrong reasons or pass for the wrong reasons. Encode constraints and test the data itself.
- Models trained on synthetic data alone
- Accuracy can drop sharply when real data is removed entirely. Keep a share of real data or validate on real data before deployment.
- A second shadow copy of production
- Synthetic datasets proliferate without owners or labels. Register them in the catalogue with purpose, lineage and expiry.
What are the risks and rules?
EU AI Act
Depends on design
A generator of synthetic tabular test data is not listed in Annex III and does not interact with people, so it is minimal risk with only the AI literacy duty of Article 4. When the system generates synthetic text, images, audio or video, such as documents or conversation transcripts, Article 50(2) requires its provider to mark the output in a machine readable format as artificially generated. When synthetic data is used to train, validate or test a high risk system, such as credit scoring, it falls under that system's data governance duties in Article 10.
Guidance
- What PETs are there? Synthetic data (UK Information Commissioner's Office, Europe). Guidance on synthetic data as a privacy enhancing technology. It notes that generating it from real data may involve processing personal information, that closer resemblance to real data raises the chance of revealing someone's information, and that biases carry through. The page says it is under review after the Data (Use and Access) Act.
- Using Synthetic Data in Financial Services (Financial Conduct Authority, Synthetic Data Expert Group, Europe). March 2024 report with use cases on system testing, model validation and data sharing, and practical advice on evaluating privacy, utility and fidelity.
Controls to put in place
- Data protection impact assessment for the generator and its training data
- Documented privacy, utility and fidelity thresholds with sign off per dataset
- Access control and logging on the environment where the generator sees real data
- Catalogue entries with lineage, purpose and expiry for every synthetic dataset
- Contract terms that forbid suppliers from attempting reidentification
Frequently asked questions
- Is synthetic data still personal data under GDPR?
- It can be. The FCA's Synthetic Data Expert Group notes that synthetic data tends to have a lower privacy risk than real data but cannot guarantee privacy, so a risk assessment per implementation is needed. Generation from real records is itself processing of personal data, and outputs should be tested for copied records and reidentification before release.
- Can synthetic data replace real data for training models?
- Not always. In a proof of concept reported by the FCA's Synthetic Data Expert Group, a model trained on half synthetic and half real data scored 2.5% below the accuracy of the real data benchmark, while one trained on synthetic data only scored 32% below it. For software testing the bar is lower, because privacy and utility matter more than exact fidelity.
- How is synthetic data different from masked production data?
- Masking keeps real records and replaces identifying fields, so relationships stay intact but people can sometimes still be identified from what remains. Synthetic data creates new records from learned patterns or rules. Many teams combine the two, as Kin Insurance did with subsetting, masking and differential privacy.
- How much faster does test data get?
- The published cases report large gains in time to data. Patterson Dental reports test data generation falling from 2.5 hours to 35 minutes per dataset, and Merkur Versicherung reports time to data falling from one month to one day with a daily automated pipeline. Both are vendor case studies.
How to cite this page
Blits.ai AI Use Case Library, "AI for synthetic test data generation", last verified 26 September 2026, https://www.blits.ai/ai-use-cases/synthetic-test-data-generation. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 27 September 2026: First published