What problem does it solve?
Most data questions in a business are small and urgent: how many applications came in from this channel last week, what is the arrears rate by region, which branches missed target. Dashboards answer the questions someone anticipated; everything else becomes a ticket for an analyst who knows which of thousands of tables holds the answer and how to join them. The queue slows decisions and consumes analysts on work that is repetitive rather than analytical.
Generic text to SQL is not the answer on its own. On a real warehouse it picks the wrong table, misreads a column or applies the wrong business definition, and returns a confident, wrong number. In a bank or insurer the second risk is access: a query tool must never let a user see rows or columns their role does not permit. The job is therefore governed text to SQL: a curated semantic layer, the user's own permissions, and the SQL always visible.
- Snowflake, citing a Forrester report, says anecdotal evidence shows a best case rate of 70% for generating accurate, executable code on simple single table queries and around 20% at worst for queries with multiple tables or complex joins.Cortex Analyst: Paving the Way to Self-Service Analytics with AI (2024)
How does it work?
- Understand the question. The assistant restates the question in business terms and asks for what is missing ("which period?", "gross or net?") instead of guessing.
- Find the right data. It retrieves from a curated semantic layer: certified tables, metric definitions, join paths and example queries for the business domain. Uber narrows the search with domain "workspaces"; LinkedIn had domain experts certify and describe key tables, and draws example queries from notebooks that users have certified.
- Write the query. A model generates SQL against those definitions only, and the query is validated (syntax, allowed tables, row limits) before it runs.
- Run it as the user. The query executes with the user's own credentials, so row and column level security in the data platform decides what comes back.
- Show the work. The answer comes with the SQL, the tables and the definitions used, plus a short explanation, so the user or an analyst can check it.
- Learn from corrections. Queries that analysts correct or certify become new examples in the semantic layer.
- Audience
- Employee facing
- Autonomy
- Assist
- Adoption
- Early adopters
- Channels
- Internal tools, Microsoft Teams
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
| KPI | Median | Reported range | Data points | Claimed by |
|---|---|---|---|---|
| Users served | Not pooled | about 300 | 1 | 1 organization |
Value drivers: Speed and cycle time, Employee productivity, Lower cost to serve.
Indicative value
A bank with 1,500 regular data consumers and a central data and BI team
USD 81,000 to USD 1.3 million
Analyst time released from routine ad hoc requests per year
How this is calculated
Formula: requests * selfServeShare * analystHours * hourlyCost. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Ad hoc data requests to the data team per year requests, requests per year | 6,000 | 12,000 | Editorial assumption, about four to eight requests per data consumer per year. Replace with your own ticket volume. |
| Share of requests the assistant answers without an analyst selfServeShare, fraction of requests | 0.15 | 0.35 | Editorial assumption, deliberately conservative because accuracy falls on complex multi table questions (see the Forrester figure cited on this page). |
| Analyst hours per ad hoc request analystHours, hours per request | 1.5 | 3 | Editorial assumption including clarification, query writing and checking. |
| Fully loaded analyst hour hourlyCost, USD per hour | 60 | 100 | Editorial assumption. Replace with your own rate. |
What it leaves out: Analyst time only. It leaves out the value of faster decisions for the business users, the cost of building and maintaining the semantic layer, and the cost of any wrong answers that are not caught, which is why the self serve share is kept low.
Who already uses it?
3 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
United States · Technology and software · 2024
LinkedIn's SQL Bot, built into its DARWIN data science platform, finds the right tables, writes the query and fixes errors so employees can answer data questions themselves; datasets keep their own access control lists, and the bot only supplies group credentials the user is entitled to. Domain experts certified and described hundreds of key tables, which improved retrieval, and the example queries come from notebooks certified by users and those meeting recency and reliability heuristics. The interface shows the retrieved tables and the query, and hundreds of employees across business units use it. In a user survey about 95% rated its query accuracy "Passes" or above (about 40% "Very Good" or "Excellent"); LinkedIn publishes no measured accuracy rate.
No outcome disclosed.
Uber Technologies
United States · Technology and software · 2024
Uber's QueryGPT turns an English question into SQL against its data platform, which handles about 1.2 million interactive queries a month. It narrows the problem with curated "workspaces" of tables and sample queries per business domain (such as Mobility, Ads and Core Services), picks the relevant tables and columns with separate agents, and returns the generated SQL with an explanation so the user can check it. It was released to some Operations and Support teams first.
- Users served: about 300, daily active users during the limited release
"With our limited release to some teams in Operations and Support, we are averaging about 300 daily active users, with about 78% saying that the generated queries have reduced the amount of time they would’ve spent writing it from scratch."
Claimed by: organization
Bayer
Germany · Pharma and life sciences · 2024
Bayer uses Snowflake Cortex Analyst as the query generation service and a Streamlit chat interface to answer natural language questions over its enterprise data platform, alongside its existing dashboards. The first phase answered executive questions from sales vice presidents, such as the market share of a product last month, and it has since been extended to business unit analysts with row level data. No outcome figures are published.
No outcome disclosed.
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- A curated set of certified tables or a semantic model with business definitions
- Row and column level security defined in the data platform per role
- A library of example questions with correct, reviewed SQL
- Data classification, so sensitive columns can be excluded or masked
Systems to integrate
- Data warehouse or lakehouse (SQL endpoint)
- Semantic layer or data catalog
- Identity provider for passing the user's identity to the data platform
- BI tool or chat front end such as Microsoft Teams
Complexity: Medium
The model is the easy part. The work is the semantic layer (certified tables, metric definitions and example queries), enforcing each user's data permissions, and an evaluation set of real questions with known answers.
- 1
Pick one domain with a clean model
Start where definitions are settled and tables are few, for example sales pipeline or contact centre volumes. Write down the metric definitions before any model sees a question.
- 2
Build the evaluation set first
Collect 100 to 200 real questions from the request queue with correct SQL and results. Run every change of model, prompt or semantic layer against it and publish the pass rate.
- 3
Enforce permissions in the data platform, not the prompt
Run each query with the user's own identity so row and column security applies. Never rely on instructions to the model to hide data.
- 4
Always show the SQL and the definitions
Users and analysts must be able to see how a number was produced. Label answers built on uncertified tables clearly.
- 5
Close the loop with analysts
Route questions the assistant cannot answer confidently to the data team, and turn their corrected queries into new certified examples.
- 6
Widen by domain, not by table count
Add the next domain only when the first meets its accuracy target, and track accuracy per domain separately.
Guardrails
- Queries run with the user's own credentials; row and column level security is enforced by the data platform
- Read only access, an allow list of schemas and a row limit on every query
- The generated SQL, tables and definitions are shown with every answer
- The assistant asks for clarification or declines when confidence is low instead of guessing
- Every question, query and result is logged for audit and evaluation
- Sensitive columns (personal data, account numbers) excluded or masked unless the role needs them
KPIs to instrument
- Execution accuracy on the evaluation set, per domain
- Share of questions answered without analyst involvement, and share later flagged as wrong
- Weekly active users and repeat usage
- Time from question to answer compared with the request queue
- Denied or blocked queries by reason (permission, schema, row limit)
Human in the loop
Analysts own the semantic layer and certify example queries. Any number that goes into a board pack, regulatory report or customer communication is checked by an analyst, not taken straight from the assistant. Users can flag a wrong answer, which goes to the data team for review.
Common failure modes
- Confident but wrong numbers
- The query runs and returns a plausible figure from the wrong table or definition. Show the SQL, restrict to certified tables, and measure accuracy on real questions.
- Permission leaks through a service account
- The assistant queries with a powerful technical account and bypasses row level security. Pass the user's identity to the data platform.
- Definitions drift
- The business changes a metric definition but the semantic layer does not. Give each definition an owner and a review date.
- Adoption stalls on trust
- One visible wrong answer and users return to the queue. Start with a narrow domain, label uncertainty and publish the accuracy number.
What are the risks and rules?
EU AI Act
Limited risk (transparency)
Article 50(1) requires providers to design AI systems that interact directly with people so that those people are informed they are dealing with AI, unless this is obvious from the context, as it usually is for an internal assistant. An analytics assistant that makes no decisions about people is not a prohibited practice under Article 5 and is not listed in Annex III. It would be high risk only if it were intended for an Annex III purpose, such as assessing the creditworthiness of natural persons (point 5(b)).
Rules that apply
Guidance
- General Data Protection Regulation, Article 25 data protection by design and by default (European Union, Europe). Access to personal data must be limited to what each purpose needs, which argues for running every generated query under the user's own permissions and masking personal data by default.
- Article 50, transparency obligations for providers and deployers of certain AI systems (European Union, Europe). Providers must design AI systems that interact directly with people so that those people are informed they are interacting with AI, unless this is obvious from the context. Applies from 2 August 2026.
Controls to put in place
- Data access enforced by the data platform per user, with no shared privileged service account
- Query and answer logs retained and reviewable by the data owner
- Evaluation set pass rate recorded for every release of model, prompt or semantic layer
- Owner and review date for every certified table and metric definition
- Inventory entry for the assistant with an accountable owner
Frequently asked questions
- How accurate is text to SQL on real company data?
- It depends heavily on the question and the data model. Snowflake, citing anecdotal evidence in a Forrester report, gives a best case of 70% on simple single table queries and around 20% at worst on complex joins. LinkedIn reports that about 95% of surveyed users rated SQL Bot's query accuracy "Passes" or above, a user rating rather than a measured accuracy rate. A curated semantic layer and a narrow scope make the difference.
- How do you stop users seeing data they are not entitled to?
- Run every query with the user's own identity so the data platform's row and column level security applies, give the assistant read only access to an allow list of schemas, and mask sensitive columns. Instructions in a prompt are not an access control.
- Does this replace the BI team?
- No. It takes routine, repetitive questions off the queue. Analysts still own the definitions, certify example queries and handle complex or novel analysis, and anything that goes into a regulatory report or board pack is checked by an analyst.
How to cite this page
Blits.ai AI Use Case Library, "Governed text to SQL analytics assistant", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/governed-text-to-sql-analytics. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 27 September 2026: First published