AI use case

Governed text to SQL analytics assistant

An assistant that turns a business user's plain language question into a query against governed data, runs it under that user's own data permissions and returns the table or chart together with the SQL and the tables used, so routine ad hoc questions no longer queue for the data team.

By Len Debets · Last verified 27 September 2026 · 3 public deployments

About 300
Users served
Uber Technologies (organization claim).
USD 81,000 to USD 1.3 million
Indicative value per year
A bank with 1,500 regular data consumers and a central data and BI team. Worked example, see how it is calculated.

What problem does it solve?

Most data questions in a business are small and urgent: how many applications came in from this channel last week, what is the arrears rate by region, which branches missed target. Dashboards answer the questions someone anticipated; everything else becomes a ticket for an analyst who knows which of thousands of tables holds the answer and how to join them. The queue slows decisions and consumes analysts on work that is repetitive rather than analytical.

Generic text to SQL is not the answer on its own. On a real warehouse it picks the wrong table, misreads a column or applies the wrong business definition, and returns a confident, wrong number. In a bank or insurer the second risk is access: a query tool must never let a user see rows or columns their role does not permit. The job is therefore governed text to SQL: a curated semantic layer, the user's own permissions, and the SQL always visible.

How does it work?

  1. Understand the question. The assistant restates the question in business terms and asks for what is missing ("which period?", "gross or net?") instead of guessing.
  2. Find the right data. It retrieves from a curated semantic layer: certified tables, metric definitions, join paths and example queries for the business domain. Uber narrows the search with domain "workspaces"; LinkedIn had domain experts certify and describe key tables, and draws example queries from notebooks that users have certified.
  3. Write the query. A model generates SQL against those definitions only, and the query is validated (syntax, allowed tables, row limits) before it runs.
  4. Run it as the user. The query executes with the user's own credentials, so row and column level security in the data platform decides what comes back.
  5. Show the work. The answer comes with the SQL, the tables and the definitions used, plus a short explanation, so the user or an analyst can check it.
  6. Learn from corrections. Queries that analysts correct or certify become new examples in the semantic layer.
Audience
Employee facing
Autonomy
Assist
Adoption
Early adopters
Channels
Internal tools, Microsoft Teams

What is it worth?

Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.

Value benchmarks for Governed text to SQL analytics assistant
KPIMedianReported rangeData pointsClaimed by
Users servedNot pooled
about 300
11 organization

Value drivers: Speed and cycle time, Employee productivity, Lower cost to serve.

Indicative value

A bank with 1,500 regular data consumers and a central data and BI team

USD 81,000 to USD 1.3 million

Analyst time released from routine ad hoc requests per year

How this is calculated

Formula: requests * selfServeShare * analystHours * hourlyCost. The low scenario uses every low input, the high scenario every high input.

InputLowHighBasis
Ad hoc data requests to the data team per year requests, requests per year6,00012,000Editorial assumption, about four to eight requests per data consumer per year. Replace with your own ticket volume.
Share of requests the assistant answers without an analyst selfServeShare, fraction of requests0.150.35Editorial assumption, deliberately conservative because accuracy falls on complex multi table questions (see the Forrester figure cited on this page).
Analyst hours per ad hoc request analystHours, hours per request1.53Editorial assumption including clarification, query writing and checking.
Fully loaded analyst hour hourlyCost, USD per hour60100Editorial assumption. Replace with your own rate.

What it leaves out: Analyst time only. It leaves out the value of faster decisions for the business users, the cost of building and maintaining the semantic layer, and the cost of any wrong answers that are not caught, which is why the self serve share is kept low.

Who already uses it?

3 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.

LinkedIn

United States · Technology and software · 2024

ProductionGrade B

LinkedIn's SQL Bot, built into its DARWIN data science platform, finds the right tables, writes the query and fixes errors so employees can answer data questions themselves; datasets keep their own access control lists, and the bot only supplies group credentials the user is entitled to. Domain experts certified and described hundreds of key tables, which improved retrieval, and the example queries come from notebooks certified by users and those meeting recency and reliability heuristics. The interface shows the retrieved tables and the query, and hundreds of employees across business units use it. In a user survey about 95% rated its query accuracy "Passes" or above (about 40% "Very Good" or "Excellent"); LinkedIn publishes no measured accuracy rate.

No outcome disclosed.

Uber Technologies

United States · Technology and software · 2024

PilotGrade B

Uber's QueryGPT turns an English question into SQL against its data platform, which handles about 1.2 million interactive queries a month. It narrows the problem with curated "workspaces" of tables and sample queries per business domain (such as Mobility, Ads and Core Services), picks the relevant tables and columns with separate agents, and returns the generated SQL with an explanation so the user can check it. It was released to some Operations and Support teams first.

  • Users served: about 300, daily active users during the limited release
    "With our limited release to some teams in Operations and Support, we are averaging about 300 daily active users, with about 78% saying that the generated queries have reduced the amount of time they would’ve spent writing it from scratch."
    Claimed by: organization

Bayer

Germany · Pharma and life sciences · 2024

ProductionGrade C

Bayer uses Snowflake Cortex Analyst as the query generation service and a Streamlit chat interface to answer natural language questions over its enterprise data platform, alongside its existing dashboards. The first phase answered executive questions from sales vice presidents, such as the market share of a product last month, and it has since been extended to business unit analysts with row level data. No outcome figures are published.

No outcome disclosed.

How do you implement it?

A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.

Data you need

  • A curated set of certified tables or a semantic model with business definitions
  • Row and column level security defined in the data platform per role
  • A library of example questions with correct, reviewed SQL
  • Data classification, so sensitive columns can be excluded or masked

Systems to integrate

  • Data warehouse or lakehouse (SQL endpoint)
  • Semantic layer or data catalog
  • Identity provider for passing the user's identity to the data platform
  • BI tool or chat front end such as Microsoft Teams

Complexity: Medium

The model is the easy part. The work is the semantic layer (certified tables, metric definitions and example queries), enforcing each user's data permissions, and an evaluation set of real questions with known answers.

  1. 1

    Pick one domain with a clean model

    Start where definitions are settled and tables are few, for example sales pipeline or contact centre volumes. Write down the metric definitions before any model sees a question.

  2. 2

    Build the evaluation set first

    Collect 100 to 200 real questions from the request queue with correct SQL and results. Run every change of model, prompt or semantic layer against it and publish the pass rate.

  3. 3

    Enforce permissions in the data platform, not the prompt

    Run each query with the user's own identity so row and column security applies. Never rely on instructions to the model to hide data.

  4. 4

    Always show the SQL and the definitions

    Users and analysts must be able to see how a number was produced. Label answers built on uncertified tables clearly.

  5. 5

    Close the loop with analysts

    Route questions the assistant cannot answer confidently to the data team, and turn their corrected queries into new certified examples.

  6. 6

    Widen by domain, not by table count

    Add the next domain only when the first meets its accuracy target, and track accuracy per domain separately.

Guardrails

  • Queries run with the user's own credentials; row and column level security is enforced by the data platform
  • Read only access, an allow list of schemas and a row limit on every query
  • The generated SQL, tables and definitions are shown with every answer
  • The assistant asks for clarification or declines when confidence is low instead of guessing
  • Every question, query and result is logged for audit and evaluation
  • Sensitive columns (personal data, account numbers) excluded or masked unless the role needs them

KPIs to instrument

  • Execution accuracy on the evaluation set, per domain
  • Share of questions answered without analyst involvement, and share later flagged as wrong
  • Weekly active users and repeat usage
  • Time from question to answer compared with the request queue
  • Denied or blocked queries by reason (permission, schema, row limit)

Human in the loop

Analysts own the semantic layer and certify example queries. Any number that goes into a board pack, regulatory report or customer communication is checked by an analyst, not taken straight from the assistant. Users can flag a wrong answer, which goes to the data team for review.

Common failure modes

Confident but wrong numbers
The query runs and returns a plausible figure from the wrong table or definition. Show the SQL, restrict to certified tables, and measure accuracy on real questions.
Permission leaks through a service account
The assistant queries with a powerful technical account and bypasses row level security. Pass the user's identity to the data platform.
Definitions drift
The business changes a metric definition but the semantic layer does not. Give each definition an owner and a review date.
Adoption stalls on trust
One visible wrong answer and users return to the queue. Start with a narrow domain, label uncertainty and publish the accuracy number.

What are the risks and rules?

EU AI Act

Limited risk (transparency)

Article 50(1) requires providers to design AI systems that interact directly with people so that those people are informed they are dealing with AI, unless this is obvious from the context, as it usually is for an internal assistant. An analytics assistant that makes no decisions about people is not a prohibited practice under Article 5 and is not listed in Annex III. It would be high risk only if it were intended for an Annex III purpose, such as assessing the creditworthiness of natural persons (point 5(b)).

Guidance

Controls to put in place

  • Data access enforced by the data platform per user, with no shared privileged service account
  • Query and answer logs retained and reviewable by the data owner
  • Evaluation set pass rate recorded for every release of model, prompt or semantic layer
  • Owner and review date for every certified table and metric definition
  • Inventory entry for the assistant with an accountable owner

Frequently asked questions

How accurate is text to SQL on real company data?
It depends heavily on the question and the data model. Snowflake, citing anecdotal evidence in a Forrester report, gives a best case of 70% on simple single table queries and around 20% at worst on complex joins. LinkedIn reports that about 95% of surveyed users rated SQL Bot's query accuracy "Passes" or above, a user rating rather than a measured accuracy rate. A curated semantic layer and a narrow scope make the difference.
How do you stop users seeing data they are not entitled to?
Run every query with the user's own identity so the data platform's row and column level security applies, give the assistant read only access to an allow list of schemas, and mask sensitive columns. Instructions in a prompt are not an access control.
Does this replace the BI team?
No. It takes routine, repetitive questions off the queue. Analysts still own the definitions, certify example queries and handle complex or novel analysis, and anything that goes into a regulatory report or board pack is checked by an analyst.

How to cite this page

Blits.ai AI Use Case Library, "Governed text to SQL analytics assistant", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/governed-text-to-sql-analytics. Licensed under CC BY 4.0. Method: how we verify use cases.

Changelog
  • 27 September 2026: First published

Related use cases

Cross industryBanking

AI enterprise knowledge search for employees

An assistant that lets any employee ask a question in plain language and get a synthesized answer from the organization's own policies, procedures, product manuals and research, with citations to the source documents and only from documents the employee is allowed to see.

Deployments
4 public, best grade B
Autonomy
Assist
Cross industryRetail and ecommerce

AI for voice of the customer and feedback analysis

AI that reads every piece of free text customer feedback, such as survey verbatims, NPS comments, reviews, social posts, chat and call transcripts, and turns it into themes, sentiment, drivers and suggested actions that a named owner can act on, so the organization hears all of its customers instead of a sample.

Deployments
5 public, best grade B
Reported accuracy
84%
SBF Group, vendor claim
BankingInsurance

AI for regulatory report assembly

AI that assembles periodic and data driven regulatory filings and returns, such as prudential and statistical returns, threshold and transaction reports and disclosure packs, by pulling data into the regulator's schema, validating it, reconciling figures to source, explaining movements against prior periods and drafting commentary, before a named officer reviews and submits. Narratives for individual suspicious activity cases are a separate use case.

Deployments
2 public, best grade B
Autonomy
Copilot
Cross industryBanking

Generative AI copilot for internal audit

A copilot for internal auditors that drafts planning memos and document request lists from prior audits, summarises large evidence sets, builds risk and control matrices from policies and process documents, and drafts findings and reports, with every statement traceable to its evidence and a qualified auditor accountable for every conclusion.

Deployments
3 public, best grade C
Reported handling time reduction
55%
Banco Bradesco, vendor claim
BankingCross industry

AI cash flow forecasting for corporate treasury

Machine learning and conversational analytics, offered by some banks inside their cash management platforms, that categorise a company's cash flows, forecast positions across accounts and currencies, and answer treasurers' questions in plain language, so the treasury team decides on funding and idle balances with better information and less spreadsheet work.

Deployments
5 public, best grade B
Reported productivity gain
about 90%
JPMorgan Chase, organization claim