AI use case

AI spend classification and spend analytics for procurement

AI that reads purchase orders, invoices, card transactions and contracts and assigns each line of spend to a category in the organization's taxonomy, and to the right supplier, so that procurement can see what is bought, from whom and where to consolidate or negotiate.

By Len Debets · Last verified 27 September 2026 · 4 public deployments

USD 66,667 to USD 746,667
Indicative value per year
An organization with 2 million purchase lines a year across several ERP and card systems. Worked example, see how it is calculated.

What problem does it solve?

Procurement can only manage what it can see, and raw spend data does not show it on its own. Purchases can arrive from several ERP systems, purchasing cards and expense tools, with free text descriptions ("gloves blk L 100"), inconsistent supplier names and general ledger codes that describe the budget, not the thing bought. Building a reliable view of spend by category means cleaning and classifying large volumes of lines. Where that is done by hand in a periodic exercise, the view can be out of date by the time it is finished.

The consequences are practical. Category managers cannot tell how much the organization spends on a category across units, so they cannot consolidate demand or negotiate on volume, contract compliance and maverick buying go unmeasured, and savings claims are hard to prove. The US federal government manages its buying through a government wide category management taxonomy. In the 2025 inventory, USDA says that to plan for the upcoming fire season all of the previous year's incident related purchases are categorized by hand, a process it describes as time consuming; it has piloted a machine learning classifier for this since December 2024.

How does it work?

  1. Gather and clean. Purchase order lines, invoice lines, card transactions and contract records are extracted from each source system, and supplier names are normalized and matched to one supplier record.
  2. Classify each line. A model trained on lines already labeled by buyers, or a language model given the taxonomy with definitions and examples, assigns each line to a category and a subcategory and returns a confidence score.
  3. Read the contracts. For contract spend, the AI reads the contract document and summarizes what goods or services it covers, which gives category managers a view of what each contract buys. The IRS uses generative AI to summarize its contracts in support of category management.
  4. Review the uncertain. Low confidence lines and high value lines go to a category analyst, whose corrections are fed back as training examples.
  5. Analyze. The classified spend feeds dashboards by category, supplier, unit and period, showing consolidation opportunities, contract coverage and trends.
Audience
Back office
Autonomy
Supervised agent
Adoption
Early adopters
Channels
Internal tools, API and system to system

What is it worth?

Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.

No public deployment has disclosed a measurable outcome yet.

Value drivers: Employee productivity, Lower cost to serve, Speed and cycle time, Compliance quality.

Indicative value

An organization with 2 million purchase lines a year across several ERP and card systems

USD 66,667 to USD 746,667

Classification labor avoided per year

How this is calculated

Formula: lines * manualShare * automatedShare * minutesPerLine / 60 * hourlyCost. The low scenario uses every low input, the high scenario every high input.

InputLowHighBasis
Purchase, invoice and card lines per year lines, lines per year2,000,0002,000,000The reference organization. Replace with your own volumes.
Share of lines that need manual classification today manualShare, fraction of lines0.20.4Editorial assumption, replace with the share your rules or suppliers do not classify.
Share of that manual work the AI takes over automatedShare, fraction of manual lines0.50.8Editorial assumption; analysts still review low confidence and high value lines.
Analyst minutes per manually classified line minutesPerLine, minutes per line0.51Editorial assumption for a trained analyst working in batches.
Fully loaded analyst hour hourlyCost, USD per hour4070Editorial assumption.

What it leaves out: Counts only the labor of classifying spend. It leaves out the larger but harder to attribute value of better category decisions (consolidation, negotiation, contract compliance), the cost of the platform and model, and the one time work of building the taxonomy and training data.

Who already uses it?

4 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.

Internal Revenue Service

United States · Government and public sector · 2025

ProductionGrade B

Since April 2025 the IRS runs an analytics hub for contract documents and contract spending data in which generative AI reads contract PDFs and writes a summary of the products or services bought under each agency contract into a table. The aim is better category management and a better view of what the agency buys. Outputs go to contracting officials for review, and the inventory classifies the use as not high impact because it is not the principal basis for significant decisions. No outcome figures are published.

No outcome disclosed.

U.S. General Services Administration

United States · Government and public sector · 2025

ProductionGrade B

The General Services Administration runs an Acquisition Analytics capability that uses natural language processing to classify transactions within the Government-wide Category Management Taxonomy. The classification lets category managers see how obligations are distributed across categories and decide where to aggregate spend across agencies. The inventory lists it as deployed; no operational date or outcome figures are published.

No outcome disclosed.

Veterans Health Administration

United States · Healthcare · 2025

ProductionGrade B

The Veterans Health Administration uses generative AI to enrich purchase order data from its Integrated Funds Control, Accounting, and Procurement system (IFCAP) by assigning each line item to predefined spend categories that are more useful for analysis. The result feeds an executive dashboard with total spend, line items, categories and trends that can be filtered by region, budget object code, fund control point and vendor, so that network and national leaders can find areas to drive spending efficiencies. The inventory lists it as deployed and not high impact; no outcome figures are published.

No outcome disclosed.

U.S. Department of Agriculture

United States · Government and public sector · 2024

PilotGrade B

To plan for each fire season, USDA staff in Natural Resources and Environment categorized all of the previous year's incident related purchases by hand, which took a long time and could miss subtle purchasing patterns. An in house classical machine learning model, trained on 2023 item descriptions (such as protective work gloves) and validated on human labeled item categories from the same year, assigns purchases to common purchase categories identified by procurement staff in the pilot. The goal is to find purchasing patterns and cost savings opportunities, which the agency says could lead to quicker resource allocation to incident locations. The inventory lists it as a pilot since December 2024; no outcome figures are published.

No outcome disclosed.

How do you implement it?

A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.

Data you need

  • Purchase order, invoice and card line data with free text descriptions
  • A spend taxonomy with definitions and examples per category
  • A supplier master, or at least a supplier normalization table
  • A labeled sample of lines classified by experienced buyers

Systems to integrate

  • ERP and purchasing systems
  • Purchasing card and expense platforms
  • Contract repository
  • Analytics or business intelligence tools for the spend dashboards

Complexity: Medium

The model is the easy part. The work is in extracting and joining data from several source systems, agreeing a taxonomy with clear definitions, cleaning supplier names and building a labeled sample that reflects the organization's own language.

  1. 1

    Fix the taxonomy first

    Choose the taxonomy (your own, UNSPSC or a public sector category structure) and write a definition and examples for each category. A model cannot be more consistent than the definitions it is given.

  2. 2

    Build a labeled benchmark

    Have buyers classify a random sample of a few thousand lines, stratified by source and value, and keep it aside to measure accuracy on every change.

  3. 3

    Classify with confidence thresholds

    Let the model classify all lines, accept high confidence lines automatically and route low confidence and high value lines to analysts. Tune the threshold on the benchmark.

  4. 4

    Close the loop

    Feed analyst corrections back into training data or prompt examples, and rerun the benchmark before each release.

  5. 5

    Put the result to work

    Connect classified spend to category plans and supplier negotiations, and report savings by category, so the classification earns its keep.

Guardrails

  • Every line keeps its original description and source, so a classification can be traced and corrected
  • Confidence score on every line, with low confidence and high value lines reviewed by a person
  • Accuracy measured on a fixed benchmark before each change to the model or taxonomy
  • Taxonomy changes versioned, with reclassification of history when definitions change

KPIs to instrument

  • Share of spend value and of lines classified automatically
  • Accuracy on the labeled benchmark, by category and source
  • Analyst hours spent on classification per period
  • Time from period end to an updated spend view
  • Savings identified and realized per category

Human in the loop

Category analysts own the taxonomy, review low confidence and high value lines and sample accepted lines each period. Classifications inform decisions but do not trigger purchases or payments.

Common failure modes

High accuracy by line, low accuracy by value
The model classifies many small lines well and a few large ones badly. Measure accuracy weighted by spend value and review large lines by hand.
Taxonomy drift
Categories are renamed or split but old spend is not reclassified, so trends break. Version the taxonomy and reclassify history.
Garbage descriptions
Lines with empty or generic descriptions cannot be classified reliably by any model. Use the supplier, contract and ledger code as extra signals and fix the capture at the source.
Dashboards nobody uses
Spend is classified but category plans do not change. Tie the output to specific sourcing decisions from the start.

What are the risks and rules?

EU AI Act

Minimal risk

Classifying the organization's own purchase lines into categories is not listed in Annex III and is used internally by procurement staff, so no specific obligations apply beyond AI literacy. The data can still contain personal data, for example in purchasing card and expense lines, which brings GDPR duties. Using the classified card and expense lines to monitor or evaluate individual employees would move the system towards Annex III point 4 (employment and worker management) and a high risk assessment.

Guidance

  • M-19-13: Category Management: Making Smarter Use of Common Contract Solutions and Practices (Office of Management and Budget, North America). The OMB memorandum that implements category management government wide, defines the role of the Category Management Leadership Council and asks agencies to use spending data and data analytics tools to make data driven buying decisions. It is not specific to AI.
  • Category management (U.S. General Services Administration, North America). Describes category management as identifying categories of spend and using data to consolidate contracts, manage suppliers and demand, and reduce total cost of ownership.

Controls to put in place

  • Named owner for the taxonomy and for classification quality
  • Benchmark results and model versions kept with each release
  • Masking or exclusion of personal data in expense and card lines before classification
  • Periodic sample audit of automatically accepted lines

Frequently asked questions

Who uses AI for spend classification?
In the US federal government, GSA classifies transactions into the government wide category management taxonomy, the Veterans Health Administration uses generative AI to categorize purchase order lines for an executive spend dashboard, and the IRS has generative AI summarize what each contract buys. None of the inventory entries reports accuracy or savings figures, so measure your own.
Should we use a language model or a trained classifier?
Both work. A classifier trained on your own labeled lines is cheap and fast at volume; a language model given the taxonomy and examples needs less training data and copes better with new categories. One option is to use the classifier for routine lines and a language model for uncertain ones; whichever you choose, compare both on the same benchmark.
How accurate does it need to be?
Accurate enough by value for the decisions it supports. Measure accuracy weighted by spend, review high value lines by hand, and report the share of spend classified with high confidence rather than a single accuracy number.

How to cite this page

Blits.ai AI Use Case Library, "AI spend classification and spend analytics for procurement", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/procurement-spend-classification. Licensed under CC BY 4.0. Method: how we verify use cases.

Changelog
  • 27 September 2026: First published

Related use cases

Cross industryBanking

AI for supplier invoice processing in accounts payable

AI that captures supplier invoices from any format, extracts header and line data, matches them to purchase orders and goods receipts, proposes tax and cost centre coding, flags duplicates and suspected fraud, and routes them for approval and posting, leaving only exceptions to accounts payable staff.

Deployments
4 public, best grade B
Reported productivity gain
80%
Kingfisher, organization claim
Cross industryBanking

AI assistant for procurement and supplier contract review

An assistant for procurement and vendor management that reads supplier contracts and proposals, extracts the key terms, flags deviations from the organization's standard positions, drafts requests for proposal and evaluation matrices, and prepares negotiation positions, with a procurement or legal owner approving every conclusion.

Deployments
5 public, best grade B
Autonomy
Copilot
Cross industryBanking

AI for third party and vendor risk due diligence

AI that reviews a vendor's security questionnaires, SOC and assurance reports, contracts and model documentation against the organization's control requirements, researches the vendor's ownership, sanctions, financial health and adverse media, drafts the risk assessment for a human to approve and keeps the register of material service providers current with ongoing monitoring.

Deployments
4 public, best grade B
Autonomy
Copilot
BankingPayments and cards

AI for ledger and payment reconciliation

AI that matches entries across nostro and vostro statements, card and scheme settlement files, the general ledger and suspense accounts, proposes matches and clearing journals, and routes only the genuine breaks to an operator with a plain language explanation.

Deployments
4 public, best grade B
Autonomy
Supervised agent