What problem does it solve?
Cloud bills grow with usage nobody is watching in real time: instances sized for a launch that never get downsized, storage nobody deleted, reserved capacity that no longer matches the workload. A platform team can see the bill at the end of the month, but by then the waste has already been paid for, and working out which of thousands of resources across dozens of accounts are the problem is its own project.
Committed discounts make this harder, not easier. AWS Savings Plans commit a consistent amount of usage for a one or three year term, so many teams either skip them and pay full on demand price, or commit too much and pay for capacity they stop using when a workload changes. AWS reports that as of May 2026, across a sample of more than 71,000 anonymized, opted in AWS customers, the median Cost Efficiency score is 83 while the mean is 79, a gap it attributes to a long tail of accounts that are less optimized than the typical one.
- Savings Plans offer low prices on Amazon EC2, AWS Lambda and AWS Fargate usage in exchange for a commitment to a consistent amount of usage for a 1 or 3 year term.AWS Savings Plans compute pricing (2026)
- AWS reports that as of May 2026, across more than 71,000 anonymized, opted in AWS customers, the median Cost Efficiency score is 83 while the mean is 79, a gap it attributes to a long tail of less optimized accounts.The AWS State of Cost Efficiency Report (2026)
How does it work?
- Read usage and billing continuously. The agent ingests cost and usage data across accounts, regions and workload types, instead of a monthly export someone reviews by hand.
- Learn the normal pattern. Machine learning models the expected usage per resource and flags spend that deviates from it, including idle, oversized and orphaned resources.
- Model the commitment portfolio. For reserved capacity, the agent forecasts usage and recommends or automatically manages the mix of on demand, reserved and spot capacity that covers it, adjusting as usage changes instead of locking in a single upfront bet.
- Act within limits. Low risk changes that are reversible, such as stopping an idle resource or buying a commitment that carries a buyback or exchange option, execute automatically under a policy; changes that touch running production workloads go to an engineer as a proposal with the expected saving and the risk.
- Report the outcome. Savings, coverage and utilization are tracked back to the team and workload that owns the cost, so the finance and engineering view of spend stays the same one.
- Audience
- Employee facing
- Autonomy
- Supervised agent
- Adoption
- Early adopters
- Channels
- Internal tools, API and system to system
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
| KPI | Median | Reported range | Data points | Claimed by |
|---|---|---|---|---|
| Cost reduction | Too few to pool | at least 40% | 1 | 1 organization |
Value drivers: Lower cost to serve, Employee productivity.
Indicative value
A technology company spending USD 5 million a year on public cloud
USD 300,000 to USD 900,000
Annual cloud cost avoided per year
How this is calculated
Formula: annualCloudSpend * wasteShare * capturedSavings. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Annual public cloud spend annualCloudSpend, USD per year | 5,000,000 | 5,000,000 | The reference organization. Replace with your own annual cloud bill. |
| Share of spend that is waste or poorly optimized wasteShare, fraction of spend | 0.15 | 0.3 | Editorial assumption. AWS reports a median Cost Efficiency score of 83 and a mean of 79 across more than 71,000 opted in customers (AWS State of Cost Efficiency Report). That score measures the share of optimizable spend that is already well optimized, not the share of the total bill that is waste, so it does not map directly to this input; it is used only as a signal that most accounts still leave savings on the table. |
| Share of the identified waste the agent captures capturedSavings, fraction of identified waste | 0.4 | 0.6 | Editorial assumption, not derived directly from the evidence on this page. VERMEG's reported figure is a cut in on demand AWS spend; part of that reduction moves into committed (Reserved Instance) spend rather than disappearing as a net saving, so it does not map cleanly onto "share of identified waste captured". Akamai's 40 to 70% figure covers only the workloads Cast AI manages, not an organization's total cloud bill. Both are used only as a signal that a well run agent captures a meaningful share of identified waste, not as a direct calibration of this range; replace with your own track record. |
What it leaves out: Gross savings only. It leaves out the subscription cost of the optimization platform, the engineering time to configure and monitor it, and the performance risk of aggressive rightsizing if guardrails are too loose.
Who already uses it?
2 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
Akamai Technologies
United States · Technology and software · 2024
Akamai runs a large, complex Kubernetes estate on Microsoft Azure (AKS) behind services with strict SLAs. It uses Cast AI's automated bin packing, selection of the most cost efficient compute instances, Spot instance lifecycle automation and Kubernetes cost analytics to optimize the cost of its core production infrastructure. Akamai reports savings of 40 to 70% depending on the workload, its engineering team reports reclaiming time previously spent manually tuning capacity, and the interviewee describes the value of being able to "turn on and forget" the platform.
- Cost reduction: at least 40%, ongoing
"The core savings we got are just brilliant, falling between 40-70%, depending on the workload."
Claimed by: organization
VERMEG
Netherlands · Technology and software · 2024
VERMEG, a provider of software solutions to the worldwide financial services industry, used nOps to manage its AWS commitments (Reserved Instances) with a buyback guarantee across its accounts. nOps uses AI and analytics to monitor usage continuously and automatically adjusts those commitments to match it, instead of VERMEG locking in a fixed multi year forecast. Over ten months VERMEG cut its on demand AWS costs by more than 39% with no change to its own infrastructure or configuration, and gained continuous visibility into cloud spend.
No outcome disclosed.
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- Billing and usage export from every cloud account in scope
- Resource tags that map spend to a team, product or workload owner
- Existing reserved capacity and commitment inventory
- A policy for what may change automatically versus what needs approval
Systems to integrate
- Cloud billing and cost management APIs (AWS, Google Cloud, Azure)
- Container orchestration platform for rightsizing (Kubernetes)
- Infrastructure as code or the compute provisioning API, for executing approved changes
- Chat tool for savings reports and approval requests
- Finance or FP&A system for chargeback and budget tracking
Complexity: Medium
Reading usage and billing data is a fast, low risk integration. The work is in tagging resources to the right owner and workload, setting the policy for what the agent may change on its own, and building trust before commitment purchases run unattended.
- 1
Start with visibility, not action
Connect billing and usage data, tag spend to owners, and publish a shared view of waste before the agent changes anything. This is also what earns engineering trust.
- 2
Automate the reversible changes first
Idle resource cleanup, storage tier changes and commitment purchases that come with a buyback or exchange guarantee are safe to automate early; changes to a running production workload are not.
- 3
Write the action policy
Decide by resource type and environment what the agent may do alone, what needs an engineer's approval, and what stays out of scope entirely (for example anything touching a regulated workload).
- 4
Rightsize with a rollback plan
Every automated rightsizing change should be reversible within minutes; monitor error rates and latency after a change and roll back automatically if they move.
- 5
Attribute savings to the team that owns the workload
Report savings and remaining waste by team and product, not only as one company wide number, so the incentive to keep costs down sits with the people who can act on it.
Guardrails
- Read only access to billing and usage data; write access limited to an explicit allow list of actions
- A policy layer that separates automatic actions from actions that need an engineer's approval
- Every automated change is reversible, logged and attributed to the agent and the policy that allowed it
- Commitment purchases carry a buyback or exchange option rather than a fixed, unchangeable term
- Alerting on cost, error rate and latency after every automated change, with automatic rollback on regression
KPIs to instrument
- Realized savings versus identified savings, by team and workload
- Share of recommendations executed automatically versus approved manually versus rejected
- Commitment coverage and utilization, to catch both under and over commitment
- Incidents or rollbacks caused by an automated change
Human in the loop
Engineers approve any change to a running production workload and any commitment above a set size. Platform and finance teams review the action policy periodically and adjust what the agent may do alone as trust builds.
Common failure modes
- Savings that break performance
- An aggressive rightsizing change starves a workload at its next traffic peak. Keep a margin on autoscaled resources and monitor after every change.
- Commitment lock in
- A long term commitment is bought against a workload that later moves or shuts down. Prefer commitments with an exchange or buyback option and reforecast regularly.
- Tag debt hides the owner
- Untagged or mistagged resources cannot be attributed, so savings and accountability stall. Enforce tagging at resource creation and flag untagged spend separately.
- Automation nobody trusts
- Engineers turn off the automatic actions after one bad surprise. Start narrow, publish a clear action log, and expand scope only after a track record.
What are the risks and rules?
EU AI Act
Minimal risk
An internal tool that optimizes infrastructure spend and makes no decision about a natural person, so the default case falls outside Annex III. The relevant Annex III entry to check against is point 2, AI safety components in the management and operation of critical digital infrastructure: a cost agent stays outside it as long as its policy keeps it to cost actions (rightsizing, commitment purchases, idle cleanup) rather than acting as a safety component of the infrastructure itself. The Akamai deployment on this page shows the scope can extend to core production infrastructure, so an operator should confirm this against its own policy rather than assume it.
Rules that apply
Controls to put in place
- Inventory entry for the agent with an owner, its action policy and the resource scope it may touch
- Change log of every automated action, with the resource, the saving and who or what approved it
- Periodic review of the action policy as workloads and risk tolerance change
- Rollback tested and documented for every category of automated action
Frequently asked questions
- Can an AI agent be trusted to change production infrastructure on its own?
- It depends on the policy the operator sets. Akamai uses Cast AI's automated bin packing, instance rightsizing and Spot instance automation to optimize the cost of its core Kubernetes infrastructure, and the interviewee describes the value of being able to "turn on and forget" the platform. nOps, in VERMEG's deployment, held only read only permissions and managed VERMEG's AWS commitments (Reserved Instances) with a buyback guarantee, without altering any infrastructure or configuration. Neither case study states whether it kept a separate human approval step for individual changes; a safe general design restricts unattended action to reversible changes and routes anything that touches a running production workload to an engineer for approval.
- How much can an AI agent actually save on cloud costs?
- It depends heavily on how optimized the starting point already is. Akamai reports savings of 40 to 70% depending on the workload using Cast AI, covering only the workloads Cast AI manages, not Akamai's entire cloud bill. nOps reports that VERMEG cut its on demand AWS costs by more than 39% over ten months; part of that cut moved into committed Reserved Instance spend rather than disappearing as a net saving, so it is not directly comparable to Akamai's figure.
- Does this replace a FinOps team?
- No. It removes the manual, repetitive part of the work, tracking usage, spotting waste and managing commitments, so a smaller team can cover more infrastructure and spend its time on architecture decisions and the policy that governs what the agent may do.
How to cite this page
Blits.ai AI Use Case Library, "AI agent for cloud cost optimization and FinOps", last verified 28 September 2026, https://www.blits.ai/ai-use-cases/cloud-cost-optimization-agent. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 28 September 2026: First published