What problem does it solve?
After the checkout, the questions start: where is my order, can I change the delivery date, how do I return this, when do I get my money back. BARK, which ships tens of thousands of boxes a month, describes these questions as quick to answer one by one but overwhelming together. They peak together (a sale, the holidays, a carrier delay) and the answer sits in three or four systems: the order management system, the carrier's tracking feed, the warehouse and the returns platform.
Returns are also expensive and emotional. A clumsy returns experience costs the next sale, and policy answers that are wrong (a refund promised outside the return window, a label for an item that cannot be returned) create cost and disputes. First generation chatbots pointed customers to a tracking link or a returns page. The step change is an agent that identifies the order, reads the live status, and completes the change or the return itself, inside the retailer's rules.
- The National Retail Federation and Happy Returns projected that total retail returns in the United States would reach USD 890 billion in 2024, with retailers expecting 16.9% of annual sales to be returned.NRF and Happy Returns Report: 2024 Retail Returns to Total $890 Billion (2024)
- In the same NRF and Happy Returns survey, 67% of consumers said a negative return experience would discourage them from shopping with a retailer again.NRF and Happy Returns Report: 2024 Retail Returns to Total $890 Billion (2024)
How does it work?
- Identify the customer and the order. The agent matches the customer to the order from the logged in session, an email or phone number and a verification step, so the customer does not have to find an order number.
- Read the live status. It combines the order management system, the warehouse status and the carrier's tracking events, and explains them in plain words: packed, handed to the carrier, delayed at the depot, out for delivery.
- Change what can be changed. Within set rules it updates the delivery address or slot, cancels an order that has not shipped, or reschedules a delivery through the carrier's API.
- Run the return. It checks eligibility against the return policy (window, item type, condition), offers the options the retailer allows (refund, exchange, store credit, drop off or pickup), creates the return and sends the label or QR code.
- Explain the refund. It tells the customer when and how the money comes back, from the payment and refund status, not from a guess.
- Hand over exceptions. Damaged or missing items above a value threshold, suspected fraud, complaints, repeat failures and emotional conversations go to a person with the order and the conversation attached.
- Audience
- Customer facing
- Autonomy
- Supervised agent
- Adoption
- Mainstream
- Channels
- Web chat, Mobile app, WhatsApp, Phone and voice, Email
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
| KPI | Median | Reported range | Data points | Claimed by |
|---|---|---|---|---|
| Customer satisfaction | Too few to pool | 90% to 98% | 2 | 2 vendor |
| Interactions handled | Not pooled | 2.3 million | 1 | 1 organization |
| Response time reduction | Too few to pool | 82% | 1 | 1 organization |
Value drivers: Lower cost to serve, Customer experience, Speed and cycle time, Revenue growth.
Indicative value
An online retailer that ships 5 million orders a year
USD 450,000 to USD 3 million
Human handled contact cost avoided per year
How this is calculated
Formula: orders * contactsPerOrder * resolvedShare * costPerContact. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Orders shipped per year orders, orders per year | 5,000,000 | 5,000,000 | The reference retailer. |
| Post purchase contacts per order contactsPerOrder, contacts per order | 0.1 | 0.2 | Editorial assumption for order status, delivery and returns contacts. Replace with your own contact reason data. |
| Share of those contacts the agent resolves end to end resolvedShare, fraction of contacts | 0.3 | 0.5 | Editorial assumption within the range of the evidence. BARK's agent handled roughly a quarter of all customer conversations in its first year, Klarna's assistant two thirds of its service chats in its first month, and Ingka Group reports that its Billie chatbot resolved about 47% of all enquiries it received from 2021 to 2023. End to end resolution needs write access to orders and returns. Source |
| Cost of a human handled contact costPerContact, USD per contact | 3 | 6 | Editorial assumption for a blended chat, email and phone contact. Replace with your own fully loaded cost. |
What it leaves out: Gross avoided contact cost only. It leaves out the cost of running the AI and the integrations, the effect on repeat purchase of a better returns experience, fewer refund errors, and any change in return rates.
Who already uses it?
5 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
Best Buy
United States · Retail and ecommerce · 2024
In April 2024 Best Buy announced, with Google Cloud and Accenture, a generative AI virtual assistant for BestBuy.com, its app and its customer support line, expected to launch in late summer 2024, to help customers troubleshoot product issues, change order delivery and scheduling, and manage subscriptions and memberships. Google Cloud reported in its April 2026 list that Best Buy now guides shoppers through technical specifications, issue resolution and appointment scheduling autonomously, using Agent Assist powered by Gemini Enterprise for Customer Experience. It is a retail example of the same job a bank has when it books a branch or specialist appointment. No outcome figures for the assistant were published.
No outcome disclosed.
Klarna
Sweden · Payments and cards · 2024
Klarna announced in February 2024 that its AI assistant built on OpenAI models had been live globally for a month as the first line of its customer service, handling refunds, returns, payment issues, cancellations and disputes in more than 35 languages across 23 markets. In 2025 the company said it had gone too far in replacing people and began recruiting human agents again so that customers can always reach a person; the assistant still handles the majority of inquiries. The record is useful precisely because it shows both the gain and the correction.
- Interactions handled: 2.3 million, first month after launch
"The AI assistant has had 2.3 million conversations, two-thirds of Klarna’s customer service chats"
Claimed by: organization - Response time reduction: 82%, since launch, as reported in 2025
"Since launch, response times have improved by 82%, and Klarna has seen a 25% drop in repeat issues."
Claimed by: organization
BARK
United States · Retail and ecommerce · 2026
BARK, the company behind BarkBox, ships tens of thousands of boxes a month and deployed an AI agent called Scout to take the high volume, repetitive conversations first: where is my order, what is my tracking number, when will it arrive, and basic account questions. Within the first year the agent handled roughly a quarter of all customer conversations and chat response times fell below two minutes, while emotional conversations, such as the loss of a pet, stay with the human team.
- Customer satisfaction: 98%, first year of the agent, customer service overall
"BARK maintained a 98% customer satisfaction rate, holding the same high bar as the Happy Team."
Claimed by: vendor
Next
United Kingdom · Retail and ecommerce · 2026
Next, the British fashion retailer that operates across 83 countries and also serves customers of brands such as Gap, Victoria's Secret and Fat Face, launched an AI agent in six weeks across two use cases. The agent handles arrange return queries, matches the customer to the order without asking for an order number and manages identity verification. It adapts to regional preferences, so Next can add languages and processes as it grows, and it meets customers across chat, voice and WhatsApp. No outcome figures were published.
No outcome disclosed.
Sun & Ski Sports
United States · Retail and ecommerce · 2025
Sun & Ski Sports, a Texas based outdoor retailer with a strongly seasonal business, started its AI agent Sunny on basic returns and order status questions and then extended it to expert product advice on skis, boards, boots and bindings on its product pages. Sierra, the vendor, reports higher satisfaction on conversations the agent handles than on those transferred to humans, higher conversion for shoppers who engage with it, and a winter season without hiring temporary service staff.
- Customer satisfaction: 90%, conversations handled by the agent, as reported in October 2025
"Sunny achieves 90% customer satisfaction compared to 68% for conversations transferred to human agents."
Claimed by: vendor - Satisfaction uplift: 50%, as reported in October 2025, three years after the CMO joined in 2022
"Three years later, Sunny, their AI agent, has improved CSAT by 50% and tripled product page conversion rates"
Claimed by: vendor - Conversion uplift: 3x
"Customers who engage with Sunny convert at triple the rate of those who don't."
Claimed by: vendor
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- Order, shipment and return data reachable through APIs, with carrier tracking events normalised
- The return and refund policy as explicit rules (windows, excluded items, conditions, exceptions)
- Contact reason data per intent to choose the first scope
- Approved content for delivery, returns and warranty questions
Systems to integrate
- Order management system and ecommerce platform
- Carrier tracking and delivery management APIs
- Returns platform or warehouse returns process, including label generation
- Payment service provider for refund status
- Contact centre or helpdesk for handover with context
Complexity: Medium
Reading the order status is easy; acting safely is the work. The agent needs reliable APIs into order management, carriers and the returns platform, identity matching without an order number, and policy rules that are encoded rather than paraphrased.
- 1
Start with the contact reason report
Order status, delivery changes and return requests are often a large share. Pick the intents with the highest volume and the clearest rules, and leave damaged goods and disputes with people at first.
- 2
Encode the policy, do not paraphrase it
Turn the return window, excluded categories, condition rules and refund methods into rules the agent's tools enforce. The model explains the outcome; the rule decides it.
- 3
Solve order matching
Most customers do not have the order number at hand. Match on the logged in session, email or phone with a one time code, and confirm the item before acting.
- 4
Give the agent a narrow set of actions
Address or slot change before dispatch, cancellation before dispatch, return creation, label sending and refund status. Each action has its own limits, such as a maximum order value for an automatic refund.
- 5
Plan for peaks and carrier incidents
When a carrier has a regional delay, publish one explanation the agent uses for every affected order instead of letting it improvise, and scale the channel before the peak.
- 6
Measure resolution, not deflection
Count a conversation as resolved only if the customer did not come back on the same order within seven days, and read transcripts of the ones that did.
Guardrails
- Refund and exchange decisions come from policy rules in the tools, never from the model's own reading
- Value thresholds above which refunds, reshipments or goodwill gestures need a person
- Identity verification before any change to address, delivery or refund destination
- The agent never promises a delivery date the carrier has not given
- Automatic handover for complaints, suspected fraud, damaged goods above a threshold and distress
KPIs to instrument
- Resolution rate per intent, counting repeat contacts on the same order within seven days as unresolved
- Share of returns created end to end by the agent and their error rate
- Satisfaction on agent conversations versus human conversations for the same intents
- Average response time and time to refund
- Refund disputes and complaints that mention the assistant
Human in the loop
People own exceptions and judgment calls: damaged or missing items above a threshold, goodwill gestures, suspected return fraud and complaints. A team lead reviews a weekly sample of resolved conversations and every policy answer that led to a refund dispute, and signs off each new action before it goes live.
Common failure modes
- Promising what policy does not allow
- A fluent answer that grants a refund outside the window or for an excluded item. Keep eligibility in deterministic rules and have the agent quote the rule it applied.
- Tracking data the agent cannot interpret
- Carrier events are cryptic and sometimes wrong. Normalise them into a small set of states and let the agent say "we do not know yet" rather than guess.
- Return fraud through an easy channel
- An agent that issues refunds without checks becomes a target. Use value limits, customer history signals and a person for high value or repeat claims.
- Deflection dressed up as resolution
- Customers who give up look like contained conversations. Track repeat contacts and satisfaction per intent.
What are the risks and rules?
EU AI Act
Limited risk (transparency)
A customer facing service agent must disclose that the customer is interacting with AI (Article 50). It is not high risk: it does not decide on access to essential services, credit or employment.
Rules that apply
Guidance
- Article 50, transparency obligations for providers and deployers of certain AI systems (European Union, Europe). People must be informed that they are interacting with an AI system unless this is obvious from the context.
- Consumer rights directive (European Commission, Europe). Sets the EU right of withdrawal for distance purchases and the refund rules an agent's answers on returns must respect.
Controls to put in place
- AI disclosure at the start of each conversation
- Return and refund rules versioned with an owner, and regression tests on every policy change
- Audit log of every order change, return and refund the agent initiates
- Masking of payment data and personal data in logs and model prompts
- Monitoring of refund value and return volume initiated through the agent, with alerts on outliers
When it went wrong elsewhere
- Incident 639: Air Canada chatbot reportedly provides inaccurate bereavement fare information, leading to customer overpayment. Air Canada's website chatbot gave a customer inaccurate information about bereavement fare refunds. A Canadian small claims tribunal held the airline responsible for what its chatbot said and ordered it to pay damages. The same risk applies to any agent that answers return and refund questions without the policy encoded as rules.
Frequently asked questions
- What share of order and returns contacts can an AI agent resolve?
- It depends on whether the agent can act. Agents that only link to a tracking page resolve little; agents that can read the live status and create returns can do much more. On this page, BARK's agent handled roughly a quarter of all customer conversations in its first year, and Klarna's assistant, which handles refunds, returns and disputes, took two thirds of its service chats in its first month (in 2025 Klarna began recruiting human agents again).
- Should the AI decide refunds?
- No. Put eligibility and refund rules in deterministic tools with value limits, and let the agent explain the outcome. A Canadian tribunal held Air Canada responsible for what its website chatbot told a customer about bereavement fare refunds, so a refund answer must come from the policy itself.
- Does a better returns agent hurt sales?
- No source on this page measures the sales effect of a returns agent specifically. What is documented: in the NRF and Happy Returns survey, 67% of consumers said a negative return experience would discourage them from shopping with a retailer again, a measure of stated intent rather than an AI agent's effect. Separately, Sun & Ski Sports extended its agent from returns and order status into product advice, and its vendor reports that shoppers who engage with it convert at three times the rate of those who do not, a self selected comparison, not a controlled measure of the returns agent's effect on sales.
How to cite this page
Blits.ai AI Use Case Library, "AI agent for order status, delivery changes and returns", last verified 26 September 2026, https://www.blits.ai/ai-use-cases/order-status-and-returns-agent. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 27 September 2026: First published