The famous failure stat is making the rounds again. This time it's in Linas Beliūnas's newsletter, under a headline I'd happily frame and hang on the office wall: 95 percent of bank AI pilots fail because chatbots can't run compliance workflows. The number goes back to MIT's GenAI Divide report and has been picked up across banking research since. I've quoted it myself more than once. Guilty.
What's new is the diagnosis. And for once it's not "wrong model" or "bad data." It's the unit of work. A chatbot automates conversations. A bank doesn't run on conversations. It runs on cases.
A conversation is a session. It opens when the customer types, it answers, it closes when the customer leaves, and it remembers nothing. Perfect for "what are your opening hours." Useless for almost everything a bank actually gets paid to do.
Think about what real banking work looks like. An onboarding is a document collection that runs for days and ends in a compliance decision. A collections arrangement is a promise you have to track for three months. A KYC review touches the core, the screening provider, the document store and a human approver. It all runs long, it all crosses systems, and it all has to end in something an auditor can reconstruct.
Take one business account opening and you'll see it. Day zero: the owner of a small company starts an application in chat. That's the conversational part, and it takes about six minutes. Day one: the registry check and sanctions screening run, and the bank asks for proof of who actually owns the company. Day three: the documents come in and get checked, and one passport scan turns out to be expired. Day eight: nobody has uploaded a new one, and the application is about to lapse, whether anyone types or not. Day twelve: decision, account opened in the core, and a record the compliance team may one day have to show a supervisor.
One case, five systems, a dozen state changes. And exactly one of those steps looks like a conversation. Automate those first six minutes brilliantly and you've automated the doorbell of a house you never entered.

Forrester explained why better than anyone this year: a long running agent doesn't behave like a chatbot, it behaves like a distributed system. And distributed systems need orchestration, identity and context discipline that most companies have never built. That's the 95 percent in one sentence. The pilots that die automated the conversation. The ones that survive automate the outcome.
The market is quietly admitting the same thing. Intercom builds one of the best horizontal support agents out there, and it offers a money back guarantee on its Fin agent: a million dollars if you don't reach a 65 percent resolution rate. That's a floor Intercom itself describes as the resolution rate of humans. Klarna's famous assistant handled two thirds of its support chats in month one. Two thirds. 65 percent. The best conversational AI in production keeps landing in the same band.
Coincidence? I don't think so. And it's not a model problem either. It's simply the share of work that actually is conversational. The last third isn't harder questions. It's the disputes, the onboardings, the arrangements. The cases. You can't deflect a case, because it was never a ticket to begin with. What's left is work that carries state, and no amount of prompt engineering turns a goldfish into a filing cabinet.
We see this every day in our own delivery. The onboarding workflow we built for document collection keeps track of every candidate in a database: which documents arrived, which were rejected and why, when the last reminder went out, when to escalate. The friendly chat that asks for a passport photo? Honestly, that's the easy part. Most of the build lives in state, deadlines and write backs. That matches what the newsletter reports: specialist banking AI teams spend 80 percent of their engineering on compliance and integration. It also mentions a UK card provider that rebuilt its disputes process around a case running agent: resolution went from five to seven days down to two to three, and first contact evidence rates from under 25 percent to 80. That matches what we see on every case shaped project we've shipped.
So if cases are where the value is, why does every pilot start with a session bot?
Because pilots get picked on how well they demo. A session fits in a steering meeting: ask a question, watch a beautiful answer, approve the budget. A case takes twelve days to show, and most of those days look like nothing is happening. Which is exactly the point, and exactly what nobody puts on a slide. Sound familiar? I've sat in that steering meeting more times than I can count.
So companies keep funding the work that fits the meeting instead of the work that carries the money. Then they file the result under "AI didn't deliver." The 95 percent isn't just a technology failure. A big part of it is a buying bias: we bought what demos in minutes to automate work that lives in weeks.
And even when a pilot picks the right unit, there's a second brake. I see almost every project slow down on stakeholder management. If I break down where the time really goes, it's roughly 20 percent technology, 20 percent customization and 60 percent approvals and stakeholders. That's not a complaint. It's how large organizations work, and a case touches more of them than a chatbot ever will: risk, compliance, legal, IT, the business owner.
20 percent technology, 20 percent customization, 60 percent approvals. Plan for the 60.
Over the years we've become experts in that 60 percent, and we know how to fast track it. But almost nobody puts it in the pilot plan. Which is why so many pilots never make it past the steering meeting.
Strip it down and a case running agent needs four things a session bot doesn't have. Memory that survives the conversation, so state in a database and not in a prompt. The ability to act across systems: read the core, write the ticket, call the screening API. An understanding of who may decide what, so the agent drafts and a named human approves. And a clock.
That last one is the most underrated. Cases have deadlines, waiting periods and follow ups. A session bot experiences time as one conversation. A case runner experiences it as a calendar.
Paul is writing a deeper piece on what that coordination layer looks like in engineering terms, so I'll leave the architecture to him. The strategic point is simple, and I'll keep repeating it: "which model do you use" is still the wrong first question. Models are interchangeable inside a case runner. The machinery around them isn't.
Want to know what kind of AI your pilot really is? There's a test so simple you can run it in the steering meeting. Ask what your AI does on day eight.
An application started on day zero. The owner hasn't replied since day three, when the bank asked for a valid passport. On day eight, the application lapses in seventy two hours. Does anything happen? Does the AI send a reminder, alert the relationship manager, flag the file? Or is it just sitting there, fully intelligent and completely idle, waiting for someone to type?
If your AI only acts when someone messages it, you built a chatbot. If it acts because a deadline arrived, you built a case runner.
None of this needs a better model. All of it needs state, timers, integrations and authority. The boring machinery that separates the five percent from the ninety five. So before the next pilot gets approved, change the demo script. Don't ask the vendor for a beautiful answer. Ask them to show you day eight.
And if your onboardings, disputes or collections still run on an AI with the memory of a goldfish, I'm always happy to compare notes.