Here is a useful question for any enterprise AI team: if your model provider retired your model tomorrow, what would you have to rebuild?
A year or two ago, the honest answer was often "almost everything." Today it should be "one configuration entry." For three years, enterprise AI strategy meant model strategy: which model, which vendor, which benchmark. That question matters less every quarter, because the leading models are converging and switching between them is increasingly a configuration change. What does not converge, and cannot be swapped in ten seconds, is everything around the model: the layer that coordinates people, agents and systems.
In this article I'll make that layer concrete: what it contains, why your engineering budget belongs there, and a checklist to test whether a team is building it or demoing around it.
Key message: the model is a dependency. The coordination layer is the asset.
The evidence is piling up from three independent directions.
Forrester's 2026 agentic AI assessment lands on a structural diagnosis: "a long running agent doesn't behave like a chatbot: it behaves like a distributed system, and distributed systems demand orchestration, identity, and context discipline that most companies have never built." Anthropic's 2026 State of AI Agents report finds that nearly half of organizations, 46 percent, cite integration with existing systems as a primary obstacle to deploying agents, ahead of any concern about model capability. And Gartner now predicts that by 2030, more than 60 percent of early agentic orchestration implementations will fail to meet performance or cost expectations, because enterprises underestimate the integration, governance and talent requirements of making digital workforces reliable.
Read those three together and the pattern is clear. Organizations rarely fail on the model anymore. They fail on everything the model plugs into. Or, as Linas Beliūnas put it in his newsletter this month: "Coordination between humans and agents is the scarce layer once models converge." Len covered the unit of work half of this argument two weeks ago: sessions versus cases. This piece is about the machinery that makes case running work possible at all.
Across every regulated agent build we've delivered, the same four components keep being the actual work. The newsletter reports specialist teams spending 80 percent of engineering on compliance and integration infrastructure; our own ratio is in the same neighborhood, and it divides into these four.

Not "is the user authenticated" but "which action requires which authority, per product, per amount, per jurisdiction." An agent that can read a balance, freeze a card and draft a settlement holds three very different powers, and each needs its own gate: some actions the agent takes alone, some need step up authentication from the customer, some need a named employee's approval. The common mistake is encoding this in the prompt. Prompts are suggestions. A permission model is enforced outside the model, where a confused agent cannot talk its way past it.
Memory that survives sessions, shifts and systems: what has happened in this case, what is pending, what was decided and by whom. State lives in a database with a schema, not in a context window. This is also where the clock lives: deadlines, waiting periods, retry schedules, escalation timers. The test Len proposed is the right one, so I'll repeat it: your agent should be able to act on day eight of a case in which nobody has typed anything since day one.
Human in, agent out, and back again, with full context both ways. When confidence drops or authority requires it, a person takes over and sees the whole case, not a transcript dump. When the person resolves their part, the agent resumes the routine follow up. Most builds get the first half right and forget the second: the human closes their task and the case silently dies, because nothing resumes it. A handover is a two way door, and both directions carry state.
Every automated step logged, attributable and reconstructable: which data was read, which model version answered, which check passed, who approved. Not as an afterthought export, but generated as the case runs, so that "show me how this decision happened" is a query, not a forensic project. This is where governance stops being policy and becomes code, and it's the part your auditor, and increasingly your regulator, will judge you on.
Everything below the coordination layer is commoditizing: models converge, get swapped and get repriced. Everything above it, the channels, changes with fashion: web chat today, WhatsApp tomorrow, voice next year. The layer in between is the only part that compounds. It encodes your products, your authority structure, your compliance requirements and your integrations, none of which a competitor can download.
There's a practical corollary for build versus buy. Whatever platform or vendor you choose, the question is not "how good is your model" but "show me your permission model, your case state, your handover protocol and your evidence trail." A vendor that answers with a benchmark chart is selling you the layer that's converging. That is also how we built our own platform: the four components are part of the platform itself, so a new use case starts with them in place instead of rebuilding them per project.
Five checks that tell you within an hour whether a team is building the coordination layer or a demo:
- Kill the process and restart it. Does the case resume where it was? If state lived in the context window, the case is gone.
- Ask for the permission table. Which actions can the agent take alone, and where is that enforced? "It's in the system prompt" is a failing answer.
- Trigger a handover and come back. Does the human see the case or a transcript? When they finish, does the agent resume, or does the case die in a queue?
- Pick a closed case and ask how the decision happened. If reconstructing it takes engineering time, evidence generation doesn't exist.
- Swap the model. If anything other than a configuration entry changes, the model was load bearing in ways it should never be.
A team that passes all five can survive a model deprecation, an audit and a regulator visit in the same quarter. A team that fails three of them has built an impressive conversation, and will find out the difference in production.
Models will keep converging. The coordination layer is the part you own. If you'd like to see how we have built ours, with the permission tables, the case state and the evidence trail on real use cases, we'd be glad to walk your team through it. You can reach us here.