AI Technology20-10-20266 min read

The Coordination Layer: What to Build When Models Converge

Models are converging and getting cheaper to swap, so the durable engineering in enterprise AI is moving somewhere else: the layer that decides who may act, what is remembered, when a human takes over and what gets evidenced. A build guide for the part of your AI stack you'll still own in five years.

Paul Coerkamp
Paul Coerkamp
CEO & Co-Founder
The Coordination Layer: What to Build When Models Converge

Here is a useful question for any enterprise AI team: if your model provider retired your model tomorrow, what would you have to rebuild?

A year or two ago, the honest answer was often "almost everything." Today it should be "one configuration entry." For three years, enterprise AI strategy meant model strategy: which model, which vendor, which benchmark. That question matters less every quarter, because the leading models are converging and switching between them is increasingly a configuration change. What does not converge, and cannot be swapped in ten seconds, is everything around the model: the layer that coordinates people, agents and systems.

In this article I'll make that layer concrete: what it contains, why your engineering budget belongs there, and a checklist to test whether a team is building it or demoing around it.

Key message: the model is a dependency. The coordination layer is the asset.

Why the hard part moved

The evidence is piling up from three independent directions.

Forrester's 2026 agentic AI assessment lands on a structural diagnosis: "a long running agent doesn't behave like a chatbot: it behaves like a distributed system, and distributed systems demand orchestration, identity, and context discipline that most companies have never built." Anthropic's 2026 State of AI Agents report finds that nearly half of organizations, 46 percent, cite integration with existing systems as a primary obstacle to deploying agents, ahead of any concern about model capability. And Gartner now predicts that by 2030, more than 60 percent of early agentic orchestration implementations will fail to meet performance or cost expectations, because enterprises underestimate the integration, governance and talent requirements of making digital workforces reliable.

Read those three together and the pattern is clear. Organizations rarely fail on the model anymore. They fail on everything the model plugs into. Or, as Linas Beliūnas put it in his newsletter this month: "Coordination between humans and agents is the scarce layer once models converge." Len covered the unit of work half of this argument two weeks ago: sessions versus cases. This piece is about the machinery that makes case running work possible at all.

The four components

Across every regulated agent build we've delivered, the same four components keep being the actual work. The newsletter reports specialist teams spending 80 percent of engineering on compliance and integration infrastructure; our own ratio is in the same neighborhood, and it divides into these four.

1) The permission model

Not "is the user authenticated" but "which action requires which authority, per product, per amount, per jurisdiction." An agent that can read a balance, freeze a card and draft a settlement holds three very different powers, and each needs its own gate: some actions the agent takes alone, some need step up authentication from the customer, some need a named employee's approval. The common mistake is encoding this in the prompt. Prompts are suggestions. A permission model is enforced outside the model, where a confused agent cannot talk its way past it.

2) Case state

Memory that survives sessions, shifts and systems: what has happened in this case, what is pending, what was decided and by whom. State lives in a database with a schema, not in a context window. This is also where the clock lives: deadlines, waiting periods, retry schedules, escalation timers. The test Len proposed is the right one, so I'll repeat it: your agent should be able to act on day eight of a case in which nobody has typed anything since day one.

3) The handover protocol

Human in, agent out, and back again, with full context both ways. When confidence drops or authority requires it, a person takes over and sees the whole case, not a transcript dump. When the person resolves their part, the agent resumes the routine follow up. Most builds get the first half right and forget the second: the human closes their task and the case silently dies, because nothing resumes it. A handover is a two way door, and both directions carry state.

4) Evidence generation

Every automated step logged, attributable and reconstructable: which data was read, which model version answered, which check passed, who approved. Not as an afterthought export, but generated as the case runs, so that "show me how this decision happened" is a query, not a forensic project. This is where governance stops being policy and becomes code, and it's the part your auditor, and increasingly your regulator, will judge you on.

Why this layer is the defensible one

Everything below the coordination layer is commoditizing: models converge, get swapped and get repriced. Everything above it, the channels, changes with fashion: web chat today, WhatsApp tomorrow, voice next year. The layer in between is the only part that compounds. It encodes your products, your authority structure, your compliance requirements and your integrations, none of which a competitor can download.

There's a practical corollary for build versus buy. Whatever platform or vendor you choose, the question is not "how good is your model" but "show me your permission model, your case state, your handover protocol and your evidence trail." A vendor that answers with a benchmark chart is selling you the layer that's converging. That is also how we built our own platform: the four components are part of the platform itself, so a new use case starts with them in place instead of rebuilding them per project.

The checklist

Five checks that tell you within an hour whether a team is building the coordination layer or a demo:

  1. Kill the process and restart it. Does the case resume where it was? If state lived in the context window, the case is gone.
  2. Ask for the permission table. Which actions can the agent take alone, and where is that enforced? "It's in the system prompt" is a failing answer.
  3. Trigger a handover and come back. Does the human see the case or a transcript? When they finish, does the agent resume, or does the case die in a queue?
  4. Pick a closed case and ask how the decision happened. If reconstructing it takes engineering time, evidence generation doesn't exist.
  5. Swap the model. If anything other than a configuration entry changes, the model was load bearing in ways it should never be.

A team that passes all five can survive a model deprecation, an audit and a regulator visit in the same quarter. A team that fails three of them has built an impressive conversation, and will find out the difference in production.

Models will keep converging. The coordination layer is the part you own. If you'd like to see how we have built ours, with the permission tables, the case state and the evidence trail on real use cases, we'd be glad to walk your team through it. You can reach us here.

Paul Coerkamp
Paul Coerkamp
CEO & Co-Founder
Published on 20-10-2026

Related Articles

9 Things I Really Hate About AI
AI Technology12-05-2025

9 Things I Really Hate About AI

Let's be honest: I think it's great that technology is so embedded in our daily lives. It helps us get knowledge faster, complete tasks more efficiently, gives us inspiration, and occasionally scares the hell out of us with those crazy (fake) videos. I help a lot of companies implement AI, so in the end—it pays my bills. But after spending a ridiculous amount of time with all these new technologies, I feel it's time to reflect on the things I really hate about AI.

Read More →
A Conversation Is Not a Case
AI Technology06-10-2026

A Conversation Is Not a Case

The 95 percent pilot failure stat is making the rounds again, this time with a diagnosis I actually agree with: chatbots automate sessions, but a bank runs cases. Why the unit of work decides whether your AI survives production, and the day eight test that tells you which one you built.

Read More →
Agentic AI Languages and Dialects: Why Voice Quality Is Still the Hard Part
AI Technology10-04-2026

Agentic AI Languages and Dialects: Why Voice Quality Is Still the Hard Part

Agentic systems promise autonomous workflows, but speech and dialect quality fail first. Why Arabic, Turkish, Chinese and more hit the same gap.

Read More →

Stay Updated

Get the latest insights on conversational AI, enterprise automation, and customer experience delivered to your inbox

No spam, unsubscribe at any time