NorthPeakCloud
A padlock over lines of code

Guide · 2026

AI agent governance: a maturity model for enterprises

Most organisations building AI agents today are at governance Level 0: agents deployed across business units with no central registry, no safety evaluation before release, and no runtime enforcement of what they're allowed to touch. That gap — not model capability — is the real blocker to scaling AI agents safely. This guide sets out a four-level maturity model and the concrete steps to move up it.

Photo: Pixabay
Azure AI Foundry specialists
FinOps cost engineering
Fixed-price foundations
UK · US · Canada
Key takeaways
  • Level 0 (ungoverned) is where most enterprises sit today — agents built and deployed with no central visibility, which is a bigger risk than shadow IT because agents hold API access to production systems.
  • Governance, not model capability, is what determines whether AI agents can scale past a pilot: a registry, pre-deployment safety evaluation and runtime enforcement are the three controls that separate Level 0 from Level 2+.
  • Automated adversarial testing (red teaming) finds attack variations — prompt injection, jailbreaks, data disclosure — orders of magnitude faster than manual review, and should gate any agent before general release.
  • You don't need a governance platform to start: a registry, a review board and a documented evaluation gate get an organisation to Level 1 with a spreadsheet and a process.

Why governance, not capability, is the constraint

Every cloud provider can host models. Every framework can build agents. The question that actually determines whether an organisation can run AI agents at scale is whether it can prove — to a regulator, a customer or its own board — what those agents are allowed to do, and that they only do it.

Most organisations can't answer that today. Agents get built by whichever team has the skills and the urgency, deployed with production API access, and left running with no central registry. That's a bigger problem than shadow IT ever was: shadow IT was a spreadsheet nobody knew about, shadow agents have write access to production systems and can act on it at machine speed.

A four-level maturity model

Level State What's missing
0 — Ungoverned Agents built and deployed ad hoc across business units No registry, no pre-deployment evaluation, no runtime enforcement
1 — Visible Every agent is registered somewhere central, with an owner Evaluation is inconsistent; enforcement is still manual
2 — Controlled Registration is mandatory; agents pass a documented safety evaluation before release Enforcement is policy-based but not automated at runtime
3 — Automated Runtime policy enforcement, automated red teaming in CI, full audit trail per agent action Ongoing tuning as agent capability and blast radius grow

Most enterprises we see sit at Level 0, with pockets of Level 1 where a platform team has taken the initiative unilaterally. Very few have reached Level 2. That gap is the single biggest determinant of whether an AI agent programme survives contact with security, legal and a nervous board.

What moving up actually requires

Level 0 to 1 — get visibility. Stand up a registry: what agents exist, who owns them, what systems and data they can touch. This does not require new tooling. A spreadsheet and a mandatory intake process closes most of the gap in weeks, not quarters.

Level 1 to 2 — add the evaluation gate. Before an agent goes from pilot to general release, it should clear a documented safety evaluation: what happens on bad input, what its blast radius is if it misfires, whether its actions are reversible. This is where automated red teaming earns its keep — generating adversarial inputs (prompt injection, jailbreaks, information disclosure attempts) at a scale and speed manual testing can't match, and surfacing failures before a real user or attacker does.

Level 2 to 3 — enforce at runtime, not just at release. Policy checks that ran once at deployment need to keep running: rate limits, scope restrictions, approval gates on high-risk actions, and a full audit trail per action so an incident is traceable to a specific agent, a specific decision, and a specific input.

Building the baseline without a platform

You do not need a dedicated governance platform to reach Level 1 or 2. The baseline is process, not product:

  1. A registry — even a spreadsheet — with every agent, its owner, its data access and its blast radius.
  2. A review board — platform engineering, security and whoever owns the affected business process — that signs off before general release.
  3. A documented evaluation gate — a checklist an agent has to clear, applied consistently, not ad hoc.
  4. A runtime kill switch — the ability to disable an agent's access immediately, without a deployment cycle, the day something goes wrong.

Platform tooling that automates registration, evaluation and runtime enforcement becomes worth the investment once the manual process starts to strain — usually once an organisation is running double digits of production agents. Before that point, discipline beats tooling.

Getting help

We help engineering leaders put an agent governance baseline in place before capability outruns control — usually as part of the same AI landing zone work that gets the network, identity and cost guardrails right in the first place.

Related: Is your Entra ID ready for AI agents?, Azure AI Foundry: a practical guide and Azure AI Foundry consultancy.

Questions

Asked and answered.

What is AI agent governance, and how is it different from general AI governance?+

AI agent governance specifically controls what autonomous agents are allowed to do once deployed — which systems they can call, what data they can access, and how their actions are logged and reversible. General AI governance covers model selection and content risk; agent governance covers action risk, because an agent takes real actions in production systems, not just generates text.

What is 'Level 0' AI agent governance and why is it dangerous?+

Level 0 means agents are built and deployed across an organisation with no central registry, no safety evaluation before release and no runtime enforcement. It's dangerous because an ungoverned agent typically holds API access to production data and systems — a single misconfigured agent can act at machine speed across an entire estate before anyone notices.

Do I need Microsoft Entra Agent ID or a dedicated agent governance platform to get started?+

No. A registry, a review board and a documented evaluation gate can be run with a spreadsheet and a process, and gets most organisations to a defensible governance baseline. Platform tooling becomes worth adopting once agent volume outgrows a manual process — typically double digits of production agents.

Get in touch

Let’s get you AI-ready.

Tell us what you’re trying to do with AI on Azure — a couple of sentences is plenty. You’ll get a reply within one working day with clear next steps. No obligation, UK, US and Canada welcome.