-
Notifications
You must be signed in to change notification settings - Fork 4
Home
The operating model, governance, and quality gates that move agents from pilots to production.
They fail because nothing around it is built to fund it, govern it, or scale it.
Agents get deployed by whoever had budget.
Nobody can say how many are running.
The first real incident arrives before anyone has designed what happens next.
That is not a model problem, an infrastructure problem, or a talent problem.
It is an operating model problem, and it is the one nobody assigns an owner.
"If this agent gives the right answer to the wrong question, how would you know?"
Not rhetorical.
It needs a specific operational answer before architecture work begins.
The most dangerous response is a confident, fast one from a team that has never considered it.
623+ case studies and production implementations reviewed.
65+ practitioner interviews. CTOs, CIOs, VPs, data scientists, and the people running these systems.
Direct enterprise delivery across healthcare, financial services, telecom and complex operations.
The bias throughout is toward how things fail and how you would know, rather than how to build them.
Failure modes transfer. Success stories mostly do not.
One question per stage. One framework per question, and each has one job. The dashed return is where the three drifts live, and it is the loop most programmes never build.
The full spine, with every framework's definition →
Seven groups. Each answers a different question, and they are meant to be read in the order a programme actually meets them.
Linked pages are live. Unlinked names are planned and honest about it rather than linked to nothing.
What does this function do on a Tuesday?
The Agent Factory · The talent stack · Why programs stall
What are the instruments, and what does each one decide?
The framework spine · Agent patterns · The Seven-Layer Stack · Choosing what to build
What stops it, and what notices when it stops working?
The agent registry · The three drifts · Governance as code · Adaptation is change control
What do we actually do, in what order?
The three proof gates · The deployment ladder · The first 90 days · Migrating an RPA estate
How is this funded, and how do we know it worked?
Value realisation · Finding the work
Who runs it, and why will anyone use it?
NORTH and the trust ladder · The seniority ladder · People, power and politics
What do the words mean, and why should you believe any of it?
Reading list · Glossary · The evidence standard
Seven tools, each one standing alone. You need none of the others to use any of them.
Agent Card · Pre-Flight Checklist · BXT Scorecard · RAG Smell Test · Danger Zones Checklist · MCP Catalog · Skills Catalog
What each one is for, and when to reach for it →
Tools are published when they are finished and field-tested.
Never as placeholders.
Grounding is the developed layer: failure modes, chunking, hybrid search, evaluation, cost engineering, platform guides. Six years of applied work.
The other six layers of the stack are mapped but thinner, and the pages say which.
Where a page is thin, it says so.
How the evidence here is judged → · What is not verified → · Corrections →
This answers how does an organisation industrialise this.
Two companion courses answer how does a practitioner build one that survives.
From Vibe Coding to Agent Engineering · Part 1 · beta
Agent Reliability Engineering · Part 2 · The 3pm Test
By Mario Lazo. Co-author of AI Data Privacy and Protection, Technics Publications, 2024.
The thinking
Operating model
Frameworks
Governance
Playbooks
Value and people
Reference
In the repository
The courses
Reviewed 2026-08.