-
Notifications
You must be signed in to change notification settings - Fork 4
What Is Not Verified
Every page here, and what it rests on. Read it before you quote anything.
A caveat repeated on twenty-five pages stops being read on the second one.
So the pages state their argument, and this page says what each argument rests on: what is measured, what is a design, what is drawn from practice, and what was deliberately dropped.
If you are about to cite something from this wiki in a document that matters, find it here first.
| What it means | |
|---|---|
| High | Canonical, in current use, or checkable against something already shipped |
| Medium-high | In current use, and part of the page is our reading rather than the source's claim |
| Medium | A design or a recommendation. It holds up in practice and it has not been measured |
| Low to medium | Drawn from experience and hard to check. Argued, not evidenced |
Most of this wiki is medium, and that is the honest answer for an operating model. A wiki where every page claimed high confidence would be one where nobody graded anything.
No named-company figure appears anywhere here. The source material behind this wiki carries roughly two dozen named-company cases, and every one of them reaches its number through a citation that does not resolve. They were excluded rather than re-sourced, which is safe rather than thorough. Any figure re-sourced later has to come from a primary document.
Thresholds are deliberately unstated. Accuracy bars, error-rate limits, canary percentages and durations all depend on volume and risk tier. A threshold copied from another organisation's numbers has no authority behind it, so the pages give the shape and leave the number to you.
high. Every claim here is either a link to a shipped artifact or a statement about this repository, both of which you can check.**
high for the definitions, which are in current use and shipped in the tools. The relationships drawn between frameworks are the author's own reading.**
The definitions are in current use. The relationships between them are this page's own reading. Which framework belongs at which stage, and how the companions relate to the spine, is an argument rather than a finding.
The four risk tiers are read off two shipped tools, so they are checkable. The governance question attached to each tier is a recommendation.
The framework count was retired rather than corrected. If a later page adds a framework and this table is not updated, that is a drift, and this page is where it will show.
high. This page describes artifacts you can open and read. The curation judgement about what counts as a tool is stated rather than assumed.**
high. Five failures observed repeatedly across engagements. The ordering is causal in argument, not measured.**
medium-high. The operating model is drawn from repeated delivery. The maturity levels are a working model, not a validated instrument.**
The four maturity levels are a working model, not a validated instrument. They were built from reviewed programmes rather than from a survey with a sample. They are useful for placing an organisation in conversation and they are not a score.
The time bands are reported, not measured. They come from programmes describing their own history, which is the least reliable form of evidence there is. Treat them as an order of magnitude.
The three-pillar split is an organising claim. Plenty of working functions distribute the same responsibilities differently. What is defensible is the failure-mode argument: remove any one of the three and the named failure follows.
The opening story is a composite and is labelled as one. No organisation matching that description is being referred to.
The staffing ratios are shapes, not benchmarks. They vary with how much platform already exists, and an organisation with a mature data function needs a materially different mix.
medium. The three roles are a design recommendation. The staffing ratios are the shape of one source's table and are a hypothesis to check against your own volume.**
The three roles are a design, not a study. They are drawn from what has worked and they are stated as a recommendation. No claim is made that a team without one of them fails at a measurable rate.
The staffing ratios are the shape of one source's table. The four-to-one engineer-to-architect ratio held across all four maturity levels in that table, which is suggestive and is not evidence. Treat it as a starting hypothesis to check against your own volume.
The enablement-starts-at-zero observation is this page's reading, not a claim the source makes. The source states the number without comment. The argument that it is backwards is ours and it rests on the adoption argument on another page rather than on data.
The four responses come from a chapter that illustrated them with a named company. The example is dropped; the four responses are general and survive without it.
medium-high. The eight patterns come from practice and are in current use. The default risk tiers per pattern are a starting position, not a measurement.**
The eight names and their definitions come from practice, not from a measured study. They are a working taxonomy that held up across a large body of reviewed deployments. They are not the only defensible cut and the boundary between F and A is genuinely blurry.
The tier defaults in the table are starting points, not rulings. A retrieval agent over a regulated corpus is not tier 1. The default is what to argue against, not what to accept.
The claim that most requests fit three patterns is not measured here. It is a reported observation from the source material, and the underlying counts are not published. Treat it as a reason to check your own portfolio, not as a benchmark.
The reconciliation against the current Agent Card is unresolved. The card ships seven values and three of these eight are not among them. Until that is settled, a card and this page will disagree.
medium. The seven layers come from practice, not from a standard. No industry body has ratified this ordering and the boundaries are judgements.**
The seven names come from practice, not from a standard. No industry body has ratified this stack and nothing about it is normative. Other credible cuts exist, and several of them merge Cognition into Orchestration.
The numbering direction is a choice made on this page. Counting from the user is conventional and it is not the only option. Anything citing "Layer N" from a source older than 2026-08-24 may be counting the other way.
The claim that grounding is the developed layer is a statement about this repository, not about the field. It reflects where the published work happens to be, which is a function of what the author has spent six years on.
The five-part expansion of Grounding comes from the same working material and its boundaries are soft. Retrieval and Injection blur in most real systems, and the split is useful for diagnosis rather than exact.
Layer 7 being a layer at all is a presentational decision, not an architectural one. It is argued for above and reasonable people place it as a plane instead.
medium-high. The three gates are in current use. The setup-time bands are calibratable and the split-the-process recommendation is the author's own.**
The three gates are a decision aid, not a classifier. They make the reasoning explicit and comparable across candidates. They do not settle genuinely ambiguous cases, and the split-the-process move above is a recommendation from practice rather than a measured result.
The setup-time bands are the source's and they are calibratable. They depend on integration surface and on how much of the platform already exists. A second agent on an established platform is not the same project as the first one.
The determinism axis is a heuristic. It is useful because it is fast, and its speed is exactly why it should not be the only check on a consequential decision.
Both example rows are constructed to show the placement logic, not drawn from engagements.
high. Canonical in the practice and used in client assessment since before this wiki existed. The promotion criteria are policy, not findings.**
The five levels are a synthesis from deployment work, not a survey. They have held up across assessments and they are stated at high confidence, but the boundaries are judgements about where governance changes, not measured discontinuities.
The promotion criteria are policy, not findings. The bands above are a design for how to earn a level. The specific durations and error rates in the source documents are context-dependent and are deliberately not reproduced as numbers here.
The direction is a ruling, not a consensus. The opposite convention exists in credible sources. This page states which one it uses, and why, so that a reader who meets the other one knows what happened.
The claim that most level 5 self-assessments are misclassifications is an observation from assessment work, not a measurement. It is the kind of claim that would need a sample and a rubric to establish properly.
medium-high. The three drifts and the morning routine are in current use. The fleet-level extension is this page's own reconciliation.**
The framework names three drifts, and the source's own primary case study monitors four dimensions (goal, context, model, collaboration). Model and Data appear in both. The reconciliation offered above, three per agent and the extra dimensions at fleet level, is this page's reading, not a statement made by the source.
The headline outcome figure attached to this ritual has been dropped. It belongs to a named company through a citation that does not resolve. The ritual is a better argument than the number was.
The scorecard thresholds are policy, not findings. The source states starting values. They are omitted here deliberately.
The example is reported, and it is anonymised. It is included for the mechanism, which is that a rate monitor caught what an accuracy monitor could not, and not as an outcome claim.
The four monitoring layers are a structure, not a measurement. They are useful because they make the second, third and fourth layers impossible to forget, not because the boundaries between them are precise.
medium-high. The three-layer stack and four pillars are designs in current use. No claim is made that programmes adopting them fail less often at a measured rate.**
The three-layer stack and the four pillars are designs, not measured interventions. They are stated as recommendations. No claim is made that programmes adopting them fail less often at a measured rate.
The specific thresholds in the source material are illustrative and are omitted here. Error-rate limits, cost-variance triggers and accuracy-drop percentages all depend on volume and risk tier. Set them locally, write them down before the agent runs, and treat a change to one as a change to the governance.
The regulator framing in the source is a composite and is treated as one.
The seven-versus-ten gate discrepancy is unresolved as to which is better. This page rules that the shipped artifact wins, which is a rule about authority rather than a finding about gate design.
high. The field list and the API-key rule are concrete and checkable. The rule assumes a gateway exists, and the page says what happens when it does not.**
The field list is a minimum, not a standard. It is drawn from working practice and no claim is made that it is complete, or that organisations using it have measurably fewer incidents.
The API-key enforcement rule assumes a gateway exists. In an environment where teams can call model providers directly with their own credentials, the registry is advisory and will decay. Say so out loud rather than pretending the rule holds.
The amnesty approach is a design, not a study. It is reported to work and it is the kind of claim that would need a before-and-after count to establish.
medium. The taxonomy is borrowed and cited. The gate questions and card requirements are original to this repository and are untested.**
The taxonomy is borrowed and the gate questions are not. The four categories are the paper's contribution, restated. Everything in the gate-question section is this repo's reading of what they imply operationally, and the paper does not endorse it.
The paper does not solve evaluation of adaptive systems. It names standardised evaluation as open work rather than answering it. Borrowing a taxonomy is fair; implying it arrives with an evaluation method it does not have would not be. If you are looking for how to test an adaptive agent, this page does not have it either, and neither does the source.
No claim here is measured. This is a governance argument built on someone else's classification. It earns its place by making a question askable, not by having been tested.
high. The gate structure has been run repeatedly. Every duration and threshold is a calibratable range rather than a finding.**
The durations are the source's ranges and they are calibratable. Two-to-four, four-to-eight, six-to-twelve weeks are starting shapes. A gate that always takes the maximum is telling you something about scope, not about the gate.
The hundred-case golden dataset is a convention, not a statistical claim. The right number depends on how varied the work is. The rule that survives is who builds it.
The four failure modes are drawn from practice and are stated as recurring patterns, not as a ranked or measured list. No claim is made about their relative frequency.
Every threshold on this page is deliberately unstated. Accuracy bars, tolerance bands and canary percentages depend on volume and risk tier. Set them locally, before the run.
high for the five rungs, which are in current use. All exit thresholds are deliberately unstated because they depend on volume and risk tier.**
The five rungs are a design, not a measured intervention. They are recommended, and no claim is made that programmes using them fail at a lower rate.
Every threshold is deliberately unstated. The source gives specific accuracy bands, durations and canary percentages. They depend on volume, risk tier and how costly a wrong output is, and a threshold borrowed from another organisation's numbers has no authority behind it.
The two examples are constructed to show the pattern, not drawn from named engagements.
medium-high. The score is a heuristic, not a model: its value is the argument it forces, not the number it produces.**
The formula is a heuristic, not a model. Multiplying two ordinal scores and dividing by a third is not measurement. It is a structured way to make a judgement comparable across candidates, and its value is the argument it forces, not the number it produces.
The threshold is a convention. Calibrate it against work you already know before using it on work you do not.
The Empathy claim is an observation from practice, not a measured effect. That people are more forthcoming about work they resent is reported and reasonable; it has not been tested here.
The worked example is constructed to show the method, not drawn from an engagement.
medium. The sequence is ordered by dependency rather than by calendar. The language-shift table is a communication design that has not been tested here.**
Ninety days is a structure, not a finding. The sequence is ordered by dependency: you cannot score candidates before you have walked the floor, and you cannot build before data access is confirmed. The calendar is a container for that ordering.
The eight-week build limit is a heuristic used as a scope test rather than a schedule. No claim is made that projects exceeding it fail.
The language-shift table is a communication design. It is reported to change how announcements land. It has not been tested here.
All quadrant examples are constructed. They illustrate the placement logic and are not drawn from engagements.
medium-high. The triggers are testable against your own estate today. The thresholds inside them are calibratable and the pattern preference order is this page's argument.**
The three-category split is a triage tool, not a taxonomy. Plenty of real bots sit between workhorse and zombie, and the point of the exercise is forcing a decision rather than achieving a correct classification.
The thresholds in the triggers are the source's and they are calibratable. The forty-percent maintenance ratio and the one-fifth exception rate are starting points. The direction of travel matters more than the crossing point.
The four patterns are a design. No claim is made about their relative success rates, and the preference ordering above is this page's argument rather than a measured finding.
All examples are constructed.
medium. The tiers and the unit economics are in current use. The opening scene is a composite and every illustrative percentage has been removed.**
Every percentage in the source material for this page is illustrative and has been removed rather than reproduced. Any number a reader needs has to come from their own baseline.
The three tiers are a framing, not a measurement standard. No accounting body recognises them. Their value is that they force a programme to say which kind of claim it is making before it makes one.
The conservative modelling rule is a heuristic. The specific discounts in the source are judgement, not derivation. What is defensible is the direction and the reason for it.
The opening scene is a composite and is labelled as one.
The claim that tier 3 value is systematically under-instrumented is an observation from practice, not a measured finding. It is offered as something to check rather than something to accept.
Verified against payroll is doing a lot of work in the month six row and it is the step most often skipped. Nothing on this page establishes how often it is skipped.
medium. The five dimensions and the ladder are a design. The scorecard bands are a policy, not a statistically meaningful boundary.**
NORTH is a diagnostic frame, not a validated instrument. The five dimensions were derived from observed adoption failures. No study establishes that a score of nineteen predicts anything, and the three bands are judgement calls presented as thresholds.
The scoring bands have no evidence behind the specific numbers. Twenty is not a statistically meaningful boundary. What is defensible is the ordering: a programme scoring low on Ownership fails differently and earlier than one scoring low on Habits.
The four-rung ladder is a policy design, not a finding. The promotion criteria in the source name specific durations and thresholds. Those are deliberately stated as bands here, because the right duration depends on volume, and thirty days at low volume proves nothing.
The rung direction is a ruling, not a consensus. The opposite convention exists in credible sources. This page states which one it uses and why.
The claim that over-reliance is a larger risk than resistance is an observation, not a measurement. It is reported from practice and it is the kind of claim that would need longitudinal data to establish properly.
low to medium, and deliberately so. This is the most experience-driven and least checkable page here. Research figures in the source carry citations that do not resolve and are stated as direction, not numbers.**
The research figures in the source carry citations that do not resolve. The monitoring prevalence, the share of pilots stalled by middle management, and the anxiety percentages are all stated here as direction rather than as numbers, deliberately. The direction is consistent across the material; the figures should not be quoted.
The mandate case is described from press coverage and the organisations are not named. It is included for the mechanism, not as an outcome claim.
The middle-manager argument is an interpretation. That resistance concentrates there is widely reported; the explanation offered here, that their power derives from navigating the complexity agents reduce, is a reading rather than a finding.
Nothing on this page is measured. It is the most experience-driven page in this wiki and it should be read as informed argument.
high on the definitions in current use, which are the ones with a linked page. This is a working vocabulary, not a ratified standard, and two entries are corrections rather than transcriptions.**
This is a working vocabulary, not a standard. Nobody outside this practice has ratified these definitions, and several terms have competing meanings elsewhere in the industry.
Two entries were corrected against their source rather than transcribed, and both are flagged where they appear. Carrying a definition this repo has already ruled against would have reintroduced the drift the ruling fixed.
The source glossary's examples were written for one vertical throughout. They are generalised here, which loses some vividness and gains transferability.
high. This page describes the filter actually applied to the pages published here. It is applied unevenly to older material in the repository, and it says so.**
This page describes a standard that is applied unevenly across the corpus behind it. The pages published here were written to it. Older material in the repository predates it.
The four-phase test is a filter, not a method. It rejects weak evidence. It does not tell you whether a good-looking system will work in your environment, and nothing here does.
The confidence labels are this repository's own convention. No external body defines them and other people use the same words differently.
high on the links, every one of which was checked with a live request on 2026-08-24, and not a claim about the content behind them.**
"Resolves" means the server answered, not that the content is still what the link description says. An HTTP 200 on a moved-and-redirected page is a weaker signal than it looks, and this check does not distinguish them.
The bot-blocked entries were confirmed to exist, not read. A 403 to an automated request is not evidence of anything except that the site blocks automated requests.
This list is deliberately short. The four resources/ files hold
roughly a hundred more working links. Most are fine and none of them earned
a place on a curated page by working.
---
Reviewed 2026-08. How the evidence here is judged · Back to Home
The thinking
Operating model
Frameworks
Governance
Playbooks
Value and people
Reference
In the repository
The courses
Reviewed 2026-08.