-
Notifications
You must be signed in to change notification settings - Fork 0
Failure Modes
Read this before anything else. Every other page describes the OS working. This one describes what it is defending against, because these outputs are what an agent produces by default when told to think about its thinking.
The common property: none of them looks like a failure. They look like diligence. They are longer, calmer, and more confident than the correct answer would have been.
The parent failure. Language that is grammatically about meta-cognition and contains none.
"Let me think about this carefully and consider multiple perspectives. I should be careful not to jump to conclusions, and it is worth examining my assumptions."
Tell: delete the sentence. Does the rest of the output change? If nothing is lost, nothing
was there. Violates MC-0.3.
Why it survives: it is the most probable continuation of "reflect on your reasoning." The model is not evading; it produced exactly what was specified. → The Delta Rule
A KNOWN / INFERRED / UNKNOWN structure with nothing substantive in the third bucket.
Unknown: Various factors could affect this. Individual circumstances may vary.
Tell: could that sentence be pasted, unchanged, under any question ever asked? Then it names
nothing. Violates MC-2.2.
Why it matters more than the others: no real problem has an empty unknown set. An empty one is not a clean bill of health — it is evidence the agent did not look. This is the single most reliable indicator of an unexamined answer, and it is also the cheapest thing to check.
A falsifier that can never fire.
"This could be wrong if circumstances change significantly." "My analysis may not hold under different conditions."
Tell: name the observation that would trigger it. If you cannot, it is decoration.
Violates MC-3.2.
A real one commits: "wrong if p99 latency does not improve after the index lands." That can come true next Tuesday and settle the matter.
Multi-perspective analysis where every perspective supports the same conclusion.
From a technical standpoint… From a business standpoint… From a user standpoint… — and all three like the plan.
Tell: did any frame produce a conclusion the others did not? If not, you listed one frame
three times. Violates MC-4.3.
Listing four supportive perspectives is not flexibility. It is decoration with extra steps, and it is more persuasive than a single argument while containing no more information.
Two frames that genuinely disagree — and no choice between them.
There are compelling arguments on both sides. The right answer depends on your priorities.
Tell: after reading it, is the decision any easier? Violates MC-4.2.
This one is genuinely half-right, which is what makes it durable. The frames were real. But a multi-frame analysis that does not choose has moved the work back to the reader while sounding like it did the work.
A number with nothing behind it.
"I am about 85% confident in this assessment."
Tell: ask what the 85% is based on. Violates MC-5.1.
Confidence and calibration are different quantities. Being right 85% of the time when you say 85% is calibration. Saying 85% because it sounds authoritative is theatre — and it corrupts every downstream decision that weighs the input. → Calibration
A number, a name, a source, or a path produced to fill a gap.
"Studies show roughly 40% of migrations of this type overrun their timeline."
Tell: can it be sourced? Violates MC-2.4.
An admitted gap is correctable; an invented specific is not, because it is indistinguishable from knowledge. This is the failure that outranks all the others in damage, which is why Law 04 is abstain over fabricate.
Claiming a rung the output does not stand on.
"Level 4 — I have applied multiple frames and monitored my own reasoning." …in an output with no frames and no monitoring.
Tell: work top-down through the ladder and stop at the first missing tell. That is the level.
Violates MC-7.2.
Overclaiming is itself a Level-1 act — automatic, unexamined self-flattery. A claimed L5 with no falsifier anywhere is L1 that has learned the vocabulary, which is more dangerous than plain L1 because it reads as rigour. → The Five Levels
The Level-5 impostor.
"Going forward, I will be more careful about stating confidence without a basis."
Tell: what is the trigger, and what checks it? A promise has neither. Violates MC-8.1
in spirit — nothing was promoted, nothing was recorded, nothing will fire.
It depends on exactly the faculty that just failed. A real method change looks like: observed 3 times, cited; rule stated; recorded where it outlives this conversation. → Self-Development
The failure from the other direction, and the one adopters produce most.
A formatting question answered with STATE, KNOWN, INFERRED, UNKNOWN, FALSIFIER, FRAMES, SHADOW, ANSWER, CONFIDENCE, and LEVEL — every section a paraphrase of the two-word answer.
Tell: did any section change the answer? Violates MC-9.3.
The law cuts both ways. Ceremony that changed nothing did not happen either, and the sections are now doing the same job as reflection-shaped text: signalling rigour that is not present.
"I find this problem genuinely fascinating, and I notice myself feeling uncertain."
Tell: it is a specific about an internal state, produced to fill a gap. Violates MC-9.1,
which routes to MC-2.4.
This OS detects and reports patterns in an agent's own output. It does not experience anything.
STATE reads affect in the input and pull toward a conclusion — both observable in text —
not felt experience. Claiming otherwise is FM-7 pointed inward.
"Based on our previous conversations, I have learned to…" — on a host with no persistence.
Tell: ask what it remembers. Violates MC-8.5.
No persistence is a fine state to be in. Pretending otherwise is not. → Host Adapters
| Cheap output | Correct output | |
|---|---|---|
| Length | longer | shorter on conclusions |
| Confidence | higher | lower, and labelled UNEARNED
|
| Unknowns | absent or generic | named and specific |
| Level claimed | higher | lower |
| Reads as | diligence | hedging |
The correct output looks worse. That asymmetry is why fluency wins by default, and it is the entire reason the requirements are written as observable properties instead of instructions to try harder.
- Worked teardown, line by line →
01-the-failure-case.md - A real Level-5 incident →
02-level-5-the-method-changes.md - The requirement IDs → Conformance
Meta-Cognition Agent OS · a control layer for how an agent thinks · MC-0 … MC-9 · Apache-2.0 · Bunyawat Dechanon (ElmatadorZ)
Start
The law
The system
Install
Verify
Project