Skip to content

Debugging When your instrumentation lies

pappadf edited this page Sep 17, 2026 · 3 revisions

When your instrumentation lies

This page owns: the four places this project trusted a tool and was wrong — and how to avoid being the fifth.

Debugging assumes the debugger tells the truth. Most of the time it does. The exceptions are expensive out of all proportion, because when your instrument is wrong you do not get a wrong answer — you get a confident wrong answer, and you follow it.

Four of them here, each of which cost real time.


1. Breakpoints report after the instruction executes

The emulator's breakpoints fire once the instruction at the address has run.

The consequence that matters: a breakpoint on a faulting instruction never fires for that fault. The exception is taken first, so control leaves before the breakpoint is reported.

This produced the project's most memorable dead end. A breakpoint on a crashing instruction was hit 18,171 times, every hit showing healthy registers, and every conclusion drawn from those hits was about the wrong instruction — because the one that actually faulted never reported at all.

What to do instead: break on the exception vector, not on the suspect instruction. A conditional breakpoint on the data-storage vector gives you SRR0 — the address the processor itself recorded — which is authoritative.

debug.breakpoints.add 0x300 "machine.cpu.dar == <address>" "logical"

2. Memory access is physical, not virtual

machine.memory.peek and poke do not go through the MMU.

For a KSEG0 pointer — anything beginning 0x8 — the fix is subtraction:

physical = virtual - 0x80000000

For anything else, including drivers loaded into paged system space, you must translate:

machine.cpu.mmu.translate(0xEE326BC0)     →  0x004BEBC0

The symptom of getting this wrong is not an error. Reading an unmapped physical address returns all-ones, so you get a tidy screenful of 0xffffffff and a plausible story about uninitialised memory. That happened here: a buffer was declared empty and the read declared broken, when the read had been fine and the probe was looking half a gigabyte past the end of RAM.


3. Peek and poke do not apply the address munge

The 604 in little-endian mode routes a word access at A to physical A ^ 4, and a byte access at A to A ^ 7. The shell's memory access does not do this for you.

guest word at A    →   machine.memory.peek.l(A ^ 4)
guest byte at A    →   machine.memory.peek.b(A ^ 7)

So a patch documented as "the instruction at image 0x54748" is written:

machine.memory.poke.l 0x5474c 0x4086003c

Both of those are correct, and the gap between them is a trap. Read a poke address as though it were the guest address and you will spend an hour chasing an off-by-four that does not exist. That is precisely what happened on this project, complete with a confident and wrong diagnosis of "my script has an off-by-one" — when the script was right and the reading of it was not.

→ Little-endian PowerPC


4. DAR is munged too — and so is the blue screen

The fourth one is the nastiest, because it survives all three fixes above.

The processor's DAR — the address a faulting access referenced — is reported with the munge still applied. So is the first parameter NT prints in a 0x50 bugcheck.

reported   0xEE315C98
   ^ 4  →  0xEE315C9C      ← what the code actually asked for

Four bytes. Enough to make every subsequent calculation come out almost right, which is far worse than coming out obviously wrong.

In wall 49 this produced a genuine impasse: a register value derived from DAR came out four bytes below a value plainly visible in memory, and the two refused to reconcile for an hour. Both numbers were correct; one of them was munged.

Rule: un-munge DAR and bugcheck addresses before doing arithmetic with them.


5. A grep that counts only what you want cannot tell "absent" from "never ran"

Twice in one day an experiment reported a clean negative that was infrastructure, not a finding.

An ARC-tree test reported zero floppy nodes — the daemon it needed was not running, and the whole run had failed in five lines. The rerun reported zero again — this time because rebuilding the emulator had invalidated the checkpoint (checkpoint.c refuses a build-ID mismatch; tools/restamp-ckpt.py fixes that when no checkpointed structure changed). Both times the filter counted only FloppyDiskPeripheral, so "the veneer built no node" and "nothing happened at all" produced identical output.

Count a sentinel as well as the thing you are looking for. The third attempt reported dump_node: 31 FloppyDiskPeripheral: 2 — the 31 is what makes the 2 trustworthy.

A related trap in the same session: machine.screen.checksum() with no arguments hashes stride × height, and the stride padding is off-screen video memory the display driver uses as scratch. It churns while the picture is perfectly still, so a "has the screen changed?" test built on it is always true. The visible-region form takes four arguments.

6. A negative result is only as strong as the condition you took it under

Worse than a broken filter, because it looks like knowledge. "The veneer does not enumerate the floppy" was measured with /chosen bootpath pointing at the hard disk, and written down without that condition. Point bootpath at the drive and the node appears immediately.

The observation was correct; the generalisation was wrong; and it was the generalisation that got published — where it would have sent the next person to rewrite a firmware they did not need to touch. Write the condition beside the result, especially when the result is a negative.

7. map-space exists only inside the pe-loader package

bad address to DMA-MAP-IN looks like a firmware limitation and is usually a missing mapping. map-space is a method of /packages/pe-loader; issued outside dev /packages/pe-loader it is an unknown word, the buffer is never mapped, and the subsequent read-blocks fails in a way that reads exactly like "this device cannot do DMA". The floppy reads perfectly once mapped — EB 3C 90 'MSDOS5.0', byte for byte.

I reported that failure as a hardware conclusion twice before checking the mapping.

8. A headless run that never lets the firmware touch the hardware

Every headless script in this project used to begin the same way: if the 0 > prompt did not appear on the serial port, type setenv output-device ttya / setenv input-device ttya on the ADB keyboard and restart, so the shell could read Open Firmware over machine.scc.a.sent(). Convenient — and it meant Open Firmware never opened its display driver. The Cirrus 54M30 went into NT with every register at its reset value.

A user's browser has no serial port. Open Firmware's console is the screen, its Cirrus FCode driver runs before NT, and it leaves SR07 = $F1 and SR17 = $66 behind: linear addressing on, memory-mapped BLT registers at the top of linear space. cirrus.sys read-modify-writes SR17, keeps bit 6, and then writes its BLT registers at $B8000 — where nothing decodes them any more. Every register write became a pixel, START read back $0, not one BLT ever ran, and GUI-mode Setup came up as the backdrop with one grey < Back button (wall E28, 17 September 2026).

Meanwhile every headless run drew the wizard perfectly — for a day — and each "it works headless" result was measuring a boot no user ever performs. The disk, the floppy, the HAL, the CD, build flags, pacing, retained state and wasm codegen were all eliminated one by one before anyone compared the two environments' event streams. The differential took two lines:

browser:   [54m30] 1 SR17 = $66: memory-mapped BLT registers at the top of the linear aperture
headless:  [54m30] 1 SR17 = $04: memory-mapped BLT registers at $B8000

The lesson is section 6 again in a different coat: a negative result taken under a condition the user is not in says nothing about the user. Reproduce with the console on the screen, typing through machine.adb.keyboard.type and reading progress from machine.screen.checksum() — see Debugging recipes §9 — and reach for a differential log before a hypothesis.

9. The general lesson

Calibrate a debugger against a value you already know before you trust what it tells you.

Concretely, on arriving in an unfamiliar environment:

  • Read something whose value you are certain of — a known string, a known instruction — and check that it comes back right. If it does not, you have just learned a transformation you would otherwise have learned the hard way.
  • When two measurements disagree by a small, constant amount, suspect a transformation before suspecting a bug. Four, seven and 0x80000000 are the three constants on this machine, and each corresponds to one of the sections above.
  • When a measurement is too clean — every word 0xffffffff, every hit healthy — suspect that you are measuring the wrong thing.

None of this is specific to this emulator. Every layer between you and the hardware — a JTAG probe, a hypervisor, a simulator — has its own version, and the cost of finding out empirically is always higher than the cost of one deliberate calibration.


Next

Clone this wiki locally