Repository navigation
Debugging When your instrumentation lies
This page owns: the four places this project trusted a tool and was wrong — and how to avoid being the fifth.
Debugging assumes the debugger tells the truth. Most of the time it does. The exceptions are expensive out of all proportion, because when your instrument is wrong you do not get a wrong answer — you get a confident wrong answer, and you follow it.
Four of them here, each of which cost real time.
The emulator's breakpoints fire once the instruction at the address has run.
The consequence that matters: a breakpoint on a faulting instruction never fires for that fault. The exception is taken first, so control leaves before the breakpoint is reported.
This produced the project's most memorable dead end. A breakpoint on a crashing instruction was hit 18,171 times, every hit showing healthy registers, and every conclusion drawn from those hits was about the wrong instruction — because the one that actually faulted never reported at all.
What to do instead: break on the exception vector, not on the suspect instruction. A
conditional breakpoint on the data-storage vector gives you SRR0 — the address the processor
itself recorded — which is authoritative.
debug.breakpoints.add 0x300 "machine.cpu.dar == <address>" "logical"
machine.memory.peek and poke do not go through the MMU.
For a KSEG0 pointer — anything beginning 0x8 — the fix is subtraction:
physical = virtual - 0x80000000
For anything else, including drivers loaded into paged system space, you must translate:
machine.cpu.mmu.translate(0xEE326BC0) → 0x004BEBC0
The symptom of getting this wrong is not an error. Reading an unmapped physical address
returns all-ones, so you get a tidy screenful of 0xffffffff and a plausible story about
uninitialised memory. That happened here: a buffer was declared empty and the read declared
broken, when the read had been fine and the probe was looking half a gigabyte past the end of
RAM.
The 604 in little-endian mode routes a word access at A to physical A ^ 4, and a byte
access at A to A ^ 7. The shell's memory access does not do this for you.
guest word at A → machine.memory.peek.l(A ^ 4)
guest byte at A → machine.memory.peek.b(A ^ 7)
So a patch documented as "the instruction at image 0x54748" is written:
machine.memory.poke.l 0x5474c 0x4086003c
Both of those are correct, and the gap between them is a trap. Read a poke address as though it were the guest address and you will spend an hour chasing an off-by-four that does not exist. That is precisely what happened on this project, complete with a confident and wrong diagnosis of "my script has an off-by-one" — when the script was right and the reading of it was not.
The fourth one is the nastiest, because it survives all three fixes above.
The processor's DAR — the address a faulting access referenced — is reported with the munge
still applied. So is the first parameter NT prints in a 0x50 bugcheck.
reported 0xEE315C98
^ 4 → 0xEE315C9C ← what the code actually asked for
Four bytes. Enough to make every subsequent calculation come out almost right, which is far worse than coming out obviously wrong.
In wall 49 this produced a genuine impasse: a register value derived from DAR came out four
bytes below a value plainly visible in memory, and the two refused to reconcile for an hour. Both
numbers were correct; one of them was munged.
Rule: un-munge DAR and bugcheck addresses before doing arithmetic with them.
Twice in one day an experiment reported a clean negative that was infrastructure, not a finding.
An ARC-tree test reported zero floppy nodes — the daemon it needed was not running, and the
whole run had failed in five lines. The rerun reported zero again — this time because rebuilding
the emulator had invalidated the checkpoint (checkpoint.c refuses a build-ID mismatch;
tools/restamp-ckpt.py
fixes that when no checkpointed structure changed). Both times the filter counted only
FloppyDiskPeripheral, so "the veneer built no node" and "nothing happened at all" produced
identical output.
Count a sentinel as well as the thing you are looking for. The third attempt reported
dump_node: 31 FloppyDiskPeripheral: 2 — the 31 is what makes the 2 trustworthy.
A related trap in the same session: machine.screen.checksum() with no arguments hashes
stride × height, and the stride padding is off-screen video memory the display driver uses as
scratch. It churns while the picture is perfectly still, so a "has the screen changed?" test
built on it is always true. The visible-region form takes four arguments.
Worse than a broken filter, because it looks like knowledge. "The veneer does not enumerate the
floppy" was measured with /chosen bootpath pointing at the hard disk, and written down without
that condition. Point bootpath at the drive and the node appears immediately.
The observation was correct; the generalisation was wrong; and it was the generalisation that got published — where it would have sent the next person to rewrite a firmware they did not need to touch. Write the condition beside the result, especially when the result is a negative.
bad address to DMA-MAP-IN looks like a firmware limitation and is usually a missing mapping.
map-space is a method of /packages/pe-loader; issued outside dev /packages/pe-loader it is
an unknown word, the buffer is never mapped, and the subsequent read-blocks fails in a way that
reads exactly like "this device cannot do DMA". The floppy reads perfectly once mapped —
EB 3C 90 'MSDOS5.0', byte for byte.
I reported that failure as a hardware conclusion twice before checking the mapping.
Every headless script in this project used to begin the same way: if the 0 > prompt did not
appear on the serial port, type setenv output-device ttya / setenv input-device ttya on the
ADB keyboard and restart, so the shell could read Open Firmware over machine.scc.a.sent().
Convenient — and it meant Open Firmware never opened its display driver. The Cirrus 54M30 went
into NT with every register at its reset value.
A user's browser has no serial port. Open Firmware's console is the screen, its Cirrus FCode
driver runs before NT, and it leaves SR07 = $F1 and SR17 = $66 behind: linear addressing on,
memory-mapped BLT registers at the top of linear space. cirrus.sys read-modify-writes SR17,
keeps bit 6, and then writes its BLT registers at $B8000 — where nothing decodes them any more.
Every register write became a pixel, START read back $0, not one BLT ever ran, and GUI-mode
Setup came up as the backdrop with one grey < Back button (wall E28, 17 September 2026).
Meanwhile every headless run drew the wizard perfectly — for a day — and each "it works headless" result was measuring a boot no user ever performs. The disk, the floppy, the HAL, the CD, build flags, pacing, retained state and wasm codegen were all eliminated one by one before anyone compared the two environments' event streams. The differential took two lines:
browser: [54m30] 1 SR17 = $66: memory-mapped BLT registers at the top of the linear aperture
headless: [54m30] 1 SR17 = $04: memory-mapped BLT registers at $B8000
The lesson is section 6 again in a different coat: a negative result taken under a condition
the user is not in says nothing about the user. Reproduce with the console on the screen,
typing through machine.adb.keyboard.type and reading progress from machine.screen.checksum()
— see Debugging recipes §9 — and reach for a differential log before a
hypothesis.
Calibrate a debugger against a value you already know before you trust what it tells you.
Concretely, on arriving in an unfamiliar environment:
- Read something whose value you are certain of — a known string, a known instruction — and check that it comes back right. If it does not, you have just learned a transformation you would otherwise have learned the hard way.
- When two measurements disagree by a small, constant amount, suspect a transformation before
suspecting a bug. Four, seven and
0x80000000are the three constants on this machine, and each corresponds to one of the sections above. - When a measurement is too clean — every word
0xffffffff, every hit healthy — suspect that you are measuring the wrong thing.
None of this is specific to this emulator. Every layer between you and the hardware — a JTAG probe, a hypervisor, a simulator — has its own version, and the cost of finding out empirically is always higher than the cost of one deliberate calibration.
- The emulator shell — the primitives, with these caveats applied.
- Decoding a bugcheck — where all four of these bite at once.
-
STORY.md— lesson 6, and walls 39 and 49 where these were learned.
Corrections welcome — this wiki is edited directly, so nothing here has had a review. Repository · STORY.md · GPL-2.0-only
Start here
Theory
- Why NT on a Power Mac is hard
- Open Firmware
- ARC
- The veneer
- The NT boot chain
- The HAL contract
- The NT PowerPC ABI
- Little-endian PowerPC
- How Setup chooses a HAL
- The NT video stack
Machines
Emulator
- Getting Granny Smith
- Media you must supply
- Building the HAL
- The boot floppy
- Preparing disks
- Running text-mode Setup
- Capturing the installed image
- Booting the installed system
- Iterating on the HAL
- Checkpoints and deltas
- Making an OEM CD (retired)
Real hardware
Debugging
- The emulator shell
- Reading NT binaries
- Decoding a bugcheck
- When your instrumentation lies
- Debugging recipes
Reference
- HAL exports
- ARC environment variables
- The veneer's VrDebug bitmask
- Veneer patch catalogue
- Address and interrupt map
- Error codes seen
Project