-
Notifications
You must be signed in to change notification settings - Fork 0
Chapter 6: Interrupts, Panics, and Spinlocks
The previous chapter ended with build_physmap returning cleanly: all of physical memory is mapped into virtual space at PHYSMAP_VIRTUAL_BASE, the identity map is gone, and we are fully operating in the higher half.
What we do not yet have is:
- Any way to handle (and visually communicate) failure when something goes catastrophically wrong
- Any way to (maybe) gracefully recover from faults
- Any way to stop shared data structures from being torn apart by concurrent access (more on that later)
If we divide by zero right now, the CPU will realize that this operation is catastrophic, but because we have set up nothing to help it decide what to do, it will decide to reset the entire machine.
In this chapter, we're gonna fix all of that. As per the previous chapter, concepts here will not be overexplained. Google and LLMs exist, so you can always ask follow up questions if something you read here is not clear. I will not omit any crucial details.
Note
In QEMU, resetting the entire machine looks like "freezing" the current frame, or sometimes, it is just a plain black screen with nothing happening. Keep that in mind if you are following along.
Before we can discuss any of the aforementioned concepts, we need to establish what an interrupt actually is, because we have been using the word loosely throughout these docs without ever properly defining it.
An interrupt is the CPU's mechanism for responding to events that happen outside the normal flow of instruction execution. These events fall into two broad categories.
The first is exceptions. These are signals generated by the CPU itself when it encounters a condition it cannot resolve on its own. Dividing by zero is an exception. Accessing a virtual address with no page table entry (a page fault) is an exception. Executing an instruction the CPU does not recognize is an exception.
We have actually already encountered exceptions in a practical sense: back in Chapter 3, the entire purpose of the error function was to halt the machine when something went wrong during the 32-bit bootstrap, because at that point we had no proper way to handle them. What we are building in this chapter is the proper way.
The second category is hardware interrupts. These are signals sent to the CPU from external devices. A keyboard controller raises an interrupt when a key is pressed. A timer chip raises an interrupt at a regular frequency. The CPU notices the signal, finishes the current instruction it was executing, and then suspends normal execution to jump to a predefined handler function.
Important
If you search OS dev forums online, you might see people referring to hardware interrupts as "IRQs". In common conversation, the terms are often used interchangeably, but in kernel development, they represent different parts of the communication chain.
- IRQ (Interrupt Request): This is the input signal sent from a hardware device (like a keyboard or NIC) to the Interrupt Controller. It is a request for attention.
- Hardware Interrupt: This is the asynchronous event the CPU actually processes. It occurs when the Interrupt Controller "interrupts" the CPU's current execution flow to handle the IRQ.
In both cases, the mechanism is the same: the CPU pauses whatever it was doing and transfers control to a function the operating system has registered in advance. The critical word is pauses. An interrupt can fire between any two instructions, at any point in the call stack, completely invisible to the code that was running. The handler executes, and then, if everything goes right, execution resumes exactly where it left off, as if nothing happened.
This is a remarkably powerful mechanism. It is also the source of an entire class of subtle bugs, which is why spinlocks are the first topic of this chapter.
The Interrupt Descriptor Table (IDT) is essentially a lookup table that the CPU uses whenever any interrupt, exception or otherwise, fires.
Each entry in the IDT is indexed by a number called an interrupt vector.
An interrupt vector is just an index into the table, nothing more magical than that. You can think of it as:
"If interrupt number
Nhappens, jump to handler stored atIDT[N]."
So for example, in x86:
- Vector 0 → Divide by zero exception
- Vector 14 → Page Fault
- Vector 32+ → Hardware interrupts (timer, keyboard, etc)
The CPU uses these vectors to decide exactly which function to call when an interrupt or exception occurs.
Note
The IDT stores pointers to handlers. So the value stored in IDT[N] is the pointer to the handler function for interrupt vector N. If IDT[N] is not populated (i.e. we have no handler function for that vector), we move on to a double (and potentially a triple) fault.
The IDT is declared as a fixed-size array (on x86, typically 256 entries) where each entry describes:
- which function to call (the handler address)
- what privilege level it runs at (Ring 0, 1, 2, 3)
- what kind of gate it is (interrupt gate, trap gate, etc.)
So conceptually:
IDT = [
0 → handler for divide error
1 → handler for debug exception
8 → handler for double fault
14 → handler for page fault
...
]
When an interrupt happens, the CPU does not “search” or “compute” anything. It simply:
- takes the interrupt vector number
- indexes directly into the IDT
- jumps to the corresponding handler
That’s it, it’s a hardwired dispatch mechanism.
In the x86 architecture, gates are special entries in the Interrupt Descriptor Table (IDT). They define how the CPU transfers control into kernel code when an interrupt or exception occurs.
You can think of a gate as a controlled doorway: it doesn’t just say where to go, it also defines how the CPU should behave while going there (what to save, what to disable, what kind of context switch happens, etc).
The key difference between gate types comes down to three ideas: maskable interrupts, the Interrupt Flag (IF), and whether the CPU performs a full hardware task switch or just a normal function call into the kernel.
An Interrupt Gate is the standard mechanism used for hardware interrupts (like timers, keyboards, network cards).
The most important behavior here is how it interacts with maskable interrupts.
Note
A maskable interrupt is an interrupt that the CPU is allowed to temporarily ignore or delay. These are controlled by the Interrupt Flag (IF) inside the CPU’s FLAGS register. When IF = 1, interrupts are allowed. When IF = 0, maskable interrupts are blocked.
When the CPU enters an Interrupt Gate, it automatically clears IF (sets it to 0). This means:
While the handler is running, no other maskable interrupts are allowed to interrupt it.
This is important because hardware interrupts can be frequent. If one interrupt handler gets interrupted by another interrupt, and that one gets interrupted again, you can quickly end up with deep nesting, corrupted state, or stack exhaustion.
By disabling maskable interrupts during entry, the CPU ensures:
- the current interrupt handler runs atomically (in terms of interrupt delivery)
- the kernel can safely save CPU state without re-entrancy issues
Once the handler finishes, the IRET instruction restores the previous CPU state, including the original IF value, so interrupts are re-enabled automatically if they were enabled before.
A Trap Gate is very similar, but it behaves differently with respect to the Interrupt Flag.
When the CPU enters a Trap Gate, it does not modify IF. That means:
Maskable interrupts remain enabled while the handler runs.
Trap gates are typically used for:
- CPU exceptions (like divide by zero, invalid opcode, page faults)
- software interrupts / system calls (depending on design)
The reason this works safely is that exceptions are synchronous: they are caused by the current instruction stream. Unlike hardware interrupts, they don’t arrive randomly at high frequency. So there is less risk of interrupt storms or uncontrolled re-entrancy.
Keeping interrupts enabled during a trap handler is beneficial because it reduces latency: hardware events (like timer ticks or disk interrupts) can still be serviced even while handling the exception.
A Task Gate is a legacy mechanism tied to hardware task switching in x86.
Instead of pointing directly to a function, a Task Gate points to a Task State Segment (TSS). The TSS is a special CPU-managed structure that stores the entire execution context of a task (registers, stack pointers, segment state, etc).
When a Task Gate is triggered, the CPU performs a hardware task switch, which means:
- it saves the full state of the current task into its TSS
- it loads a completely new task state from another TSS
- execution resumes in a totally different context
This is very different from a normal interrupt or function call, because the CPU itself is doing the context switching automatically in hardware.
However, modern operating systems do not use this mechanism for scheduling anymore. The reason is simple: hardware task switching is rigid, slow, and less flexible than doing it in software. Instead, modern kernels perform context switching manually by saving registers and switching stacks in controlled code.
The only place Task Gates still occasionally appear is in very specific failure paths (like Double Fault handling), where the system wants a guaranteed clean execution context.
| Gate Type | What it does | IF behavior | Typical use |
|---|---|---|---|
| Interrupt Gate | Jumps into kernel and blocks further maskable interrupts | Clears IF (disables interrupts) | Hardware interrupts (timer, keyboard) |
| Trap Gate | Jumps into kernel but keeps interrupts enabled | Leaves IF unchanged | Exceptions, syscalls |
| Task Gate | Triggers full hardware task switch via TSS | N/A (full context switch) | Legacy / special fatal cases |
If you strip everything down, the mental model is:
- Interrupt Gate → “stop everything else while I handle this carefully”
- Trap Gate → “handle this, but keep the system responsive”
- Task Gate → “abandon this execution entirely and switch to another saved CPU state”
At the lowest level, the CPU provides two instructions that directly control whether maskable interrupts are allowed to interrupt execution: CLI and STI.
These instructions operate on the Interrupt Flag (IF) inside the CPU’s FLAGS register.
CLI (Clear Interrupt Flag) disables maskable interrupts by setting IF to 0.
When CLI is executed:
- The CPU immediately stops accepting maskable interrupts
- Any hardware interrupts that occur during this time are effectively blocked (or delayed until IF is re-enabled)
- The CPU continues executing normally, but in a “protected” state where it cannot be interrupted by external devices
This is typically used in very small, critical sections of kernel code where you must guarantee atomicity without introducing locking overhead.
However, CLI is not a general-purpose synchronization tool. If interrupts are disabled for too long, the system becomes unresponsive: timer interrupts stop firing, scheduling halts, and devices may appear frozen.
STI (Set Interrupt Flag) re-enables maskable interrupts by setting IF to 1.
When STI is executed:
- The CPU allows maskable interrupts again
- Any pending interrupts that were waiting while IF was cleared can now be delivered
- Normal interrupt-driven system behavior resumes
A subtle but important detail is that interrupts do not necessarily occur immediately after STI. The CPU typically enables interrupts and then checks for pending events at the next safe instruction boundary.
Together, CLI and STI form a very primitive but powerful mechanism for critical section protection at the CPU level. Unlike locks, they operate globally per CPU core, not per data structure. This makes them extremely fast but also extremely dangerous if misused.
In modern kernels, they are usually reserved for:
- very short critical sections in low-level code
- interrupt handler coordination
- per-CPU data structure updates where locking would be overkill
They are almost never used for long-running operations, because disabling interrupts effectively pauses part of the system’s ability to respond to the outside world.
In kernel and OS development (specifically on x86), these terms describe a hierarchy of failure during CPU exception handling. Think of them as levels of escalation when the processor tries (and fails) to handle an error.
A Single Fault is a standard CPU exception that occurs during normal code execution.
- What happens: The CPU encounters an instruction it cannot complete ( dividing by zero or accessing a memory page not currently in RAM).
- Resolution: The CPU looks up the appropriate handler in the Interrupt Descriptor Table (IDT) and jumps to that code to fix the issue. This is a normal part of system operation.
A Double Fault (Interrupt Vector 8) occurs when the CPU encounters a second exception while attempting to call the handler for the first one.
- What happens: Imagine a Page Fault occurs, but the CPU discovers the Page Fault handler itself is located on a "not present" memory page or the kernel stack is corrupt.
- Resolution: The CPU gives up on the first handler and tries to call a specialized Double Fault Handler. Kernel developers often use a separate, known good stack (via the Task State Segment or Interrupt Stack Table) for this handler to ensure it can run even if the main kernel stack is blown.
Note
We will talk about the Task State Segment (TSS) and the Interrupt Stack Table (IST) in later chapters, specifically when we set up userspace.
A Triple Fault is the final stage of failure where the CPU encounters an exception while trying to invoke the Double Fault handler.
- What happens: If the Double Fault handler itself cannot be called (eg. because the IDT is completely trashed or the GDT is invalid), the CPU enters a "Shutdown" state.
- Resolution: There is no "Triple Fault Handler." On modern PCs, the hardware detects this shutdown state and immediately triggers a system reset. This often manifests as an infinite "reboot loop" during early kernel development, or a frozen screen.
So, in a nutshell, we have:
| Fault Level | Trigger | Result |
|---|---|---|
| Single | Normal instruction error (eg. Page Fault) | Executes specific handler from IDT |
| Double | Error while calling the Single Fault handler | Executes Double Fault handler (Vector 8) |
| Triple | Error while calling the Double Fault handler | Instant hardware reset/reboot |
A spurious interrupt is an interrupt signal that arrives even though no real hardware event actually occurred. In other words, the CPU is told “an interrupt happened,” but when the kernel checks the device responsible, there is nothing meaningful to handle.
This sounds strange at first, but it is a real and well-known hardware behavior. It typically happens when a device signals an interrupt, but the signal disappears before the CPU acknowledges it, or when electrical noise or timing issues briefly trigger the interrupt line.
From the CPU’s perspective, it still goes through the full interrupt entry process: it looks up the IDT entry, pushes an interrupt frame, and calls the handler. But when the handler queries the device, it finds no valid cause.
Because of this, kernels often treat spurious interrupts as no-ops: they acknowledge them and immediately return without doing any work. They are essentially “ghost interrupts”, real at the hardware signaling level, but meaningless at the software handling level.
The Programmable Interrupt Controller (PIC) is the legacy hardware component responsible for managing hardware interrupts on early x86 systems. Its job is to collect interrupt signals from devices (keyboard, timer, disk controllers, etc.) and route them to the CPU using a limited set of interrupt lines.
The classic implementation is the Intel 8259 PIC, usually configured as a master-slave pair to extend the number of available interrupt inputs. Each device is assigned an interrupt request line (IRQ), and the PIC translates these into CPU interrupt vectors.
Before the CPU can handle an interrupt, the PIC must be explicitly acknowledged. This creates a multi-step handshake:
- Device raises an IRQ line
- PIC signals the CPU
- CPU invokes the interrupt handler via the IDT
- Kernel sends an End of Interrupt (EOI) signal back to the PIC
This design worked well for early uniprocessor systems, but it has several limitations in modern environments.
Modern systems have largely replaced the PIC with the APIC (Advanced Programmable Interrupt Controller). The reason is that the legacy PIC design does not scale well with modern hardware.
Key limitations of the PIC include:
- Limited interrupt lines (only 15 usable IRQs in practice)
- No efficient support for multi-core CPUs
- Centralized design, which becomes a bottleneck on SMP systems
- Lack of flexibility in interrupt routing
In contrast, the APIC architecture is designed for modern systems:
- Supports many more interrupt vectors
- Allows interrupts to be routed to specific CPU cores
- Enables inter-processor interrupts (IPIs)
- Scales properly in multi-core and NUMA systems
Because of this, modern kernels typically disable the PIC entirely during early boot and switch to the APIC system as soon as possible. Once APIC is active, the PIC is effectively obsolete and no longer participates in interrupt delivery.
GatOS disables the PIC completely, as we will see later on in this chapter:
void disable_pic(void) {
// reinit the PIC to a known state before masking
outb(PIC_COMMAND_MASTER, ICW_1);
io_wait();
outb(PIC_COMMAND_SLAVE, ICW_1);
io_wait();
// remap IRQs above the exception range (0-31)
outb(PIC_DATA_MASTER, ICW_2_M);
io_wait();
outb(PIC_DATA_SLAVE, ICW_2_S);
io_wait();
outb(PIC_DATA_MASTER, ICW_3_M);
io_wait();
outb(PIC_DATA_SLAVE, ICW_3_S);
io_wait();
outb(PIC_DATA_MASTER, ICW_4);
io_wait();
outb(PIC_DATA_SLAVE, ICW_4);
io_wait();
outb(PIC_DATA_MASTER, 0xFF);
outb(PIC_DATA_SLAVE, 0xFF);
LOGF("[APIC] Legacy PIC disabled and masked.\n");
}This code can be found in apic.c. It uses functions we have explained in previous chapters.
Note
We will talk about the APIC, as well as how to set it up and why, in a later chapter. For now, disabling the PIC is the only thing you need to know.
Keep this in mind, because we will call disable_pic later in this chapter.
Consider a simple scenario, at a high level.
We have a simple memory system that exposes a linked list of free pages. Allocating a page from anywhere in the code works in three steps:
- Read the head of the free list
- Update the head to the next node
- Return the original head page
Now imagine an interrupt occurs between steps 1 and 2.
The interrupted code has already read the head (say page A) but has not yet updated the list. Before it can continue, an interrupt handler runs and also allocates a page. It reads the same head pointer, also sees A, updates the head to A.next, and returns A.
When the original code resumes, it continues as if nothing happened. It sets the head to A.next and returns A as well.
As a result, both the interrupt handler and the original code believe they own page A, even though it was only supposed to be allocated once. Oh no!
Everything downstream of this is undefined behavior.
This is a data race, and in a kernel it typically manifests as silent memory corruption, allocator inconsistencies, or crashes much later in unrelated code that make it exceedingly difficult to debug. The cause is always the same: two actors accessing a shared, mutable state without coordination.
Note
Fun fact! This is actually so common in low level development that entire programming languages have been purpose built to avoid this type of memory corruption.
Rust is a striking example! It is designed to prevent data races at compile time in safe code via ownership and borrowing rules.
The solution is mutual exclusion: ensuring that only one actor at a time can modify a given data structure. The simplest possible implementation is a spinlock. The more high level, more robust and more complete implementation is usually a semaphore.
We will not be implementing semaphores in GatOS, purely because it is a single core kernel, not an SMP one.
A thread that wants to acquire a spinlock checks whether it is free. If it is, it takes it. If it is not, it spins (loops continuously checking) until the holder releases it. No sleeping, no scheduler involvement. Just burning CPU cycles until the lock becomes available.
In userspace, this sounds wasteful. In a kernel, for critical sections measured in a handful of instructions, it is often the right tool. The overhead of putting a thread to sleep and waking it later is real. For updating a freelist pointer, a spin is cheaper than a yield. Spinlocks are the first and most fundamental synchronization primitive we reach for.
GatOS's kernel spinlocks live in kernel/sys/spinlock.h and kernel/sys/spinlock.c.
typedef struct {
volatile int locked;
uint32_t cpu_id;
const char* name;
} spinlock_t;volatile int locked is the lock variable itself. The volatile qualifier tells the compiler that this value may change at any moment from outside the current execution context (specifically, from an interrupt handler) and therefore every read must go directly to memory rather than using a value the compiler cached in a register. Without volatile, an optimizing compiler might read locked once, decide it will not change, store it in a register, and loop on the register value forever, never noticing that the real value in RAM was updated by another actor.
cpu_id and name are debug aids. When you encounter a deadlock and need to know which lock is stuck and which CPU is holding it, these fields save a significant amount of time.
Here is why a naive implementation does not work:
// BROKEN
while (lock->locked != 0); // Wait until free
lock->locked = 1; // Take itImagine two CPUs running this simultaneously. Both read locked and see zero. Both pass the loop. Both write one. Both believe they hold the lock. Both enter the critical section at the same time. The lock has completely failed.
The problem is that checking and setting are two separate operations. Between them, another actor can intervene. We need a single operation that both reads the old value and writes a new value atomically, as one indivisible unit that the hardware guarantees cannot be interleaved with anything else.
__atomic_test_and_set provides exactly this:
while (__atomic_test_and_set(&lock->locked, __ATOMIC_ACQUIRE)) {
__asm__ volatile("pause");
}__atomic_test_and_set writes 1 into the memory location and returns the old value, all as one atomic hardware instruction. If the old value was 0 (the lock was free) the operation returns false and the loop exits: we have claimed the lock. If the old value was 1 (already taken) it returns true and we keep spinning.
__ATOMIC_ACQUIRE is a memory ordering constraint. It instructs the compiler and the CPU not to reorder memory operations that happen after this instruction to before it. Without acquire ordering, an optimizing compiler might speculatively move code from inside the critical section to before the lock acquisition, which would defeat the entire purpose. Acquire ordering ensures that the critical section stays fenced on its entry side.
The pause instruction is a performance hint to the CPU. During a spin-wait loop, pause prevents the processor from over-speculatively executing the loop's memory reads, which would otherwise flood the cache-coherency bus and degrade performance for the CPU that is actually doing useful work inside the critical section.
spinlock_release uses the mirror instruction:
void spinlock_release(spinlock_t* lock, bool interrupts_enabled) {
lock->cpu_id = 0xFFFFFFFF;
__atomic_clear(&lock->locked, __ATOMIC_RELEASE);
...
}__ATOMIC_RELEASE ensures that all memory writes made inside the critical section are visible to other actors before the lock is marked free. Without this, another CPU acquiring the lock immediately after might see stale data from before the critical section ran.
There is a second problem, more subtle than atomicity, and it is the reason spinlock_acquire does not simply spin, but it also saves and disables interrupts:
bool spinlock_acquire(spinlock_t* lock) {
bool was_enabled = intr_save();
while (__atomic_test_and_set(&lock->locked, __ATOMIC_ACQUIRE)) {
__asm__ volatile("pause");
}
lock->cpu_id = lapic_get_id();
return was_enabled;
}Consider this sequence on a single-core system:
- The kernel acquires a spinlock and enters a critical section.
- An interrupt fires. The CPU suspends the kernel and jumps to the interrupt handler.
- The interrupt handler also tries to acquire the same spinlock.
- The lock is held by the interrupted code. The handler spins.
- The interrupted code can never resume: it is preempted and waiting for the interrupt to return. The interrupt is waiting for the interrupted code to release the lock. Neither can ever make progress.
This is a deadlock, and it requires no multiple CPUs, just one interrupt firing at the wrong moment. The fix is to save the current interrupt state and disable interrupts before taking the lock, then restore the previous state when releasing it. An interrupt that arrives while the lock is held will simply be deferred until interrupts are re-enabled after the release.
intr_save and intr_restore live in arch/x86_64/cpu/interrupts.h:
static inline bool intr_save(void) {
uint64_t rflags;
__asm__ volatile("pushfq; popq %0" : "=r"(rflags) :: "memory");
bool enabled = (rflags >> 9) & 1;
if (enabled) __asm__ volatile("cli" ::: "memory");
return enabled;
}
static inline void intr_restore(bool enabled) {
if (enabled) __asm__ volatile("sti" ::: "memory");
}pushfq pushes the RFLAGS register onto the stack. We pop it into a general-purpose register and check bit 9, which is the Interrupt Flag (IF). If it is set, interrupts are currently enabled. We save that as a boolean and, if interrupts were on, clear them with cli. We return the saved state. intr_restore is the mirror: if interrupts were enabled before, sti brings them back.
The caller is responsible for threading the interrupt state through acquire and release:
bool flags = spinlock_acquire(&lock);
// ... critical section ...
spinlock_release(&lock, flags);If interrupts were already disabled when we acquired the lock (because we were already inside an interrupt handler) intr_save returns false, we do not call cli again, and intr_restore does nothing when we release. Nesting works correctly without any special handling.
spinlock_try_acquire is a non-blocking variant for situations where spinning would cause a deadlock:
bool spinlock_try_acquire(spinlock_t* lock, bool* was_enabled) {
*was_enabled = intr_save();
if (__atomic_test_and_set(&lock->locked, __ATOMIC_ACQUIRE)) {
intr_restore(*was_enabled);
return false;
}
lock->cpu_id = lapic_get_id();
return true;
}If the lock is already taken, we restore the interrupt state immediately and return false. The caller can decide what to do — try again later, skip the operation, or take a different code path. We will see exactly why this variant is necessary when we discuss the crash console shortly.
With spinlocks in place, we have the mutual exclusion primitive that every data structure from this point forward will rely on.
Now we need the rest of the infrastructure: a way for the CPU to actually deliver exceptions to code we control, rather than triple faulting every time.
Note
Interrupt implementations that will be referenced from this point forward can be found in ISR.S, interrupts.c and interrupts.h.
Each entry in the IDT encodes the address of a handler function and some configuration:
typedef struct {
uint16_t address_low;
uint16_t selector;
uint8_t ist;
uint8_t flags;
uint16_t address_mid;
uint32_t address_high;
uint32_t reserved;
} __attribute__((packed)) idt_entry_t;The handler address is split across address_low, address_mid, and address_high. This fragmented layout is a consequence of the IDT format being extended incrementally from 16-bit to 32-bit to 64-bit over decades of x86 history, without anyone being able to clean it up without breaking backwards compatibility.
selector is the code segment selector to use when the handler runs — always the kernel's 64-bit code segment, the same one we loaded in Chapter 3.
flags encodes the gate type: we use interrupt gates, which automatically clear the interrupt flag on entry, preventing interrupts from interrupting each other.
set_idt_entry populates one entry:
void set_idt_entry(uint8_t vector, void* handler, uint8_t dpl, uint8_t ist_index)
{
uint64_t handler_addr = (uint64_t)handler;
idt_entry_t* entry = &idt[vector];
entry->address_low = handler_addr & 0xFFFF;
entry->address_mid = (handler_addr >> 16) & 0xFFFF;
entry->address_high = handler_addr >> 32;
entry->selector = KERNEL_CS;
entry->flags = INTERRUPT_GATE | ((dpl & 0b11) << 5) | (1 << 7);
entry->ist = ist_index & 0x7;
entry->reserved = 0;
}Once all 256 entries are filled, we tell the CPU where the table lives by packing its address and size into a static idtr structure and loading it into the interrupt descriptor register with lidt:
void load_idt(void* idt_addr)
{
struct {
uint16_t limit;
uint64_t base;
} __attribute__((packed)) idtr;
idtr.limit = sizeof(idt) - 1;
idtr.base = (uint64_t)idt_addr;
__asm__ volatile("lidt %0" :: "m"(idtr));
}This is the same pattern as loading the GDT with lgdt in Chapter 3. The hardware now knows exactly where to find our handlers. Yay!
Here is the first real design challenge. When the CPU jumps to a handler after an interrupt, it does not tell the handler which vector just fired. It simply jumps to whatever address the IDT entry contains. If we pointed all 256 entries at the same C function, that function would have no way to distinguish a divide-by-zero from a keyboard interrupt.
The standard solution is to create 256 individual stubs in assembly, one per vector, each of which pushes its own vector number onto the stack before jumping to a single shared handler. GatOS generates all of these automatically using a GAS macro:
.macro GENERATE_INTERRUPT_HANDLER num
.align 16
.global interrupt_handler_\num
interrupt_handler_\num:
.if \num == 8 || \num == 10 || \num == 11 || \num == 12 || \num == 13 || \num == 14 || \num == 17
push \num
.else
push 0
push \num
.endif
jmp generic_interrupt_handler
.endm
.altmacro
.set i, 0
.rept 256
GENERATE_INTERRUPT_HANDLER %i
.set i, i+1
.endr.rept 256 repeats the macro 256 times, with i counting from 0 to 255. The result is 256 individually labeled functions.
interrupt_handler_0 through interrupt_handler_255, generated entirely at compile time from a handful of macro lines.
Every stub is aligned to exactly 16 bytes with .align 16. This is not just for correctness. We are gonna do a clever trick here. If we know the address of interrupt_handler_0 and add 16 bytes, we will land on interrupt_handler_1. In the same sense, we can generalize:
interrupt_handler_i = interrupt_handler_0 + (i*16)
Therefore, idt_init can calculate the address of stub i arithmetically, without needing 256 separate global symbols:
void* handler = (void*)((uint64_t)interrupt_handler_0 + (i * 16));Address of stub 0, plus i times 16, gives the address of stub i. Clean, and it saves the linker from resolving 256 individual symbol references.
Important
The .if inside the macro handles a hardware quirk. For certain exception vectors (8, 10, 11, 12, 13, 14, and 17 ) the CPU automatically pushes a 64-bit error code onto the stack before jumping to the handler, providing additional context about what caused the exception. For every other vector, it does not. This asymmetry means the stack layout differs between vectors with error codes and vectors without them.
The fix is to push a dummy error code of 0 for every vector that does not have a real one. After either path, the stack looks identical: vector number at the top, error code (real or zero) just below it. The shared handler can always find both at the same offsets.
The shared interrupt handler must begin by saving all general-purpose registers. At the moment the interrupt occurs, those registers contain the exact execution state of the interrupted code — essentially a snapshot of what the program was doing at that instant.
However, the handler itself is also just code, and it needs to use those same registers for its own execution. If it were to use them directly, it would overwrite the saved state and destroy that snapshot.
That would be a problem: when the handler finishes, the original code would resume with corrupted register values instead of the ones it had before the interrupt occurred.
To prevent this, the handler first saves the full register state in assembly, before doing anything else. This preserved snapshot is then restored right before returning, ensuring the interrupted code continues exactly as if nothing happened.
This saved register state is called the CPU context, and GatOS represents it using cpu_context_t. It is a structured snapshot of all general purpose and special purpose registers.
It is needed to fully pause and later resume execution exactly where it was interrupted.
typedef struct {
uint64_t r15;
uint64_t r14;
uint64_t r13;
uint64_t r12;
uint64_t r11;
uint64_t r10;
uint64_t r9;
uint64_t r8;
uint64_t rbp;
uint64_t rdi;
uint64_t rsi;
uint64_t rdx;
uint64_t rcx;
uint64_t rbx;
uint64_t rax;
uint64_t vector_number;
uint64_t error_code;
uint64_t iret_rip;
uint64_t iret_cs;
uint64_t iret_flags;
uint64_t iret_rsp;
uint64_t iret_ss;
} cpu_context_t;This layout (r15 - rax) is not arbitrary. It directly matches the exact stack layout produced during interrupt entry in our kernel.
When an interrupt occurs, the CPU itself automatically pushes a minimal “return state” onto the stack: ss, rsp, rflags, cs, and rip. This bundle is known as the IRET frame, because the iretq instruction later uses it to restore execution precisely as it was before the interrupt.
Important
This is why the ss, rsp, rflags, cs, and rip appear at the bottom of cpu_context_t. Remember that the stack is a LIFO (last-in, first-out) structure, so the last values pushed end up at the lowest addresses when viewed as a continuous block.
When we reinterpret the stack as a cpu_context_t, we are effectively mapping that memory layout directly onto a struct. That means the values pushed first by the CPU (and later by our stub) must correspond to the lowest fields in memory, while the values pushed last end up at the top of the structure.
After that, the interrupt stub adds two more values: the interrupt vector number and, if applicable, an error code. These identify what kind of interrupt occurred and why.
Finally, the generic interrupt handler saves all general-purpose registers (rax through r15) by pushing them onto the stack. This is necessary because the compiler assumes it owns those registers, but an interrupt must preserve the interrupted program’s state completely.
Because everything is pushed in a well-defined order, the stack at this point is no longer an ad-hoc sequence of values: it is a perfect, contiguous memory representation of cpu_context_t. In other words:
the current value of
rspis effectively a pointer to a fully constructedcpu_context_t.
No copying or reconstruction is needed.
generic_interrupt_handler:
test qword ptr [rsp + 24], 3
jz .no_swapgs
swapgs
.no_swapgs:
push rax
push rbx
push rcx
push rdx
push rsi
push rdi
push rbp
push r8
push r9
push r10
push r11
push r12
push r13
push r14
push r15
mov rdi, rsp
call interrupt_dispatcherThe instruction test qword ptr [rsp + 24], 3 inspects the CS value inside the IRET frame. On x86, the lowest two bits of the code segment selector encode the Current Privilege Level (CPL). A value of 0 means the interrupt occurred in kernel mode, while a value of 3 means it came from user mode.
This distinction matters for swapgs. The gs register is used in x86-64 kernels to point to CPU-local data (like the current thread structure). However, user space may also use gs for its own purposes (such as thread-local storage). swapgs switches between the kernel GS base and the user GS base.
Because this only matters when crossing the kernel/user boundary, we check the privilege level first and only execute swapgs when necessary. This avoids unnecessary overhead on kernel-to-kernel interrupts.
After all registers are pushed, rsp points exactly to the base of a cpu_context_t. We pass this pointer in rdi, which is the first argument register in the System V AMD64 calling convention, and call into the C-level interrupt dispatcher.
From this point onward, the interrupt system is fully in C land. The dispatcher receives a live pointer into the interrupted execution state, meaning it can inspect or modify the context directly. Any changes made to this structure will be reflected when the interrupt returns via iretq.
Before we dive into C land, we need a few core definitions that form the backbone of the interrupt routing system.
At the lowest level, we define the Interrupt Descriptor Table (IDT) and a parallel software dispatch table:
idt_entry_t idt[IDT_SIZE] = {0};
static irq_handler_t irq_handlers[IDT_SIZE] = {0};
extern char interrupt_handler_0[];The idt array is the hardware-facing structure used directly by the CPU to enter the kernel on interrupts and exceptions. Each entry tells the CPU how to transition into kernel mode for a given interrupt vector. We will initialize it in the next section.
The irq_handlers is an array of similar size, initially just NULL for every entry. Here we will store function pointers for each of the 256 interrupts we want to handle. If we don't want to handle something, we leave the entry NULL, otherwise, we store a function pointer that will be called.
Once inside the kernel, all interrupt handling is unified through a single function type:
typedef cpu_context_t* (*irq_handler_t)(cpu_context_t*);Each handler receives a full snapshot of CPU state (cpu_context_t) and is allowed to return a modified version of it.
This means:
- Input = exact machine state at interrupt time
- Output = potentially modified state used when resuming execution
So handlers are not just “responding” to interrupts, they can also influence how execution continues after iretq.
Finally, registering interrupt handlers is straightforward:
void irq_register(uint8_t vector, irq_handler_t handler) {
irq_handlers[vector] = handler;
}
void irq_unregister(uint8_t vector) {
irq_handlers[vector] = NULL;
}This keeps the system flexible without complicating the low-level interrupt path. The IDT remains static and hardware-focused, while the handler table provides a clean, dynamic interface for kernel subsystems.
We initialize the IDT early during boot:
void idt_init(void)
{
disable_pic();
for (size_t i = 0; i < IDT_SIZE; i++)
{
void* handler = (void*)((uint64_t)interrupt_handler_0 + (i * 16));
uint8_t ist = 0;
uint8_t dpl = DPL_RING_0;
if (i == INT_DOUBLE_FAULT) { ist = 1; }
else if (i == INT_PAGE_FAULT) { ist = 2; }
else if (i == INT_BREAKPOINT || i == INT_DEBUG){ dpl = DPL_RING_3; }
set_idt_entry(i, handler, dpl, ist);
}
load_idt((void*)idt);
}The first important step is disable_pic(). This is not optional. The legacy PIC can still deliver hardware interrupts during early boot (timer ticks, keyboard input, etc). If that happens before our handlers are ready, we triple fault. Disabling the PIC ensures that, during early initialization, only CPU-generated exceptions can occur.
The loop itself populates all 256 IDT entries using the stub arithmetic described earlier. Each vector jumps to a slightly different stub, but they all funnel into a shared handler (the Interrupt Dispatcher).
Two special cases are configured:
- Double fault gets a dedicated IST stack
- Page fault also gets its own IST stack (to maximize reliability)
- Debug/breakpoint exceptions are allowed from ring 3 for userspace debugging
Finally, load_idt activates the new table.
The ist field in an IDT entry refers to the Interrupt Stack Table, a hardware feature that forces the CPU to switch to a known-good stack before executing the handler.
This exists for one critical reason: sometimes the current stack is not safe to use.
The most important case is the double fault (vector 8). A double fault occurs when the CPU fails while trying to handle another exception. A classic example is:
- a page fault occurs
- the page fault handler itself cannot run (eg. invalid stack or corrupted memory)
- the CPU raises a double fault instead
Now consider what happens if the current stack is also invalid. The CPU still needs to push an interrupt frame to begin handling the double fault. If that push itself fails, the system enters a triple fault, which immediately resets the machine with no diagnostics.
To prevent this, we assign the double fault handler a dedicated IST stack:
if (i == INT_DOUBLE_FAULT) {
ist = 1;
}
else if (i == INT_PAGE_FAULT) {
ist = 2;
}This guarantees that even if the original kernel stack is corrupted or unmapped, the CPU has a clean, pre-allocated stack to use for recovery.
Caution
Be extremely careful here, because the code above is deceptive.
We are assigning IST values (1 and 2) in the IDT entries, which implies that the CPU will switch to dedicated, pre-allocated safe stacks for those exceptions.
However, the IST mechanism is a feature inside the 64-bit TSS (Task State Segment), NOT the IDT.
At this point in the boot process, we have not even touched the TSS, and we have not allocated or configured any IST stacks it is supposed to reference.
So what we are effectively doing here is cheating. The IDT is being set up as if the system is complete, but the supporting runtime state (TSS + IST stack memory) is not yet in place.
This means interrupts are effectively unsafe. We must not enable global interrupts (sti) yet, because any exception that tries to use IST would currently point into uninitialized memory.
The correct sequence is:
- Prepare the IDT as if everything else is set up
- Initialize our allocators
- Transfer GDT control from assembly into C
- Properly define and load the TSS
- Allocate and assign valid IST stacks
- Only then enable interrupts
We are currently in step 1. Until step 6, this configuration is purely preparatory. Structurally correct, but not yet operational. If interrupts were enabled now, any IST-triggering exception (like a double fault) would lead to an immediate and unrecoverable crash.
After the CPU enters the kernel and the register state is saved, control is passed to the C dispatcher, called by our generic_interrupt_handler:
cpu_context_t* interrupt_dispatcher(cpu_context_t* context)
{
uint64_t vec = context->vector_number;
if (vec == INT_SPURIOUS_INTERRUPT) {
return context;
}
if (irq_handlers[vec] != NULL) {
context = irq_handlers[vec](context);
if (vec >= INT_FIRST_INTERRUPT) {
lapic_eoi();
}
return context;
}
if (vec < INT_FIRST_INTERRUPT) {
// panic: unhandled exception
}
return context;
}The dispatcher is the central decision point for all interrupts. Its job is to classify the vector and route it appropriately.
Spurious interrupts are hardware-generated interrupts that have no actual source. They are harmless noise and are ignored immediately.
If a handler exists for the vector, it is invoked with the current CPU context. This is where device drivers and kernel subsystems handle interrupts in a structured way.
A handler can be registered from anywhere in the source (kernel space) using irq_register and irq_unregister. The function bound to that vector is then the handler for its interrupts.
After handling a hardware interrupt (typically vectors ≥ 32), the dispatcher sends an End of Interrupt (EOI) signal via lapic_eoi(). This tells the interrupt controller that the interrupt has been fully processed and the line can now be reused.
Note
We will explain the significance of EOIs in the APIC chapter.
If the interrupt is a CPU exception (vectors < 32) and no handler exists, the kernel treats this as a fatal condition and escalates to a panic. There is no recovery path for missing exception handlers.
It is important to note that this system is still in early initialization:
-
irq_handlersis empty - APIC drivers and the scheduler do not exist
- hardware interrupts are expected to be disabled via PIC shutdown
This is intentional. Early boot should fail loudly and deterministically. If an interrupt occurs here, it indicates either a misconfiguration or a serious initialization error, and the kernel should surface it immediately rather than silently continue.
Serial output, which we set up in Chapter 4, is an excellent debugging tool. But it requires either QEMU's -serial flag or a physical serial terminal to be visible. For a kernel panic, the kind of failure you want to be impossible to miss, we want something on the screen itself.
Welcome, kids! Ready to panic? Ready to glitch into infinity? Ready to enter a place where we scream as loud as hardware allows us to?
Good. The fun stuff begins here.
Back in Chapter 3, we asked GRUB for a linear framebuffer by including the framebuffer tag in our Multiboot2 header. GRUB fulfilled that request and populated the framebuffer info in the Multiboot2 information structure, giving us the framebuffer's physical address, width, height, pitch, and bits-per-pixel. We preserved all of that during multiboot parsing in Chapter 4.
In Chapter 5, we built the physmap: a direct linear mapping of every physical address into virtual space starting at PHYSMAP_VIRTUAL_BASE. One of the things the physmap covers is the framebuffer's physical address via the fb_pd huge-page mapping we set up.
This means we can now derive a virtual address for the framebuffer simply by adding PHYSMAP_VIRTUAL_BASE to its physical address. That is exactly what console_init does:
void console_init(multiboot_parser_t* parser) {
font_init();
multiboot_framebuffer_t* mbfb = multiboot_get_framebuffer(parser);
if (!mbfb) return;
fb_phys = mbfb->addr;
fb_w = mbfb->width; fb_h = mbfb->height;
fb_pitch = mbfb->pitch; fb_bpp = mbfb->bpp;
fb_sz = fb_h * fb_pitch;
fb = (uint8_t*)PHYSMAP_P2V(fb_phys);
fh = font_get_current()->header->charsize;
cols = fb_w / fw;
rows = fb_h / (fh + PADDING_Y);
kmemset(fb, 0, fb_sz);
}PHYSMAP_P2V is the same macro from paging.h that we introduced in Chapter 5. After this call, fb is a byte pointer directly into video memory. Writing to it writes pixels to the screen.
font_init loads a PSF1 bitmap font that is baked into the kernel image at compile time. This gives us fh (the font height in pixels) and a table mapping character codepoints to glyph bitmaps.
Important
The explanation here references font.c and font.h. These are readily available glyphs copied and pasted from google right into GatOS. They basically tell us what pixels to light up for each UTF-8 character, in order to print it to the screen through the framebuffer.
The PC Screen Font (PSF) is a bitmap font format used by the Linux kernel to display text on the console. It is widely used in OS development because of its simplicity, allowing a kernel or bootloader to render text without needing complex vector font libraries. You can look into it more if you'd like, but there isn't much use explaining this individually.
cols and rows are derived from the screen dimensions and font size and tell us how many characters fit in the display.
kmemset(fb, 0, fb_sz) clears the entire screen by writing zeros across the framebuffer.
That is all console_init does. It sets up the hardware state: the framebuffer pointer, the screen dimensions, the font. There is no allocation of any kind here, because we do not have a heap yet. There is no backbuffer, no output queue, no cursor management beyond a few static variables.
The full console abstraction (double-buffering, dirty tracking, UTF-8 decoding, ANSI escape processing) requires dynamic memory and will be introduced once the allocators are online. For now, what we have is a pointer to video memory and the ability to draw pixels. Hell yeah.
Writing a pixel to the framebuffer is as simple as writing to any other memory location, because through the physmap, it is exactly that:
static inline void put_pixel(uint32_t x, uint32_t y, uint32_t color) {
if (x >= fb_w || y >= fb_h) return;
uint8_t* dst = fb + y * fb_pitch + x * (fb_bpp / 8);
if (fb_bpp == 32) *(uint32_t*)dst = color;
else { dst[0] = color & 0xFF; dst[1] = (color >> 8) & 0xFF; dst[2] = (color >> 16) & 0xFF; }
}The address calculation fb + y * fb_pitch + x * (fb_bpp / 8) is worth understanding.
-
fb_pitchis the number of bytes per row, not necessarilyfb_w * (fb_bpp / 8), because hardware and firmware often pad rows to alignment boundaries. Using pitch rather than width ensures we always land on the correct row. -
fb_bpp / 8converts bits-per-pixel to bytes-per-pixel. For 32-bit color (the overwhelmingly common case on modern hardware) each pixel is four bytes encoding blue, green, red, and an unused alpha channel.
Rendering a character glyph is a loop that reads the bitmap for each row from the font data and calls put_pixel for every column:
static void draw_glyph(uint8_t* glyph, size_t px, size_t py, uint32_t fg, uint32_t bg) {
for (size_t y = 0; y < fh; y++) {
uint8_t row = glyph[y];
for (size_t x = 0; x < fw; x++)
put_pixel(px + x, py + y, ((row >> (7 - x)) & 1) ? fg : bg);
}
}Each byte in the glyph buffer represents one row of pixels. Bit 7 is the leftmost pixel, bit 0 is the rightmost. If the bit is set, we draw the foreground color; if clear, the background. This is enough to render text to the screen.
Here is the problem we now need to solve. Imagine the kernel is in the middle of something delicate, and a bug causes an exception. The IDT catches it, the dispatcher determines there is no registered handler, and we want to display a red screen with the fault address and the instruction pointer, for debugging purposes.
Remember that GatOS supports a full on TTY subsystem (on later versions). So why does the current panic.c and panic.h code not use that, through something like printf or a high level console abstraction? Why does it use con_crash_printf? Why do we need a special "crash" printf?
Two reasons.
First, the TTY subsystem (as well as many complex subsystems) depend on a working heap. This means that if we wanted to use them for panics, we would have to wait until all memory subsystems are online and kicking. This means that we would be "blind" for a significant portion of our initialization code. Normally, we want panic to be operational as soon as possible, to make debugging easier.
Secondly, we actually have another much more sinister problem. If we try to initialize the heap, and something goes wrong (say, corruption), we will try to panic. But panicking would mean using the TTY subsystem for output, which in turn depends on the heap being operational. But the heap was just corrupted, so panic corrupts as well.
This second case especially, must NEVER happen.
The crash console solves this by having no dependencies at all. It is a set of functions (con_crash_clear, con_crash_puts, con_crash_printf) that write directly to the framebuffer through the bare fb pointer, with their state tracked in a handful of static file-scope variables:
static uint32_t ccx = 0; // crash cursor x (in characters)
static uint32_t ccy = 0; // crash cursor y (in characters)
static uint8_t cfg = CONSOLE_COLOR_WHITE;
static uint8_t cbg = CONSOLE_COLOR_RED;
static char cbuf[2048]; // formatting buffer (on stack in callers, but this is shared)No heap. No locks. No TTY. No backbuffer. The cursor position is plain integers. The color is a plain integer. The output buffer is a static array with fixed size.
crash_emit renders one byte directly to the framebuffer:
static void crash_emit(uint8_t c) {
uint32_t row_h = (uint32_t)fh + PADDING_Y;
if (c == '\n') { ccx = 0; ccy++; }
else if (c == '\r') { ccx = 0; }
else if (c == '\t') { ccx = (ccx + 4) & ~3u; }
else {
if (ccx >= (uint32_t)cols) { ccx = 0; ccy++; }
if (ccy >= (uint32_t)rows) crash_scroll();
uint32_t px = ccx * 8;
uint32_t py = ccy * row_h;
uint8_t* glyph = get_glyph(c);
if (glyph)
for (uint32_t y = 0; y < (uint32_t)fh; y++) {
uint8_t bits = glyph[y];
for (uint32_t x = 0; x < 8; x++)
crash_pix(px + x, py + y, ((bits >> (7 - x)) & 1) ? VGA_PALETTE[cfg] : VGA_PALETTE[cbg]);
}
ccx++;
}
if (ccy >= (uint32_t)rows) crash_scroll();
}crash_pix is just put_pixel inlined, with no bounds-check overhead. crash_scroll shifts the entire framebuffer up by one text row using kmemmove on the raw pixel data and fills the vacated last row with the background color:
static void crash_scroll(void) {
uint32_t row_h = (uint32_t)fh + PADDING_Y;
size_t row_bytes = row_h * fb_pitch;
size_t total = (size_t)fb_h * fb_pitch;
kmemmove(fb, fb + row_bytes, total - row_bytes);
// ... fill last row with cbg ...
if (ccy > 0) ccy = (uint32_t)rows - 1;
}con_crash_clear fills the framebuffer with the background color and resets the cursor to the top-left:
void con_crash_clear(uint8_t bg) {
if (!fb) return;
cbg = bg & 0xF; ccx = 0; ccy = 0;
uint32_t color = VGA_PALETTE[cbg];
size_t total = (size_t)fb_h * fb_pitch;
if (fb_bpp == 32) {
uint32_t* p = (uint32_t*)fb;
for (size_t i = 0; i < total / 4; i++) p[i] = color;
}
// ...
}con_crash_puts iterates over a string calling crash_emit per byte. con_crash_printf formats into the static cbuf array using kvsnprintf (the same formatting function we wired up to serial in Chapter 4) and passes the result to con_crash_puts:
void con_crash_printf(const char* fmt, ...) {
if (!fb) return;
va_list args;
va_start(args, fmt);
kvsnprintf(cbuf, sizeof(cbuf), fmt, args);
va_end(args);
con_crash_puts(cbuf);
}The entire crash console depends on exactly three things: the fb pointer set by console_init, the glyph data baked into the kernel image at compile time and always accessible, and PHYSMAP_P2V having produced a valid virtual address. All three are true the moment console_init returns. Nothing else in the kernel needs to be functional.
This mirrors the philosophy behind the static fb_pd array from Chapter 5. A single 4KB BSS-allocated page directory was enough to map any framebuffer at any resolution, with no runtime allocation required.
Here, a handful of static integers and a direct framebuffer pointer are enough to render a crash screen that will always work, regardless of how broken the rest of the kernel is. When you are building the lowest layers of a system, independence from everything else is worth more than elegance.
Note
Notice what the crash console does not attempt to do: it does not acquire any spinlock before writing. spinlock_try_acquire exists precisely because a panic might fire while the interrupted code holds the console's lock. The crash console bypasses all of that by never touching the lock in the first place. This is correct because by the time a panic fires, we have called intr_off() as the very first thing, so no other actor on this CPU can interfere.
With the crash console available, we can build the panic subsystem. The API surface is small and deliberate:
void panic(const char* message);
void panicf(const char* fmt, ...);
void panic_c(const char* message, cpu_context_t* context);
void panicf_c(cpu_context_t* context, const char* fmt, ...);panic and panicf are for calling without CPU context — from a failed assertion during initialization, for instance.
panic_c and panicf_c receive a full cpu_context_t from the interrupt dispatcher, enabling detailed register dumps on the crash screen.
PANIC_ASSERT in panic.h wraps panicf into a one-liner:
#define PANIC_ASSERT(condition) \
((condition) ? (void)0 : panicf("Assertion failed in %s, line %d\n[!] Condition: %s", __FILE__, __LINE__, #condition))All four variants funnel into panic_c:
void panic_c(const char* message, cpu_context_t* context)
{
intr_off();
panic_log(message, context);
con_crash_clear(CONSOLE_COLOR_RED);
// ... layout the crash screen ...
halt_system();
}The very first call is intr_off(). We established earlier that the crash console bypasses locks by relying on the absence of interrupts. intr_off() makes that assumption true. From this line forward, no interrupt will fire on this CPU. The panic handler runs to completion deterministically.
panic_log writes the reason and context to COM2 via serial before anything touches the framebuffer:
static void panic_log(const char* msg, cpu_context_t* ctx)
{
LOGF("\n*** KERNEL PANIC ***\n");
LOGF("REASON: %s\n", msg);
if (ctx)
LOGF("%s (#%lu) ERR=0x%lx RIP=0x%016lx\n",
exc_name(ctx->vector_number),
ctx->vector_number,
ctx->error_code,
ctx->iret_rip);
LOGF("********************\n");
}LOGF writes to COM2, which run.py redirects to debug.log on disk. This matters because if the crash screen itself somehow fails, the serial log will still contain the panic reason and the faulting instruction pointer. Serial has no dependencies beyond the UART port address we configured in Chapter 4. It is the last resort.
Then con_crash_clear(CONSOLE_COLOR_RED) paints the entire screen red and resets the cursor. Everything previously on screen is gone. If you need to know what the kernel was outputting before the crash, debug.log is the place to look.
The crash screen layout uses con_crash_puts and con_crash_printf to print the reason, the exception name, and if we have context, the detailed CPU state:
con_crash_printf("[+] Reason: %s\n", message);
if (context) {
con_crash_printf("[+] Exception: %s (#%lu)\n",
exc_name(context->vector_number),
context->vector_number);
con_crash_printf("[+] Error Code: 0x%04lx\n", context->error_code);
if (context->vector_number == INT_PAGE_FAULT) {
uint64_t cr2;
__asm__ volatile("mov %%cr2, %0" : "=r"(cr2));
con_crash_printf("[+] CR2 (fault addr): 0x%016lx\n", cr2);
con_crash_printf("[+] Access: %s Mode: %s Cause: %s\n",
(context->error_code & 0x02) ? "write" : "read",
(context->error_code & 0x04) ? "user" : "supervisor",
(context->error_code & 0x01) ? "protection" : "not-present");
// ...
}
con_crash_printf("\nInstruction Pointer:\n");
con_crash_printf(" RIP: 0x%016lx\n", context->iret_rip);
con_crash_printf(" CS: 0x%04lx\n", context->iret_cs);
con_crash_printf(" RSP: 0x%016lx\n", context->iret_rsp);
// ...
}Page faults get extra treatment because they are by far the most common kernel crash during early development.
The cr2 control register always holds the virtual address that triggered the fault, essentially the address that was unmapped. We read it directly with inline assembly because cpu_context_t does not save control registers.
The error code for a page fault is a bitfield: bit 0 distinguishes a protection violation from a page that simply was not mapped; bit 1 says whether the access was a read or a write; bit 2 says whether we were in supervisor or user mode; bit 4 flags that the fault came from an instruction fetch rather than a data access.
Displaying all of this together makes diagnosing the cause of a fault much less painful than staring at a raw hex error code.
halt_system closes everything out:
void halt_system(void) {
while (1) __asm__ volatile("hlt");
}hlt puts the CPU into a low-power sleep until the next interrupt. Since intr_off() was called at the start of panic_c, no interrupt will ever arrive. The machine is permanently stopped.
GatOS has been permanently locked down.
Here is the section of kmain.c that this chapter describes, with the reasoning made explicit:
// Serial: zero dependencies, works before everything else
serial_init_port(COM1_PORT);
serial_init_port(COM2_PORT);
QEMU_LOG("Kernel main reached, normal assembly boot succeeded", TOTAL_DBG);
// IDT before anything that could fault
idt_init();
QEMU_LOG("Initialized the IDT", TOTAL_DBG);
// Multiboot and physmap: covered in previous chapters
multiboot_parser_t multiboot = {0};
multiboot_init(&multiboot, mb_info, multiboot_buffer, sizeof(multiboot_buffer));
reserve_required_tablespace(&multiboot);
cleanup_kpt(0x0, get_kend(false));
build_physmap();
// Physmap exists, so PHYSMAP_P2V(fb_phys) is now a valid virtual address
// The crash console works the moment this returns
console_init(&multiboot);
QEMU_LOG("Initialized console", TOTAL_DBG);
// From here, panacking halts with a red screen and extra information, rather than a silent resetEach step makes the next step debuggable.
- Serial makes IDT initialization visible if it goes wrong.
- The physmap makes
console_initsafe to call. -
console_initmakespanicuseful for catching failures. - The
interrupt_dispatchercatches exceptions and interrupts. - Spinlocks protect every shared data structure from the moment we start using them.
None of these components is particularly complex in isolation. Together, they are the floor that everything else in the kernel stands on. Now that the floor is solid, we can start building upward.
I apologize if this chapter was a bit cumbersome to go through. It went through a lot of new concepts, so I had to make sure to cover everything. Next time, let's start setting up the memory subsystems, shall we?