Skip to content

Both cores of an RP2350, running one scheduler

Choose a tag to compare

@domschl domschl released this 05 Sep 05:17
· 159 commits to main since this release

Both Hazard3 cores of an RP2350 now run one scheduler, and two real workloads use the second one. That is the headline, and phases 22 and 23 are what it took. But 0.14.0 is the first release since 0.13.1 in August, so it also carries the network stack, the identity store and the clock-precision work that landed in between — summarised under Also in 0.14.0 below.

Two cores, one scheduler

Phase 22 built the locking foundation first, deliberately, before a second hart existed to need it. Per-hart identity through tp and a hart_t record; a spinlock_t and a re-entrant yielding ylock_t built on RISC-V amoswap behind an arch seam; every hand-rolled lock in the tree converted and the rest audited; and the scheduler lock held across ctx_switch() and released on the incoming stack. The rule there was inverted in the plan as written and corrected against the tree — every shared structure a second hart could touch had to be protected before that hart was ever allowed to run kernel code, because waking it first converts a latent race into an active one on day one.

Phase 23 woke the core. Two harts pull from one ready queue; driver tasks are pinned to core 0 and what that pinning does not cover is written down rather than assumed; RP2350's core 1 is launched over the SIO FIFO and needs no separate preemption timer; and the isolation suite was re-run with both harts demonstrably mid-task at the instant of each fault, so PMP domains are enforced on a non-primary Hazard3 core and not merely believed to be.

smpstart join performs the launch. cat /proc/cpuinfo reports harts_online, and ps gains a Hart column showing what each task is pinned to.

What the second core is actually for

A second core that only makes ps longer is not worth the locking. Two workloads use it:

perft, RP2350    10450 -> 5322 ms    1.96x
chess, RP2350     9287 -> 5924 ms    1.60x
smptest           locked=80000 (want 80000), harts=2, zero lost updates

(perft 4 2) splits move generation across both cores and is checked against the published node counts, not against itself — so a parallelisation bug surfaces as a wrong number rather than as a faster wrong answer. (chess 2) is a Lazy SMP search over a lock-free transposition table, 1.60x to a fixed depth. (chess) and (chess 1) are the single-core engine, byte-for-byte unchanged.

Three things emulation could not have shown

This project's own history says a second QEMU hart is not the same claim as a second real Hazard3 core. It was right again:

  • A tight test-and-set spin starves the other core when both share one bus, and deadlocks the machine. Invisible under emulation, fatal on silicon.
  • Hazard3's per-core interrupt force array (meifa) is never cleared by this kernel — survivable for core 0, which boot_header.S brings up from reset, and fatal for core 1, which arrives from the bootrom.
  • A transposition table entry torn between two cores yields a move that is legal in the current position but belongs to another one. Nothing in the engine rejects it.

Phase 23's own §1 premise — that RP2350 boots both of its Hazard3 cores — was falsified on hardware and corrected where it was made.

Opt-in, and what it costs when you don't opt in

CONFIG_ENABLE_SMP is set by exactly two presets, rv64-smp and rp2350-smp. Every other persona boots on a single hart exactly as it did before, with the second-core code compiled out entirely, so a regression in secondary bring-up cannot reach a board that never asked for it. On RP2350 the launch stays an explicit shell command rather than happening at boot, because a board that boots is a board that can be reflashed — which cost two BOOTSEL recoveries to learn.

One cost is not opt-in and is stated rather than hidden: core 1's own 16 KB stack costs 4 pages on every RP2350 persona, SMP-enabled or not, because a linker script cannot see the generated config header.

Also in 0.14.0

Everything below shipped between 0.13.1 and this tag and has had no release of its own.

An IP stack of our own, over two different wires

ARP, IPv4, ICMP, UDP and a server-side TCP, written here rather than bought in silicon — about 2,100 lines under net/, sized for an RP2350 and developed against a packet-level QEMU peer before either piece of hardware was in hand, which is why the same code came up on both wires without a per-part IP path. Two frame sources feed it through one netif_t seam: a wired ENC28J60 (SPI, MAC-only, no closed firmware anywhere) and the CYW43439 radio on a Pico 2 W, joining WPA2 and carrying 9P over the air.

The phase began by cancelling a W5500, and the reasoning is in the README because it decides the roadmap: the distinction that matters is not blob size, it is what is left to implement. A part whose closed firmware ends at the MAC layer leaves the network to us; a part whose closed firmware ends at TCP does not. The corollary is worth stating too — the CYW43439 is strictly more work than the W5500 was, not less.

Above the stack: 9P over TCP on port 564 with the same authentication and grants a serial link uses, host/fuse-p9 mounting a board's whole namespace onto a Linux host over either wire, and an SNTP client so a board can set its own clock from the segment. What the stack deliberately does not do — no IPv6, no DHCP, no TCP options past MSS — is listed in the README, because an unstated limit gets credited as a feature.

An identity that belongs to the silicon

A node's identity, its device key, its peer grants, its address and its WLAN credential now survive a firmware reflash: on RP2350 the device UID is read from OTP CHIPID, and the 4 KB record lives in its own reserved flash sector. /flash0 became its own independently flashable segment for the same reason, and the OS image halved as a side effect. Two boards were checked to report two different UIDs, and a provisioned identity was verified to survive a UF2 reflash.

Grants turn authentication into authorization: each entry names a peer, its key, the one subtree it may attach at, and whether it is read-only — where before, any peer that proved it held some configured key received the entire exported namespace, including a directory that runs Lisp programs by design. The rule that shapes the record is that a value used to prove who a node is must never also be the value used to decide who else may attach; an earlier milestone conflated the two, and the split (p9_auth_own_key() vs p9_auth_key_for()) exists specifically to close that gap.

WLAN credentials are stored as the derived PSK, never the passphrase. wifi join with no arguments and netcfg read from the record, so a board brings its own network up after a power cut with nothing typed.

One verify item is deliberately still open and not hand-waved: an interrupted flash write leaving the store readable as corrupt has not been attempted.

A clock that knows how wrong it is (phase 24, in progress)

Not finished, but well past the interesting part. The DCF-77 receiver's delay is now measured — CONFIG_DCF77_DELAY_US = 37886, against a GPS module's own PPS wired to the board, which removes the network from the measurement entirely — rather than fitted. The clock is disciplined between syncs instead of stepped once a night, the discipline is measured against the pulse and not only against the network, and the board can serve NTP to the segment and refuses to when it does not know the time. The GPS is a transfer standard: attached for the calibration, removed afterwards, and nothing in the shipped appliance depends on it.

Fixes worth naming

  • Seven bugs in cc and ed, behind one failure that the suite could not see. The root of several was sizeof(buf) left behind on buffers that had moved to the heap — ed was silently destroying files on save, and cc was reading three bytes of any header.
  • The RTC: OSF is sticky, so it means unverified, not unusable; a DS1307 is not a DS3231 past the clock registers; and the U-mode driver task never cleared OSF, so the lamp never went out.
  • 9P: FAT32's own . and .. were leaking into directory reads, and p9srv was overflowing its stack in a way ps could not report.
  • A real double-dispatch race in net/tcp.c, found while wiring up the ENC28J60.

Verified

QEMU 343/343 across rv32-nommu, rv64-mmu and the two-hart rv64-smp target. On real RP2350 silicon: the 24/24 core hardware suite on both personas, 15 more for the wired gateway, 6 over the radio, and 3 against a GPS-disciplined reference clock.

Downloads

Prebuilt UF2s for the rp2350-chess, rp2350-smp and rp2350-clock personas are attached, with SHA256SUMS and their own README.md. Flash two files: the persona image and lugalos-0.14.0-flashfs.uf2, which is now its own flash segment. The rp2350-smp image is the only attached one that can use the second core, and it does so only after smpstart join.