Skip to content

v0.10.0

@ParkMyCar ParkMyCar tagged this 13 Jul 04:16
* perf: route heap construction through a cold alloc_copy core (fixes x86 inline regression)

#488 built the heap variant via a helper that returned a full HeapBuffer. Returning
that 24-byte aggregate across the outlined call boundary forced its result through
the stack, which made LLVM route the whole constructor -- including the hot inline
arm -- through a stack temporary and scatter-copy it into the return slot on x86:
misaligned 1+8+4+2+1 loads overlapping the 16-byte build stores, a store-to-load
forwarding stall on every construction. This regressed inline creation on x86 (only
ever wall-clock benched on aarch64, where the reload stays clean).

Replace HeapBuffer::build with HeapBuffer::alloc_copy, the single #[cold]
#[inline(never)] allocate+copy core that returns the two-word (ptr, capacity) pair
(passed back in registers, not a stack aggregate). new and alloc_copy_panic become
#[inline(always)] wrappers that assemble the HeapBuffer from that pair at the
callsite. So every construction path -- new_panic and the fallible Repr::new behind
ToCompactString/From/number formatting -- keeps its clean 3x movq / ldp+ldr copy,
lays the heap arm out cold, and only (ptr, capacity) ever crosses the boundary.

Verified: lib tests pass; Miri clean across inline/full-24/multibyte/heap lengths
plus clone/drop under strict provenance.

* perf: construct the inline buffer in registers on 64-bit little-endian

The overlapping-copy ladder's stores straddle 8-byte word boundaries, so
whole-word reloads of a freshly built repr fail store-to-load forwarding
on x86 (a ~12-cycle stall per blocked load; inline creation measured a
flat ~10ns for every length except 24). Build the buffer's three words in
registers instead: overlapping loads from the source merged with shifts,
so the value can stay in registers entirely. Inline creation drops to
~1.35ns on Sapphire Rapids; aarch64 lowers to ldp/lsr/csel/stp with no
stack traffic. copy_small now only backs the fallback path, so gate it.

* perf: clone heap strings through new_panic

clone_heap was #[inline(never)] and returned the 24-byte Repr by value,
forcing an sret return that routed even the plain-copy inline arm through
a scatter-copied stack temporary on x86 -- the same shape fixed for
construction by alloc_copy. new_panic's heap arm already is that cold
out-of-line core, so call it directly. clone small -53%, clone large -32%
on Sapphire Rapids.

* perf: drop from_string's by-value cold helpers

The #[cold] empty()/capacity_on_heap() helpers returned Result<Repr>
through an sret slot, inviting the hot buffer-steal arm into a shared
stack temporary. Inline the empty case (it is one compare) and route the
over-16MB-on-32-bit case through the cold alloc_copy core.

* perf: route with_capacity/with_additional through cold register-pair cores

Same protocol as alloc_copy: the out-of-line allocating core returns
(ptr, capacity) in registers and an #[inline(always)] wrapper packs the
HeapBuffer at the callsite, instead of returning the 24-byte buffer by
value across the cold call boundary.

* perf: build Default::default from const_new

const_new("") folds the empty string to a constant instead of running
the runtime constructor.

* cargo fmt
Assets 2
Loading