ferrodec 1.11.0 — perf pass
~17 % aggregate speedup across the core_ops bench suite vs 1.10.1, with the headline operations seeing 23–27 % wall-time reductions. No public API change.
Performance (cumulative vs 1.10.1, Apple Silicon, rustc 1.95.0, release with thin LTO + 1 codegen-unit)
| Bench | Before | After | Δ |
|---|---|---|---|
add |
7.98 µs | 5.79 µs | −27.5 % |
sub |
39.92 µs | 30.53 µs | −23.5 % |
mul |
42.45 µs | 31.36 µs | −26.1 % |
div |
50.69 µs | 44.78 µs | −11.7 % |
fma |
488 µs | 415 µs | −14.9 % |
sub_alignment_heavy |
7.98 µs | 5.80 µs | −27.3 % |
mul_full_precision |
6.60 µs | 5.23 µs | −20.8 % |
parse_str |
3.68 µs | 2.91 µs | −20.9 % |
from_i128 |
2.98 µs | 2.24 µs | −24.8 % |
Load-bearing optimizations
round_and_pack_finitenow cachesdecimal_digit_countonce per call instead of recomputing it 3× across the rounding / overflow-check / preferred-quantum branches. This single change is the bulk of the aggregate uplift (commit 15a7b98).U256::mul_pow10looks up10^kfrom a precomputed[u128; 39]table instead of running an iterativemul10loop. Hot consumers: alignment shifts inaddsub, the rounding pipeline's overflow-renormalize step, the up-renormalize infinalize_finite(commit a53ddb4).- A third commit (84e4598) unified two duplicated digit-extraction loops in the rounding path.
Audit log
The full per-candidate audit — including three optimization candidates that were tested and reverted as no-op or noise-floor — is in docs/decisions/0008-perf-results.md. The pre-pass baseline lives in docs/decisions/0007-perf-baseline.md.
This release also introduces the ADR audit log under docs/decisions/. Eight ADRs (0001–0008) backfill the design log: BID over DPD, per-op status threading, method-only API, the Verus pilot outcome, the will-not-fix non-IEEE rounding directives, and the perf pass results pair.
Bench coverage additions
benches/core_ops.rs gains alignment-heavy / full-precision / magnitude-extreme variants. New benches/comparison.rs covers partial_cmp / total_cmp / compare_total_magnitude / 64-element sort. benches/conversions.rs adds from_i32 / from_u32 / from_u64 / to_i32 / to_u64 / to_u128. Permanent regression-watching shapes.
Verification
- 425 lib tests pass (
--features=transcendentals,binary-float,serde,ops,num-traits). - 63/63 Kani harnesses verify in ~2 minutes total.
- decTest conformance: 8 622 / 0 / 99 — unchanged from 1.10.1.
🤖 Generated with Claude Code