Skip to content

ferrodec 1.11.0 — perf pass

Choose a tag to compare

@ixmatus ixmatus released this 08 May 05:12
· 749 commits to main since this release
v1.11.0
807de38

~17 % aggregate speedup across the core_ops bench suite vs 1.10.1, with the headline operations seeing 23–27 % wall-time reductions. No public API change.

Performance (cumulative vs 1.10.1, Apple Silicon, rustc 1.95.0, release with thin LTO + 1 codegen-unit)

Bench Before After Δ
add 7.98 µs 5.79 µs −27.5 %
sub 39.92 µs 30.53 µs −23.5 %
mul 42.45 µs 31.36 µs −26.1 %
div 50.69 µs 44.78 µs −11.7 %
fma 488 µs 415 µs −14.9 %
sub_alignment_heavy 7.98 µs 5.80 µs −27.3 %
mul_full_precision 6.60 µs 5.23 µs −20.8 %
parse_str 3.68 µs 2.91 µs −20.9 %
from_i128 2.98 µs 2.24 µs −24.8 %

Load-bearing optimizations

  • round_and_pack_finite now caches decimal_digit_count once per call instead of recomputing it 3× across the rounding / overflow-check / preferred-quantum branches. This single change is the bulk of the aggregate uplift (commit 15a7b98).
  • U256::mul_pow10 looks up 10^k from a precomputed [u128; 39] table instead of running an iterative mul10 loop. Hot consumers: alignment shifts in addsub, the rounding pipeline's overflow-renormalize step, the up-renormalize in finalize_finite (commit a53ddb4).
  • A third commit (84e4598) unified two duplicated digit-extraction loops in the rounding path.

Audit log

The full per-candidate audit — including three optimization candidates that were tested and reverted as no-op or noise-floor — is in docs/decisions/0008-perf-results.md. The pre-pass baseline lives in docs/decisions/0007-perf-baseline.md.

This release also introduces the ADR audit log under docs/decisions/. Eight ADRs (0001–0008) backfill the design log: BID over DPD, per-op status threading, method-only API, the Verus pilot outcome, the will-not-fix non-IEEE rounding directives, and the perf pass results pair.

Bench coverage additions

benches/core_ops.rs gains alignment-heavy / full-precision / magnitude-extreme variants. New benches/comparison.rs covers partial_cmp / total_cmp / compare_total_magnitude / 64-element sort. benches/conversions.rs adds from_i32 / from_u32 / from_u64 / to_i32 / to_u64 / to_u128. Permanent regression-watching shapes.

Verification

  • 425 lib tests pass (--features=transcendentals,binary-float,serde,ops,num-traits).
  • 63/63 Kani harnesses verify in ~2 minutes total.
  • decTest conformance: 8 622 / 0 / 99 — unchanged from 1.10.1.

🤖 Generated with Claude Code