Skip to content

MathFormer v1.2.0

Choose a tag to compare

@JeremySu0818 JeremySu0818 released this 10 Feb 21:32
· 12 commits to main since this release

MathFormer v1.2.0

Added

  • Batch Inference: New batch_predict method on MathFormer class that batches multiple expressions into a single model forward pass, significantly improving throughput on both CPU and GPU.
  • Batch Raw Predict: New _batch_raw_predict method on MathFormerAPI that leverages batch inference for internal digit-level operations.
  • Parallel Multiplication: Multi-digit multiplication (_multi_mul) now computes partial products in parallel using ThreadPoolExecutor, with each row's single-digit multiplications batched into one model call.
  • Parallel Trial Division: _trial_division now tests all 10 candidate quotients (0–9) in parallel instead of sequentially, reducing division latency by up to 10x.
  • Parallel Batch Operations: batch_predict on MathFormerAPI now processes multiple expressions concurrently using multi-threading.
  • max_workers Parameter: New max_workers parameter on MathFormerAPI to control the number of threads used for parallel computation (defaults to min(32, os.cpu_count() + 4)).

Changed

  • _multi_mul Strategy: For multi-digit multiplications (≥2 digits in multiplier), partial product rows are now computed in parallel threads with batched model inference, then merged with a single-pass carry propagation.
  • _trial_division Strategy: All 10 candidate products (divisor × 0 through divisor × 9) are computed in parallel, then the correct quotient is selected from the pre-computed results.
  • batch_predict Implementation: Replaced sequential loop with ThreadPoolExecutor-based parallel execution while preserving result ordering.

Improved

  • Multiplication Performance: For an N-digit × M-digit multiplication, the M rows of partial products are now computed simultaneously instead of sequentially, achieving up to M× speedup on multi-core systems.
  • Division Performance: Each trial division step is up to 10× faster due to parallel candidate evaluation.
  • Batch Throughput: Multiple independent operations can now fully utilize all available CPU cores.

Backward Compatibility

  • All existing APIs remain unchanged; the optimizations are internal.
  • The new max_workers parameter is optional with a sensible default.
  • All existing integer and decimal operations produce identical results.