You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Batch Inference: New batch_predict method on MathFormer class that batches multiple expressions into a single model forward pass, significantly improving throughput on both CPU and GPU.
Batch Raw Predict: New _batch_raw_predict method on MathFormerAPI that leverages batch inference for internal digit-level operations.
Parallel Multiplication: Multi-digit multiplication (_multi_mul) now computes partial products in parallel using ThreadPoolExecutor, with each row's single-digit multiplications batched into one model call.
Parallel Trial Division: _trial_division now tests all 10 candidate quotients (0–9) in parallel instead of sequentially, reducing division latency by up to 10x.
Parallel Batch Operations: batch_predict on MathFormerAPI now processes multiple expressions concurrently using multi-threading.
max_workers Parameter: New max_workers parameter on MathFormerAPI to control the number of threads used for parallel computation (defaults to min(32, os.cpu_count() + 4)).
Changed
_multi_mul Strategy: For multi-digit multiplications (≥2 digits in multiplier), partial product rows are now computed in parallel threads with batched model inference, then merged with a single-pass carry propagation.
_trial_division Strategy: All 10 candidate products (divisor × 0 through divisor × 9) are computed in parallel, then the correct quotient is selected from the pre-computed results.
batch_predict Implementation: Replaced sequential loop with ThreadPoolExecutor-based parallel execution while preserving result ordering.
Improved
Multiplication Performance: For an N-digit × M-digit multiplication, the M rows of partial products are now computed simultaneously instead of sequentially, achieving up to M× speedup on multi-core systems.
Division Performance: Each trial division step is up to 10× faster due to parallel candidate evaluation.
Batch Throughput: Multiple independent operations can now fully utilize all available CPU cores.
Backward Compatibility
All existing APIs remain unchanged; the optimizations are internal.
The new max_workers parameter is optional with a sensible default.
All existing integer and decimal operations produce identical results.