You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Pure Python LLaMA Inference Engine: Introduced a brand-new llama.py module implementing the LLaMA architecture entirely from scratch using only Python standard library modules (math, struct, operator), with native safetensors weight loading and a dependency-free tokenizer.
Changed
Zero External Dependencies: Removed all heavy third-party dependencies (torch, transformers, safetensors) from requirements.txt and pyproject.toml; the package now runs on pure Python with no external runtime dependencies.
Internal Model Backend: Replaced the transformers.LlamaForCausalLM backend with the custom TinyLlama pure Python engine for all model inference.
Project Description: Updated the package description to "A transformer-based math library (Pure Python)" to reflect the new zero-dependency architecture.
Context Manager Usage: The MathFormerAPI context manager and __enter__/__exit__ have been removed; users should now use the MathFormer class directly as a context manager for single-model lifecycle management.
Removed
MathFormerAPI.batch_predict(): Removed the batch prediction method and its underlying ThreadPoolExecutor-based concurrency from MathFormerAPI.
MathFormerAPI.get_model_info(): Removed the model introspection method that previously exposed path, load state, and device information.
MathFormerAPI.__enter__ / __exit__: Removed the context manager protocol from MathFormerAPI.
Improved
Model Architecture: Compressed model configurations — reduced hidden_size from 16 to 8, intermediate_size from 64 to 32, num_attention_heads from 8 to 2, and num_hidden_layers from 2 to 1 — further shrinking the combined artifact size of all four models from 154,560 bytes to 28,288 bytes (approx. 150.9 KB to 27.6 KB).
Documentation: Added comprehensive reStructuredText-style docstrings (:param, :type, :return, :rtype) to all public functions and classes across __init__.py, api.py, tokenizer.py, and llama.py.
README: Corrected example outputs, fixed lazy_load parameter usage, and updated the context manager section to demonstrate the MathFormer class instead of MathFormerAPI.
Attention Optimization: Fused QKV projections and streamlined attention score computation in the pure Python engine, using pre-bound operators and __slots__ on all model classes for reduced memory overhead.
Backward Compatibility
Breaking: MathFormerAPI no longer supports batch_predict(), get_model_info(), or the context manager protocol (with MathFormerAPI() as api). Migrate batch workloads to explicit loops and use MathFormer for context-managed single-model usage.
Public API functions (add, sub, mul, div, calculate, unload_models) remain fully compatible with the previous version.