Repository navigation
Releases: itsoltech/basal-rs
Release list
basal-rs v0.1.9
What's Changed
basal benchmarknow skips whole requests whose full prompts exceed the selected model's context limit, allowing mini to run the remaining supported workload instead of stopping during planning.- Filtering uses actual prompt validation, so shorter requests in the same document-size class remain eligible. It applies to embedded workloads and
--requestsreplay files. Other planning errors still stop the benchmark; an empty workload after filtering produces an explicit error. - JSON reports include source and skipped request counts plus skipped IDs, classes and error details. Console summaries show skipped counts. Timing and throughput statistics cover only retained requests.
- Documented filtering and the need to use a common supported request set when comparing models.
Validation
Local formatting, workspace Clippy, Rustdoc and dependency checks passed. GPU inference and a full benchmark run were not performed for this fix; static checks do not establish runtime performance or model parity.
Full Changelog: v0.1.8...v0.1.9
basal-rs v0.1.8
What's Changed
- CUDA packages now embed 27 batch-invariant f16 GEMM tables for 10 GPU names, including 24 newly generated tables. Matching GPU/model/cuBLASLt 12.9.1 configurations can use them without searching for algorithms at startup.
- Mini, 4.5B and max are covered on H100 PCIe, RTX 6000 Ada Generation, L40S, RTX A6000, A100-SXM4-80GB, L40 and RTX PRO 6000 Blackwell Server Edition.
- Mini and 4.5B are covered on L4, A10 and RTX A5000.
- Added a cloud runner and reports documenting table generation, batch invariance, FP32 comparisons and benchmarks.
Validation
All 24 new tables produced byte-identical exports across single/tree/budget batching on the checked cases. Each was compared with upstream FP32 on 44 examples and 900 questions. FP32 decision agreement is 44/44 for all models; on the larger set it is 899/900 for mini and 4.5B, and 900/900 for max. This does not imply bitwise equality with FP32.
Reports: initial campaign, additional GPUs.
Full Changelog: v0.1.7...v0.1.8
basal-rs v0.1.7
What's Changed
- ci: bump the actions group across 1 directory with 6 updates by @dependabot[bot] in #1
- perf(cuda): optimize Hopper attention with measured SW128 layout by @nixuuu in #3
- feat(cli): add a decision client for shell scripts by @nixuuu in #4
- feat(cli): add deterministic hardware benchmark by @nixuuu in #5
New Contributors
- @dependabot[bot] made their first contribution in #1
Full Changelog: v0.1.6...v0.1.7
basal-rs v0.1.6
Full Changelog: v0.1.5...v0.1.6
basal-rs v0.1.5
basal-rs v0.1.4
Full Changelog: v0.1.3...v0.1.4
basal-rs v0.1.3
Full Changelog: v0.1.2...v0.1.3
basal-rs v0.1.2
Full Changelog: v0.1.1...v0.1.2
basal-rs v0.1.1
Full Changelog: v0.1.0...v0.1.1
basal-rs v0.1.0
Full Changelog: https://github.com/itsoltech/basal-rs/commits/v0.1.0