Skip to content

Releases: itsoltech/basal-rs

basal-rs v0.1.9

Choose a tag to compare

@github-actions github-actions released this 10 Oct 17:17
e9444bb

What's Changed

  • basal benchmark now skips whole requests whose full prompts exceed the selected model's context limit, allowing mini to run the remaining supported workload instead of stopping during planning.
  • Filtering uses actual prompt validation, so shorter requests in the same document-size class remain eligible. It applies to embedded workloads and --requests replay files. Other planning errors still stop the benchmark; an empty workload after filtering produces an explicit error.
  • JSON reports include source and skipped request counts plus skipped IDs, classes and error details. Console summaries show skipped counts. Timing and throughput statistics cover only retained requests.
  • Documented filtering and the need to use a common supported request set when comparing models.

Validation

Local formatting, workspace Clippy, Rustdoc and dependency checks passed. GPU inference and a full benchmark run were not performed for this fix; static checks do not establish runtime performance or model parity.

Full Changelog: v0.1.8...v0.1.9

basal-rs v0.1.8

Choose a tag to compare

@github-actions github-actions released this 10 Oct 14:55
24ae958

What's Changed

  • CUDA packages now embed 27 batch-invariant f16 GEMM tables for 10 GPU names, including 24 newly generated tables. Matching GPU/model/cuBLASLt 12.9.1 configurations can use them without searching for algorithms at startup.
  • Mini, 4.5B and max are covered on H100 PCIe, RTX 6000 Ada Generation, L40S, RTX A6000, A100-SXM4-80GB, L40 and RTX PRO 6000 Blackwell Server Edition.
  • Mini and 4.5B are covered on L4, A10 and RTX A5000.
  • Added a cloud runner and reports documenting table generation, batch invariance, FP32 comparisons and benchmarks.

Validation

All 24 new tables produced byte-identical exports across single/tree/budget batching on the checked cases. Each was compared with upstream FP32 on 44 examples and 900 questions. FP32 decision agreement is 44/44 for all models; on the larger set it is 899/900 for mini and 4.5B, and 900/900 for max. This does not imply bitwise equality with FP32.

Reports: initial campaign, additional GPUs.

Full Changelog: v0.1.7...v0.1.8

basal-rs v0.1.7

Choose a tag to compare

@github-actions github-actions released this 10 Oct 07:47
85ca875

What's Changed

  • ci: bump the actions group across 1 directory with 6 updates by @dependabot[bot] in #1
  • perf(cuda): optimize Hopper attention with measured SW128 layout by @nixuuu in #3
  • feat(cli): add a decision client for shell scripts by @nixuuu in #4
  • feat(cli): add deterministic hardware benchmark by @nixuuu in #5

New Contributors

Full Changelog: v0.1.6...v0.1.7

basal-rs v0.1.6

Choose a tag to compare

@github-actions github-actions released this 09 Oct 13:50
1614bda

Full Changelog: v0.1.5...v0.1.6

basal-rs v0.1.5

Choose a tag to compare

@github-actions github-actions released this 09 Oct 09:55
190137e

What's Changed

  • feat(cli): add opt-in Prometheus observability for model serving by @nixuuu in #2

New Contributors

  • @nixuuu made their first contribution in #2

Full Changelog: v0.1.4...v0.1.5

basal-rs v0.1.4

Choose a tag to compare

@github-actions github-actions released this 08 Oct 10:51
db87ea0

Full Changelog: v0.1.3...v0.1.4

basal-rs v0.1.3

Choose a tag to compare

@github-actions github-actions released this 07 Oct 19:20
e140be9

Full Changelog: v0.1.2...v0.1.3

basal-rs v0.1.2

Choose a tag to compare

@github-actions github-actions released this 07 Oct 18:21
8f0c5c4

Full Changelog: v0.1.1...v0.1.2

basal-rs v0.1.1

Choose a tag to compare

@github-actions github-actions released this 07 Oct 17:46
0f0cf9e

Full Changelog: v0.1.0...v0.1.1

basal-rs v0.1.0

Choose a tag to compare

@github-actions github-actions released this 07 Oct 17:11
50d43b4