Releases: luxine/CalcKernel
Release list
CalcKernel v0.15.2
Changelog
All notable user-visible changes to CalcKernel are recorded here.
0.15.2 — WebAssembly Performance and Correctness Update
This patch release improves eligible WebAssembly kernels while preserving the
existing language, CLI, public ABI, strict floating-point behavior, and
baseline/simd128 feature boundaries. Unsupported cases retain their safe
scalar paths.
- Adds independently verified vector and scalar lowering for eligible matrix, stencil, and
other numerical kernels. The optimizations preserve input order, bounds and
alias guards, integer wraparound, and strict floating-point evaluation. - Development-candidate evidence from the strict 11-kernel suite records two
independent seven-round runs. Each run contains 66 backend/profile/workload
results with all correctness checks passing and all profile comparisons
eligible. Against the faster equivalent Clang/Rust WebAssembly
implementation, baseline throughput geometric means were 1.117× and 1.118×
(slowest cases 0.920× and 0.914×); SIMD128 means were 1.016× and 1.008×
(slowest cases 0.923× and 0.920×). The reports declare a clean source checkout at commit
08292f18b6374f9636bae5ab07fbfcb8e9e30091andckcbinary SHA-256
5db7c4db4146dda00020941f1047a1ade794741374e9446af4b2669de8979e3b, and
report the candidate asckc 0.15.1; report SHA-256 values are
ed56abd762d3d02404ae92b09ce213f8e8ef95f378020dea0f4b7355f302351eand
4384dc2262a4a967a40d1b6f04096235def2b51408c12ae9c363c7946243378b. The
binary-to-source relation is not independently verifiable; the exact binary
hash is authoritative. These are development-candidate results, not measurements of the v0.15.2 release
archives. See the performance guide
for methodology and currently published measurements. - The hot-call comparison excludes compilation, process startup, fixture I/O,
WebAssembly instantiation, memory growth, and report generation. In the same
candidate runs, baseline and SIMD128 modules were 28,325 and 34,941 bytes;
measured module compilation was about 0.16–0.18 ms. First-call costs remain
workload-dependent (for example, baseline matmul was about 43.95 ms and
SIMD128 matmul about 20.58 ms). These cold costs and module sizes are reported
separately from hot throughput and are not claimed as parity results.
0.15.1
The first published 0.15 release carries forward the compiler and WebAssembly
features from the immutable v0.15.0 tagged source snapshot. It also fixes the
release workflow's validation of annotated tags while retaining the exact
tagged source and all existing release gates.
- Includes the opt-in WebAssembly
simd128profile and O3 Bulk Memory paths
for eligible copies and byte fills. Unsupported cases retain scalar code. - Includes direct verified-KIR lowering to WAT and binary modules, while
preserving exported function signatures and caller-owned memory. The
ck.wasm.targetcustom-section metadata uses schema 2. - In local Node.js 24.14.0/V8 tests on Apple M5 Max, O3
simd128maps
measured 3.1–6.0× Node.js fori32and 2.2–2.8× forf64across
4,096–65,536 elements. These hot-kernel measurements exclude compilation,
instantiation, and data preparation; they describe one machine and runtime,
not general Wasm performance.
0.15.0 — Tagged source snapshot; not released
The immutable v0.15.0 tag records this source history. Its workflow did not
create an official GitHub Release or publish compiler archives or checksums;
the source is carried forward in 0.15.1.
- Added an opt-in WebAssembly
simd128profile. At O3, eligible contiguous
slice maps and modulo-2^32i32/u32sum and product reductions can use SIMD;
unsupported cases and scalar remainders retain their original behavior. - Enabled Bulk Memory in the default
baselineand opt-insimd128profiles.
O3 can use guardedmemory.copyandmemory.fillpaths for eligible copies
and byte fills, with the original scalar loop as fallback. - WebAssembly emission now lowers verified KIR directly to WAT and binary
modules, with structured control flow, direct binary encoding, and improved
loop-address lowering. - Preserved exported function signatures and caller-owned memory. The
ck.wasm.targetcustom-section metadata advances to schema 2 to identify the
selected Wasm feature profile. - In local Node.js 24.14.0/V8 tests on Apple M5 Max, O3
simd128maps
measured 3.1–6.0× Node.js fori32and 2.2–2.8× forf64across
4,096–65,536 elements. These
hot-kernel measurements exclude compilation, instantiation, and data
preparation; they describe one machine and runtime, not general Wasm
performance.
0.14.0
- Corrected the PGO-generation library flush status when its collection
directory becomes unwritable after build: it now reports directory failure
(43) instead of retrying sixteen false name collisions and reporting write
failure (44). Successful shards and genuine collision retries remain intact;
public Native C ABI 1, Runtime ABI 2, and profile wire formats do not change. - Retained 0.13.0 ordinary compilation, PGO, and multiversion behavior on the
same six native platforms. Offline Auto-Tuning is deferred:ckc tuneand
ckc build --tune-usefail explicitly without output side effects. This
version does not claim a new optimizer speedup or passing tuning gates.
0.13.0
- Added deterministic CK-owned
CKPART01shards andCKPROF01workload
profiles, with directory-safe collection, canonical merge/inspection, and the
transactionalckc pgo buildconvenience workflow. - Added non-proof profile analysis and independently checked O2 late machine
layout plus O3 guarded specialization, inlining, unrolling, SLP, and Loop SIMD
decisions. Profiles may affect profitability but never establish safety. - Added explicit
--cpu multiversionNative builds with one portable baseline,
bounded verified feature variants, a baseline-safe process-local detector,
stable public thunks, and executable/dynamic/static named-object assembly. - Added real library-generation workflows with the full-identity
ck_profile_flush_*control symbol. Final profile-use artifacts contain no
counters, profile paths, writer, or generation runtime. - Advanced the private LLVM bridge to ABI 4, KIR to v3, and the Native object
cache toCKCOBJ03key/manifest schema 4. Public Native C ABI 1 and Runtime
ABI 2 remain unchanged; 0.12 source and observable semantics remain accepted. - Added closed profile/target/dispatch/cache identities, corruption and mutation
tests, transactional multi-file output, and schema-8 performance and exact-SHA
release audits. - Auto-Tuning, indirect-call promotion, scalable KIR vectors, and adaptive JIT
PGO remained future work.
0.12.0 - Unreleased
- Added KIR v2 fixed-vector and mask instructions plus deterministic Native
KirTargetProfilecapability/cost data derived from pinned LLVM 22.1.8. - Added transactional, independently checked O3 specialization, controlled
full/partial unrolling, SLP, and Loop SIMD frontiers with monotonic analysis
budgets and stable optimization explanations. - Added unit-stride Loop SIMD for integer and strict element-wise f64 arithmetic,
supported casts, pure compare/select diamonds, splats, and contiguous memory;
strict f64 keeps per-element ordering and never enables fast math. - Added total runtime alias versioning with the unchanged scalar loop as the
fallback, scalar epilogues, and exact unchecked modular u32 add/multiply
reductions. Checked failures, effects, unsupported recurrences, scans, C,
and WebAssembly remain scalar. - Advanced the private LLVM bridge to ABI 3, KIR identity to
kir-v2, and the
Native object cache toCKCOBJ02key/manifest schema 3 while retaining public
Native C ABI 1, Runtime ABI 2, source syntax, diagnostics, and checked
first-error behavior. - Added KIR/pre-LLVM/object structural evidence, fixed-seed O0/O3 differential
coverage, mutation tests, target-feature containment, and schema-7 performance
gate inputs. PGO/multiversioning and Auto-Tuning remain future work.
0.11.0 - Unreleased
- Added explicit
unsafe fncontracts for affine range requirements,
multiple_of,noalias, alignment, and slice memory-effect ceilings. Unsafe
calls require anunsafe { ... }statement and executablemainremains safe. - Added deterministic
emit-kirinspection, verified facts/effect summaries,
proof-carrying guard elimination explanations, and opt-in Native contract
sanitization withCKR0007. - Replaced the former target-neutral MIR optimizer with one verified KIR
pipeline shared by C, WebAssembly, and Native LLVM. Semantic MIR and stable
emit-miroutput remain the source-order and first-error boundary. - Added scalar/path, region alias and Memory SSA, interprocedural effect, loop,
GVN/load-forwarding/dead-store, LICM, and evidence-audited backend facts. - Kept Native C ABI 1 while advancing the private LLVM bridge and runtime ABI
to 2 and using the KIR v1 native cache/code-generation identity. - Added fixed-seed differential and mutation suites, pre-LLVM fact audits, and
performance gates against both pinned Clang and exact CalcKernel 0.10.
0.10.0 - 2026-08-27
- Added parameterless internal
main,ckc run, and Native executable output. - Added deterministic Native output builtins for signed/unsigned integers,
f64, booleans, and newline; library, C, and WASM roots reject reachable print. - Replaced product Clang subprocesses with pinned LLVM 22.1.8 structural code
generation, ORC execution, archive writing, and in-process LLD linking. - Expanded
ckc build --kindto executable, dynamic, static, and object outputs;
dynamic remains the default andbuild-llvmis a deprecated compatibility alias. - Unified Native object/static/dynamic exports under one generated-header Native
C ABI with ...
CalcKernel v0.15.1
Changelog
All notable user-visible changes to CalcKernel are recorded here.
0.15.1
The first published 0.15 release carries forward the compiler and WebAssembly
features from the immutable v0.15.0 tagged source snapshot. It also fixes the
release workflow's validation of annotated tags while retaining the exact
tagged source and all existing release gates.
- Includes the opt-in WebAssembly
simd128profile and O3 Bulk Memory paths
for eligible copies and byte fills. Unsupported cases retain scalar code. - Includes direct verified-KIR lowering to WAT and binary modules, while
preserving exported function signatures and caller-owned memory. The
ck.wasm.targetcustom-section metadata uses schema 2. - In local Node.js 24.14.0/V8 tests on Apple M5 Max, O3
simd128maps
measured 3.1–6.0× Node.js fori32and 2.2–2.8× forf64across
4,096–65,536 elements. These hot-kernel measurements exclude compilation,
instantiation, and data preparation; they describe one machine and runtime,
not general Wasm performance.
0.15.0 — Tagged source snapshot; not released
The immutable v0.15.0 tag records this source history. Its workflow did not
create an official GitHub Release or publish compiler archives or checksums;
the source is carried forward in 0.15.1.
- Added an opt-in WebAssembly
simd128profile. At O3, eligible contiguous
slice maps and modulo-2^32i32/u32sum and product reductions can use SIMD;
unsupported cases and scalar remainders retain their original behavior. - Enabled Bulk Memory in the default
baselineand opt-insimd128profiles.
O3 can use guardedmemory.copyandmemory.fillpaths for eligible copies
and byte fills, with the original scalar loop as fallback. - WebAssembly emission now lowers verified KIR directly to WAT and binary
modules, with structured control flow, direct binary encoding, and improved
loop-address lowering. - Preserved exported function signatures and caller-owned memory. The
ck.wasm.targetcustom-section metadata advances to schema 2 to identify the
selected Wasm feature profile. - In local Node.js 24.14.0/V8 tests on Apple M5 Max, O3
simd128maps
measured 3.1–6.0× Node.js fori32and 2.2–2.8× forf64across
4,096–65,536 elements. These
hot-kernel measurements exclude compilation, instantiation, and data
preparation; they describe one machine and runtime, not general Wasm
performance.
0.14.0
- Corrected the PGO-generation library flush status when its collection
directory becomes unwritable after build: it now reports directory failure
(43) instead of retrying sixteen false name collisions and reporting write
failure (44). Successful shards and genuine collision retries remain intact;
public Native C ABI 1, Runtime ABI 2, and profile wire formats do not change. - Retained 0.13.0 ordinary compilation, PGO, and multiversion behavior on the
same six native platforms. Offline Auto-Tuning is deferred:ckc tuneand
ckc build --tune-usefail explicitly without output side effects. This
version does not claim a new optimizer speedup or passing tuning gates.
0.13.0
- Added deterministic CK-owned
CKPART01shards andCKPROF01workload
profiles, with directory-safe collection, canonical merge/inspection, and the
transactionalckc pgo buildconvenience workflow. - Added non-proof profile analysis and independently checked O2 late machine
layout plus O3 guarded specialization, inlining, unrolling, SLP, and Loop SIMD
decisions. Profiles may affect profitability but never establish safety. - Added explicit
--cpu multiversionNative builds with one portable baseline,
bounded verified feature variants, a baseline-safe process-local detector,
stable public thunks, and executable/dynamic/static named-object assembly. - Added real library-generation workflows with the full-identity
ck_profile_flush_*control symbol. Final profile-use artifacts contain no
counters, profile paths, writer, or generation runtime. - Advanced the private LLVM bridge to ABI 4, KIR to v3, and the Native object
cache toCKCOBJ03key/manifest schema 4. Public Native C ABI 1 and Runtime
ABI 2 remain unchanged; 0.12 source and observable semantics remain accepted. - Added closed profile/target/dispatch/cache identities, corruption and mutation
tests, transactional multi-file output, and schema-8 performance and exact-SHA
release audits. - Auto-Tuning, indirect-call promotion, scalable KIR vectors, and adaptive JIT
PGO remained future work.
0.12.0 - Unreleased
- Added KIR v2 fixed-vector and mask instructions plus deterministic Native
KirTargetProfilecapability/cost data derived from pinned LLVM 22.1.8. - Added transactional, independently checked O3 specialization, controlled
full/partial unrolling, SLP, and Loop SIMD frontiers with monotonic analysis
budgets and stable optimization explanations. - Added unit-stride Loop SIMD for integer and strict element-wise f64 arithmetic,
supported casts, pure compare/select diamonds, splats, and contiguous memory;
strict f64 keeps per-element ordering and never enables fast math. - Added total runtime alias versioning with the unchanged scalar loop as the
fallback, scalar epilogues, and exact unchecked modular u32 add/multiply
reductions. Checked failures, effects, unsupported recurrences, scans, C,
and WebAssembly remain scalar. - Advanced the private LLVM bridge to ABI 3, KIR identity to
kir-v2, and the
Native object cache toCKCOBJ02key/manifest schema 3 while retaining public
Native C ABI 1, Runtime ABI 2, source syntax, diagnostics, and checked
first-error behavior. - Added KIR/pre-LLVM/object structural evidence, fixed-seed O0/O3 differential
coverage, mutation tests, target-feature containment, and schema-7 performance
gate inputs. PGO/multiversioning and Auto-Tuning remain future work.
0.11.0 - Unreleased
- Added explicit
unsafe fncontracts for affine range requirements,
multiple_of,noalias, alignment, and slice memory-effect ceilings. Unsafe
calls require anunsafe { ... }statement and executablemainremains safe. - Added deterministic
emit-kirinspection, verified facts/effect summaries,
proof-carrying guard elimination explanations, and opt-in Native contract
sanitization withCKR0007. - Replaced the former target-neutral MIR optimizer with one verified KIR
pipeline shared by C, WebAssembly, and Native LLVM. Semantic MIR and stable
emit-miroutput remain the source-order and first-error boundary. - Added scalar/path, region alias and Memory SSA, interprocedural effect, loop,
GVN/load-forwarding/dead-store, LICM, and evidence-audited backend facts. - Kept Native C ABI 1 while advancing the private LLVM bridge and runtime ABI
to 2 and using the KIR v1 native cache/code-generation identity. - Added fixed-seed differential and mutation suites, pre-LLVM fact audits, and
performance gates against both pinned Clang and exact CalcKernel 0.10.
0.10.0 - 2026-08-27
- Added parameterless internal
main,ckc run, and Native executable output. - Added deterministic Native output builtins for signed/unsigned integers,
f64, booleans, and newline; library, C, and WASM roots reject reachable print. - Replaced product Clang subprocesses with pinned LLVM 22.1.8 structural code
generation, ORC execution, archive writing, and in-process LLD linking. - Expanded
ckc build --kindto executable, dynamic, static, and object outputs;
dynamic remains the default andbuild-llvmis a deprecated compatibility alias. - Unified Native object/static/dynamic exports under one generated-header Native
C ABI with target ABI classification and checked status thunks. - Added Native checked overflow and slice bounds, preserving the C
CK_Status
meanings and first-error order. - Added an isolated
runchild, secure persistent object cache, fixed runtime
diagnostics/statuses, eager symbol resolution, and audited JIT page permissions. - Added checked/unchecked C-oracle performance gates and six-host functional,
artifact, dependency, provenance, and immutable release gates. - Reserved
mainand the seven print builtin names, limited Native target output
to the host, retired the standalone LLVM exported-shape promise, and kept
emit-csource-only. See the compatibility policy for migration guidance.
0.9.0 - 2026-08-26
- Added
breakandcontinuefor structured control insidewhileloops. - Added explicit
voidprocedures, emptyreturn;, and procedure-call statements. - Added non-owning
slice<T>values,slice(data, len), indexing,.data/.len,
and half-open sub-slices written asitems[start..end]. - Added optional checked slice bounds to the C backend through
--bounds checked;
unchecked bounds remain the default, while WASM and LLVM reject checked bounds. - Stabilized native C, WebAssembly, and LLVM output paths and their V0.9 ABIs.
- Reorganized the repository around durable compiler, contract, example, benchmark,
and test responsibilities without changing the compiler's public behavior. - Froze the V0.9 compatibility boundary: patch releases in the
0.9.xline preserve
accepted source, diagnostic identifiers, CLI behavior, textual MIR, and documented
ABI contracts. A later0.10.0may make documented breaking changes with migration
guidance; long-term compatibility begins with a future1.0.0release. - Added signed-off native
ckcrelease archives and SHA-256 checksums for macOS,
Linux, and Windows on both arm64 and x64.
CalcKernel v0.14.0
Changelog
All notable user-visible changes to CalcKernel are recorded here.
0.14.0
- Corrected the PGO-generation library flush status when its collection
directory becomes unwritable after build: it now reports directory failure
(43) instead of retrying sixteen false name collisions and reporting write
failure (44). Successful shards and genuine collision retries remain intact;
public Native C ABI 1, Runtime ABI 2, and profile wire formats do not change. - Retained 0.13.0 ordinary compilation, PGO, and multiversion behavior on the
same six native platforms. Offline Auto-Tuning is deferred:ckc tuneand
ckc build --tune-usefail explicitly without output side effects. This
version does not claim a new optimizer speedup or passing tuning gates.
0.13.0
- Added deterministic CK-owned
CKPART01shards andCKPROF01workload
profiles, with directory-safe collection, canonical merge/inspection, and the
transactionalckc pgo buildconvenience workflow. - Added non-proof profile analysis and independently checked O2 late machine
layout plus O3 guarded specialization, inlining, unrolling, SLP, and Loop SIMD
decisions. Profiles may affect profitability but never establish safety. - Added explicit
--cpu multiversionNative builds with one portable baseline,
bounded verified feature variants, a baseline-safe process-local detector,
stable public thunks, and executable/dynamic/static named-object assembly. - Added real library-generation workflows with the full-identity
ck_profile_flush_*control symbol. Final profile-use artifacts contain no
counters, profile paths, writer, or generation runtime. - Advanced the private LLVM bridge to ABI 4, KIR to v3, and the Native object
cache toCKCOBJ03key/manifest schema 4. Public Native C ABI 1 and Runtime
ABI 2 remain unchanged; 0.12 source and observable semantics remain accepted. - Added closed profile/target/dispatch/cache identities, corruption and mutation
tests, transactional multi-file output, and schema-8 performance and exact-SHA
release audits. - Auto-Tuning, indirect-call promotion, scalable KIR vectors, and adaptive JIT
PGO remained future work.
0.12.0 - Unreleased
- Added KIR v2 fixed-vector and mask instructions plus deterministic Native
KirTargetProfilecapability/cost data derived from pinned LLVM 22.1.8. - Added transactional, independently checked O3 specialization, controlled
full/partial unrolling, SLP, and Loop SIMD frontiers with monotonic analysis
budgets and stable optimization explanations. - Added unit-stride Loop SIMD for integer and strict element-wise f64 arithmetic,
supported casts, pure compare/select diamonds, splats, and contiguous memory;
strict f64 keeps per-element ordering and never enables fast math. - Added total runtime alias versioning with the unchanged scalar loop as the
fallback, scalar epilogues, and exact unchecked modular u32 add/multiply
reductions. Checked failures, effects, unsupported recurrences, scans, C,
and WebAssembly remain scalar. - Advanced the private LLVM bridge to ABI 3, KIR identity to
kir-v2, and the
Native object cache toCKCOBJ02key/manifest schema 3 while retaining public
Native C ABI 1, Runtime ABI 2, source syntax, diagnostics, and checked
first-error behavior. - Added KIR/pre-LLVM/object structural evidence, fixed-seed O0/O3 differential
coverage, mutation tests, target-feature containment, and schema-7 performance
gate inputs. PGO/multiversioning and Auto-Tuning remain future work.
0.11.0 - Unreleased
- Added explicit
unsafe fncontracts for affine range requirements,
multiple_of,noalias, alignment, and slice memory-effect ceilings. Unsafe
calls require anunsafe { ... }statement and executablemainremains safe. - Added deterministic
emit-kirinspection, verified facts/effect summaries,
proof-carrying guard elimination explanations, and opt-in Native contract
sanitization withCKR0007. - Replaced the former target-neutral MIR optimizer with one verified KIR
pipeline shared by C, WebAssembly, and Native LLVM. Semantic MIR and stable
emit-miroutput remain the source-order and first-error boundary. - Added scalar/path, region alias and Memory SSA, interprocedural effect, loop,
GVN/load-forwarding/dead-store, LICM, and evidence-audited backend facts. - Kept Native C ABI 1 while advancing the private LLVM bridge and runtime ABI
to 2 and using the KIR v1 native cache/code-generation identity. - Added fixed-seed differential and mutation suites, pre-LLVM fact audits, and
performance gates against both pinned Clang and exact CalcKernel 0.10.
0.10.0 - 2026-08-27
- Added parameterless internal
main,ckc run, and Native executable output. - Added deterministic Native output builtins for signed/unsigned integers,
f64, booleans, and newline; library, C, and WASM roots reject reachable print. - Replaced product Clang subprocesses with pinned LLVM 22.1.8 structural code
generation, ORC execution, archive writing, and in-process LLD linking. - Expanded
ckc build --kindto executable, dynamic, static, and object outputs;
dynamic remains the default andbuild-llvmis a deprecated compatibility alias. - Unified Native object/static/dynamic exports under one generated-header Native
C ABI with target ABI classification and checked status thunks. - Added Native checked overflow and slice bounds, preserving the C
CK_Status
meanings and first-error order. - Added an isolated
runchild, secure persistent object cache, fixed runtime
diagnostics/statuses, eager symbol resolution, and audited JIT page permissions. - Added checked/unchecked C-oracle performance gates and six-host functional,
artifact, dependency, provenance, and immutable release gates. - Reserved
mainand the seven print builtin names, limited Native target output
to the host, retired the standalone LLVM exported-shape promise, and kept
emit-csource-only. See the compatibility policy for migration guidance.
0.9.0 - 2026-08-26
- Added
breakandcontinuefor structured control insidewhileloops. - Added explicit
voidprocedures, emptyreturn;, and procedure-call statements. - Added non-owning
slice<T>values,slice(data, len), indexing,.data/.len,
and half-open sub-slices written asitems[start..end]. - Added optional checked slice bounds to the C backend through
--bounds checked;
unchecked bounds remain the default, while WASM and LLVM reject checked bounds. - Stabilized native C, WebAssembly, and LLVM output paths and their V0.9 ABIs.
- Reorganized the repository around durable compiler, contract, example, benchmark,
and test responsibilities without changing the compiler's public behavior. - Froze the V0.9 compatibility boundary: patch releases in the
0.9.xline preserve
accepted source, diagnostic identifiers, CLI behavior, textual MIR, and documented
ABI contracts. A later0.10.0may make documented breaking changes with migration
guidance; long-term compatibility begins with a future1.0.0release. - Added signed-off native
ckcrelease archives and SHA-256 checksums for macOS,
Linux, and Windows on both arm64 and x64.