0.7: Frontend finalization and code-movement-based optimizations
Release note: 0.6.4 is skipped as a tagged release (last release before 0.7 was 0.6.3). The
work originally planned for 0.6.4/0.6.5/0.7.0 (frontend finalization, concatenation,
position embeddings, transformer toy) and for 0.7.2 (compiler optimizations, pool
allocator) is consolidated into 0.7. See ROADMAP.md.
Added
-
Removed the hosted tensor mode (gh-ocannl-333): dropped the
arrayfield of
Tnode.tand the "hosted" memory mode. Tensor value access and printing are now
context-mediated; host-init nodes self-initialize at link time. Removed the dead
automatic_host_transferssetting and theuse_host_memoryhook. -
Tensor saving, loading, and restoring (gh-ocannl-373).
-
Ternary einsum notation:
Einsum_ternin shape inference, PPX dispatch, aMul3
ternary scalar op across all backends, andeinsum3/ where-with-spec (gh-ocannl-305). -
Loop-invariant code motion (loop hoisting) prior to visit counting (gh-ocannl-350).
-
Common subexpression elimination after inlining (gh-ocannl-351).
-
Virtual-node inlining extended to non-scalar constants and ranges (gh-ocannl-142).
-
Uint32/Uint64precisions, with index-embedding operations selecting precision per
thelarge_modelssetting (formerlybig_models, still accepted as a deprecated alias)
to avoid unnecessary conversions and to select Metal pool-slot width
(gh-ocannl-349, gh-ocannl-177, gh-ocannl-344). -
-march=nativeC-compiler flag (gh-ocannl-311); restored CUDA pre-loaded builtins
referenced by pointer via a cudajit helper (gh-ocannl-353). -
Sasha Rush Tensor Puzzles expressed in the extended einsum notation (gh-ocannl-308).
-
Universal pool allocator across backends (gh-ocannl-344): tensors are addressed by
{ pool_id; offset }; working tensors are bump-packed into per-context-delta pools,
constants into per-device constant pools, and merge buffers stay in their reserved pool.
Metal now binds pool slabs plus a slot table instead of one buffer per tensor node, staying
below the backend's binding limit for large routines. CUDA and C keep per-tnode pointer
kernel parameters resolved from the same pooled locations. -
Data-parallel training:
shard_along/gathersharding primitives and a driver
with merge-buffer gradient all-reduce (part of gh-ocannl-293). -
Zero-copy slice views (
@|/Fetch.Slice, gh-ocannl-293 subtask 293a): an
alias-eligible leading-axis slice no longer materializes a copy — it is lowered as a
view that redirects reads/writes to the parent buffer at the (runtime) batch index.
Semantics change: writing through such a slice now mutates the parent (and vice
versa). Slices fall back to the materializing copy loop when ineligible (padded,
precision-converting, virtual/constant parent, or non-leading-axis). Host-side value
access (Context.get_values/set_values) of an alias view is rejected with a clear
error — read or write the parent tensor instead. -
Benchmark tables partitioned by
result_labelinto per-group sub-records (gh-ocannl-140). -
Nn_blocks.batch_norm1d— MLP batch normalization that normalizes over the
batch axis only. Mirrorsbatch_norm2d; inherits its running-statistics
FIXME (inference uses the learnedgamma/betaon batch statistics rather
than population estimates). -
test/training/mlp_names.ml— Bengio-style MLP (makemore Part 2): learned
character embeddings,block_size = 3context, einsum-contracted hidden
layer, deterministic 80/10/10 train/dev/test split. Prints numeric final
train/dev/test NLL plus threshold booleans, then three generated names. -
test/training/mlp_bn_names.ml— MLP +batch_norm1d(makemore Part 3).
Same data pipeline as Part 2 with BatchNorm between the hidden linear and
tanh. Documents the single-example-inference collapse from the
running-stats FIXME. -
docs/makemore_tutorial.md— walk-through of the makemore progression
(Parts 1–4 + cross-link to the transformer variant) mirroring Andrej
Karpathy's Neural Networks: Zero to Hero lectures, with a README entry
and Part 4 instructions for inspecting the generated backward code. -
Axis concatenation/block tensor support in einsum notation (
a^bsyntax)- Tensor concatenation (
a; b => a^b) - Axis slicing to extract prefix/suffix (
a^b => a,a^b => b) - Block tensor construction with n-ary einsum specs
invalid_varstracking for determining which dimension variables can be 0 in Block specs- New
++^operator for concatenation in DSL - Concat projection unification in shape inference (
solve_proj_equations)
- Tensor concatenation (
-
Pointwise operations now optionally accept einsum/permute specs via
?specand?capture_dimsparameters- Binary ops:
add,sub,pointmul,pointpow,pointdiv,lt,eq,ne - Unary ops:
relu,sat01,exp,log,exp2,log2,sin,cos,sqrt,recip,recip_sqrt,tanh,neg,not,stop_gradient
- Binary ops:
-
Common gotchas and idioms section in CLAUDE.md documentation
-
Rev_sidessupport in lowering for reverse-direction Block operations -
Completed the workshop article in Markdown and LaTeX, with a rendered PDF published
underdocs/html/pdfs/ocannl_workshop_article_human.pdf. -
Added the standalone formal core technical report
(docs/ocannl-formal-core-technical-report.md/.latex) covering the core
shape/projection inference proof effort: dimensions, rows, broadcasting,
flat row equality, solving, closing, and projection inference. -
Added
docs/shape-constraint-generation.md, documenting howtensor/shape.ml
generates the core constraints and projection metadata used by inference. -
Added regression coverage for shape-inference counterexamples, closing order,
and row-rank-cycle behavior used while validating the formalization.
Changed
test/training/bigram_mlp.mlrenamed totest/training/mlp_names.mland
rewritten as a true multi-character-context Bengio MLP. The old file's name
misrepresented its architecture (bigram-width input but "MLP" label); the
new file is the makemore Part 2 example (seedocs/makemore_tutorial.md).- Default
%opparameter initialization now uses centered, scaleduniform1over
[-0.25, 0.25)while preservinguniform1's flexible shape behavior. - Parser updated to allow n-ary einsum specs (e.g.,
a;b;c;d=>result) - Concat symbols are now grouped into connected components for iteration using union-find
- Product space and product iterators now use list arrays to handle concatenated dimensions
- Moved
datasets/to separatedatapreppackage - Relaxed the required
ocannl_prefix on commandline arguments; config keys are now
validated, andocannl_config.examplewas renamed toocannl_config.reference
(gh-ocannl-409). - Renamed routine/kernel parameters from
param/paramstokparam/kparams
(gh-ocannl-356). - Extended the identifier blacklist with C keywords, primitive-operator names, and
backend-specific reserved words (gh-ocannl-383);debug_namenow collapses
consecutive identical label components, e.g.ident×3 →ident3(gh-ocannl-281). - Removed remaining unnecessary buffer zeroing-out in backend code (gh-ocannl-382).
- Upgraded slipshow presentation rendering to v0.11.0, with Mermaid diagrams
(gh-ocannl-425). - Heavy training integration tests are now gated behind Dune's
slowalias, while
backend-divergent goldens and CUDA/Metal generated-source expectations were normalized
for release testing. - Documentation for the formal core now uses the direct row-subtyping/refinement
presentation consistently, including the closed-row equality and dimension-closing
policy clarifications made while preparing the workshop article. - Breaking:
Backend.device_to_devicenow returnscontext routine optioninstead ofbool.
Instead of scheduling the copy as a side effect, it builds a transfer routine: callers run
r.schedule(or link a consumer againstr.context).Nonereplaces the oldfalse("nothing
to transfer": node absent fromsrc; or, forinto_merge_buffer:No, node absent fromdstor
identical source/destination buffers). The transfer routine's context records the produced
merge-buffer node in the newBackend_intf.context.merge_buffer_nodefield, so that linking a
consumer of the merge buffer against it statically verifies the node at link time (raising
Utils.User_errorfromlink/link_batchon a mismatch), "in the right direction" — transfer
-> consumer (gh-ocannl-288). The runtimecheck_merge_buffercheck is kept as a defensive
backstop.
Fixed
- Detect rank cycles among row variables during shape inference (gh-ocannl-247).
- CUDA
Whereexpression codegen now parenthesizes ternaries correctly, and
Uint32/Uint64touint4x32PRNG-counter conversions spread bits rather than
collapsing entropy. - Prohibit
~logic:"@"(Compose) with/and**in the%cdextension, and fixed
ternary~logicbeing mapped tocompose_typeinstead ofternary_type
(gh-ocannl-192). - C-syntax tracing
printfstatements no longer produce bad line breaks / indentation
(gh-ocannl-179). - Removed a duplicate
fsm_transformertest stanza intest/training/dune
that trippeddune build @checkwithExecutable "fsm_transformer" appears for the second time in this directory. - Missing on-device margin initialization for
Fetchcases invalid_varscomputation now uses correct four-quantifier logic- Zero-dimension components filtered out in
s_dim_onesubstitution - Single remaining Concat component properly closed in
close_dim_terminal - Inequality constraints preserved for Concat with
invalid_vars - Dimension 0 (instead of 1) guessed for
invalid_varsin all guessing locations - Concat lowering for unit dims
- Concat index probing guarded against unknown projections
- Concat lowering resolves indices with cumulative offsets correctly
d=1handling in Concat projection components