Skip to content

0.7: Frontend finalization and code-movement-based optimizations

Choose a tag to compare

@lukstafi lukstafi released this 13 Jul 12:52
· 4508 commits to master since this release
bf85d03

Release note: 0.6.4 is skipped as a tagged release (last release before 0.7 was 0.6.3). The
work originally planned for 0.6.4/0.6.5/0.7.0 (frontend finalization, concatenation,
position embeddings, transformer toy) and for 0.7.2 (compiler optimizations, pool
allocator) is consolidated into 0.7. See ROADMAP.md.

Added

  • Removed the hosted tensor mode (gh-ocannl-333): dropped the array field of
    Tnode.t and the "hosted" memory mode. Tensor value access and printing are now
    context-mediated; host-init nodes self-initialize at link time. Removed the dead
    automatic_host_transfers setting and the use_host_memory hook.

  • Tensor saving, loading, and restoring (gh-ocannl-373).

  • Ternary einsum notation: Einsum_tern in shape inference, PPX dispatch, a Mul3
    ternary scalar op across all backends, and einsum3 / where-with-spec (gh-ocannl-305).

  • Loop-invariant code motion (loop hoisting) prior to visit counting (gh-ocannl-350).

  • Common subexpression elimination after inlining (gh-ocannl-351).

  • Virtual-node inlining extended to non-scalar constants and ranges (gh-ocannl-142).

  • Uint32/Uint64 precisions, with index-embedding operations selecting precision per
    the large_models setting (formerly big_models, still accepted as a deprecated alias)
    to avoid unnecessary conversions and to select Metal pool-slot width
    (gh-ocannl-349, gh-ocannl-177, gh-ocannl-344).

  • -march=native C-compiler flag (gh-ocannl-311); restored CUDA pre-loaded builtins
    referenced by pointer via a cudajit helper (gh-ocannl-353).

  • Sasha Rush Tensor Puzzles expressed in the extended einsum notation (gh-ocannl-308).

  • Universal pool allocator across backends (gh-ocannl-344): tensors are addressed by
    { pool_id; offset }; working tensors are bump-packed into per-context-delta pools,
    constants into per-device constant pools, and merge buffers stay in their reserved pool.
    Metal now binds pool slabs plus a slot table instead of one buffer per tensor node, staying
    below the backend's binding limit for large routines. CUDA and C keep per-tnode pointer
    kernel parameters resolved from the same pooled locations.

  • Data-parallel training: shard_along / gather sharding primitives and a driver
    with merge-buffer gradient all-reduce (part of gh-ocannl-293).

  • Zero-copy slice views (@| / Fetch.Slice, gh-ocannl-293 subtask 293a): an
    alias-eligible leading-axis slice no longer materializes a copy — it is lowered as a
    view that redirects reads/writes to the parent buffer at the (runtime) batch index.
    Semantics change: writing through such a slice now mutates the parent (and vice
    versa). Slices fall back to the materializing copy loop when ineligible (padded,
    precision-converting, virtual/constant parent, or non-leading-axis). Host-side value
    access (Context.get_values / set_values) of an alias view is rejected with a clear
    error — read or write the parent tensor instead.

  • Benchmark tables partitioned by result_label into per-group sub-records (gh-ocannl-140).

  • Nn_blocks.batch_norm1d — MLP batch normalization that normalizes over the
    batch axis only. Mirrors batch_norm2d; inherits its running-statistics
    FIXME (inference uses the learned gamma/beta on batch statistics rather
    than population estimates).

  • test/training/mlp_names.ml — Bengio-style MLP (makemore Part 2): learned
    character embeddings, block_size = 3 context, einsum-contracted hidden
    layer, deterministic 80/10/10 train/dev/test split. Prints numeric final
    train/dev/test NLL plus threshold booleans, then three generated names.

  • test/training/mlp_bn_names.ml — MLP + batch_norm1d (makemore Part 3).
    Same data pipeline as Part 2 with BatchNorm between the hidden linear and
    tanh. Documents the single-example-inference collapse from the
    running-stats FIXME.

  • docs/makemore_tutorial.md — walk-through of the makemore progression
    (Parts 1–4 + cross-link to the transformer variant) mirroring Andrej
    Karpathy's Neural Networks: Zero to Hero lectures, with a README entry
    and Part 4 instructions for inspecting the generated backward code.

  • Axis concatenation/block tensor support in einsum notation (a^b syntax)

    • Tensor concatenation (a; b => a^b)
    • Axis slicing to extract prefix/suffix (a^b => a, a^b => b)
    • Block tensor construction with n-ary einsum specs
    • invalid_vars tracking for determining which dimension variables can be 0 in Block specs
    • New ++^ operator for concatenation in DSL
    • Concat projection unification in shape inference (solve_proj_equations)
  • Pointwise operations now optionally accept einsum/permute specs via ?spec and ?capture_dims parameters

    • Binary ops: add, sub, pointmul, pointpow, pointdiv, lt, eq, ne
    • Unary ops: relu, sat01, exp, log, exp2, log2, sin, cos, sqrt, recip, recip_sqrt, tanh, neg, not, stop_gradient
  • Common gotchas and idioms section in CLAUDE.md documentation

  • Rev_sides support in lowering for reverse-direction Block operations

  • Completed the workshop article in Markdown and LaTeX, with a rendered PDF published
    under docs/html/pdfs/ocannl_workshop_article_human.pdf.

  • Added the standalone formal core technical report
    (docs/ocannl-formal-core-technical-report.md / .latex) covering the core
    shape/projection inference proof effort: dimensions, rows, broadcasting,
    flat row equality, solving, closing, and projection inference.

  • Added docs/shape-constraint-generation.md, documenting how tensor/shape.ml
    generates the core constraints and projection metadata used by inference.

  • Added regression coverage for shape-inference counterexamples, closing order,
    and row-rank-cycle behavior used while validating the formalization.

Changed

  • test/training/bigram_mlp.ml renamed to test/training/mlp_names.ml and
    rewritten as a true multi-character-context Bengio MLP. The old file's name
    misrepresented its architecture (bigram-width input but "MLP" label); the
    new file is the makemore Part 2 example (see docs/makemore_tutorial.md).
  • Default %op parameter initialization now uses centered, scaled uniform1 over
    [-0.25, 0.25) while preserving uniform1's flexible shape behavior.
  • Parser updated to allow n-ary einsum specs (e.g., a;b;c;d=>result)
  • Concat symbols are now grouped into connected components for iteration using union-find
  • Product space and product iterators now use list arrays to handle concatenated dimensions
  • Moved datasets/ to separate dataprep package
  • Relaxed the required ocannl_ prefix on commandline arguments; config keys are now
    validated, and ocannl_config.example was renamed to ocannl_config.reference
    (gh-ocannl-409).
  • Renamed routine/kernel parameters from param/params to kparam/kparams
    (gh-ocannl-356).
  • Extended the identifier blacklist with C keywords, primitive-operator names, and
    backend-specific reserved words (gh-ocannl-383); debug_name now collapses
    consecutive identical label components, e.g. ident ×3 → ident3 (gh-ocannl-281).
  • Removed remaining unnecessary buffer zeroing-out in backend code (gh-ocannl-382).
  • Upgraded slipshow presentation rendering to v0.11.0, with Mermaid diagrams
    (gh-ocannl-425).
  • Heavy training integration tests are now gated behind Dune's slow alias, while
    backend-divergent goldens and CUDA/Metal generated-source expectations were normalized
    for release testing.
  • Documentation for the formal core now uses the direct row-subtyping/refinement
    presentation consistently, including the closed-row equality and dimension-closing
    policy clarifications made while preparing the workshop article.
  • Breaking: Backend.device_to_device now returns context routine option instead of bool.
    Instead of scheduling the copy as a side effect, it builds a transfer routine: callers run
    r.schedule (or link a consumer against r.context). None replaces the old false ("nothing
    to transfer": node absent from src; or, for into_merge_buffer:No, node absent from dst or
    identical source/destination buffers). The transfer routine's context records the produced
    merge-buffer node in the new Backend_intf.context.merge_buffer_node field, so that linking a
    consumer of the merge buffer against it statically verifies the node at link time (raising
    Utils.User_error from link/link_batch on a mismatch), "in the right direction" — transfer
    -> consumer (gh-ocannl-288). The runtime check_merge_buffer check is kept as a defensive
    backstop.

Fixed

  • Detect rank cycles among row variables during shape inference (gh-ocannl-247).
  • CUDA Where expression codegen now parenthesizes ternaries correctly, and
    Uint32/Uint64 to uint4x32 PRNG-counter conversions spread bits rather than
    collapsing entropy.
  • Prohibit ~logic:"@" (Compose) with / and ** in the %cd extension, and fixed
    ternary ~logic being mapped to compose_type instead of ternary_type
    (gh-ocannl-192).
  • C-syntax tracing printf statements no longer produce bad line breaks / indentation
    (gh-ocannl-179).
  • Removed a duplicate fsm_transformer test stanza in test/training/dune
    that tripped dune build @check with Executable "fsm_transformer" appears for the second time in this directory.
  • Missing on-device margin initialization for Fetch cases
  • invalid_vars computation now uses correct four-quantifier logic
  • Zero-dimension components filtered out in s_dim_one substitution
  • Single remaining Concat component properly closed in close_dim_terminal
  • Inequality constraints preserved for Concat with invalid_vars
  • Dimension 0 (instead of 1) guessed for invalid_vars in all guessing locations
  • Concat lowering for unit dims
  • Concat index probing guarded against unknown projections
  • Concat lowering resolves indices with cumulative offsets correctly
  • d=1 handling in Concat projection components