Skip to content

Releases: raws-labs/tigris-runtime

v0.11.4

Choose a tag to compare

@asteinh asteinh released this 29 Sep 14:07

Patch release. Accepts plan schema 2 through 9 and pairs with compiler v0.11.4.

Fixed

  • Plans with schema 2 or 3 skipped operator semantic validation, so a record
    no kernel can execute, such as a Conv with an unsupported group, loaded
    anyway. The loader now validates the built-in operators those schemas
    already defined. Transpose, MatMul and opcodes outside that table stay open
    to custom dispatch, as before.

v0.11.3

Choose a tag to compare

@asteinh asteinh released this 27 Sep 04:36

Patch release. Accepts plan schema 2 through 9 and pairs with compiler v0.11.3.

No runtime implementation changes; released with the matching compiler version.

v0.11.2

Choose a tag to compare

@asteinh asteinh released this 26 Sep 19:09

Patch release. Accepts plan schema 2 through 9 and pairs with compiler v0.11.2.

Fixed

  • v0.11.1 failed to build with GCC 11 at -O3 when warnings are errors,
    because the fixed-point Softmax parameters were reported as maybe
    uninitialized. They are now set on every path. Behavior is unchanged from
    v0.11.1.

v0.11.1

Choose a tag to compare

@asteinh asteinh released this 26 Sep 18:37

Patch release. Accepts plan schema 2 through 9. Its source does not build with GCC 11 at -O3 when warnings are errors; use v0.11.2, which pairs with compiler v0.11.2.

Fixed

  • The int8 Softmax kernel computed in float, so its output could differ by one
    LSB from TFLite, TFLite Micro and CMSIS-NN. With TFLite's int8 Softmax output
    quantization (scale 1/256, zero point -128) it now runs their fixed-point
    algorithm and matches them bit for bit. Other output quantizations keep the
    float computation.
  • An int8 global average pool that stands for a whole-map AveragePool can
    now round like TFLite's AVERAGE_POOL_2D: the integer sum divided with ties
    away from zero, instead of the MEAN requantization that could differ by one
    LSB. Plans from compiler v0.11.1 select it with a new operator attribute.
    Older runtimes refuse such plans instead of running them with the other
    rounding. The CMSIS-NN backend runs plain global average pools on the
    reference kernel, since arm_avgpool_s8 rounds like AVERAGE_POOL_2D.

v0.11.0

Choose a tag to compare

@asteinh asteinh released this 26 Sep 10:48

First paired release: compiler and runtime now release together with the same version number. Accepts plan schema 2 through 9 and pairs with compiler v0.11.0.

Added

  • tigris_host, a shared library for desktop and server machines that runs
    compiled plans on the portable float32 and int8 reference kernels through a
    small allocating C API (host/tigris_host.h). It backs tigris run in
    tigris-ml. Build it with -DTIGRIS_BUILD_HOST=ON.
  • Release archives of the host library for Linux x86-64 and ARM64
    (manylinux_2_28), macOS 11 or newer on x86-64 and ARM64, and Windows
    x86-64, each with its header, license, a manifest and a SHA-256 file.
  • The runtime and the host library build with MSVC.

Fixed

  • A tile's top or bottom padding outside 0 to 65535 is refused with
    TIGRIS_EXEC_ERR_TILE instead of being truncated when stored.

Changed

  • The installed CMake package accepts requests within its own minor release
    line only, so CMake consumers change find_package(tigris_runtime 0.10) to
    find_package(tigris_runtime 0.11).

v0.10.2

Choose a tag to compare

@asteinh asteinh released this 25 Sep 15:07
5078c33

Patch release. Accepts plan schema 2 through 9 and pairs with compiler v0.9.0, as v0.10.1 does.

Fixed

  • On ESP-IDF the executor called vTaskDelay(1) after every stage and every
    tile, so each inference waited up to one FreeRTOS tick per stage and tile.
    The call is removed and tigris_run() now runs to completion on the calling
    task. Plans, outputs and the plan format are unchanged.

Documentation

  • The README describes how to service a task watchdog when an inference
    outlasts its timeout.

v0.10.1

Choose a tag to compare

@asteinh asteinh released this 25 Sep 11:46
0838dd0

Patch release. Accepts plan schema 2 through 9 and pairs with compiler v0.9.0, as v0.10.0 does.

Fixed

  • An Add, Mul or Sub inside a tiled chain could combine different image
    rows when one operand kept halo rows the other path had already consumed. Each
    operand is now aligned to the operation's row range before dispatch, and row
    bookkeeping resets for every chain. Plans and the plan format are unchanged.

Documentation

  • The README generates test fixtures with the compiler repository's
    scripts/gen_fixtures.py.

v0.10.0

Choose a tag to compare

@asteinh asteinh released this 25 Sep 05:21
0e3b2fa

Accepts plan schema 2 through 9 and pairs with compiler v0.9.0.

Added

  • Schema 8 operator attributes and schema 9 row bands for matrix pipelines.
  • Kernels: LayerNormalization, Erf, HardSwish, Sub, Split, a maximum
    pool to a global maximum, and bilinear Resize, float and int8.
  • Tiled execution for global reductions along their input, layout conversions,
    rank-2 matrix pipelines along their rows, rank-3 tensors in 2D tiles and
    Resize stages.
  • Add and Mul with a per-channel second operand.
  • uint8 model inputs and outputs, converted at the boundary.
  • tigris_esp_nn_enable_dual_core() gives the ESP-NN kernels the ESP32-S3's
    second core with unchanged results.

Changed

  • The ESP-IDF component requires esp-nn 1.4.0 or later. Releases below 1.3
    miscompute the mult8 1x1 convolution on some channel counts.
  • tigris_esp_nn_prepare() returns -2 when a Conv or depthwise weight or bias
    is not 16-byte aligned. ESP-NN reads filters in place and computes wrong
    results from misaligned ones, so the plan buffer has to start on a 16-byte
    boundary. The getting-started example embeds its plan through main/plan.S.
  • TIGRIS_TENSOR_ALIGN is 16 on Xtensa.

Fixed

  • ESP-NN Conv calls passed a filter channel count of 0, which ESP-NN 1.3 and
    later read as a grouped convolution and run on their ANSI C path. With
    esp-nn 1.4 on ESP32-S3, MobileNetV1 at a 128K fast arena goes from 12.5 s to
    1.19 s.
  • ESP-NN depthwise writes stay within the output buffer.
  • A tiled stage writes each output at its own rows.
  • Half-pixel bilinear bands start at the first input row they read.
  • The two counts a plan gives for its quantization parameters are reconciled.

v0.9.1

Choose a tag to compare

@asteinh asteinh released this 15 Sep 10:38
a949529

Patch release. No runtime or plan-format change: accepts plan schema 2 through 7 and pairs with compiler v0.8.0, exactly as v0.9.0 does.

The release tag's CI job required a compiler tag of the same name. Compiler and runtime version independently and have never shared one, so a release tag could only ever fail that check, and did on v0.8.0 and v0.9.0. A tag now pairs against the released compiler, which is the tip of its main.

The same job resolved a ref by bare name, and ls-remote patterns match a ref's tail, so a tag could pair against an unrelated branch whose name merely ended in it. Both sides match a full refs/heads path now.

v0.9.0

Choose a tag to compare

@asteinh asteinh released this 15 Sep 04:35
181ee6c

Pairs with compiler v0.8.0, which emits plan schema 7. This runtime accepts 2 through 7; earlier runtimes reject a schema-7 plan at load, so the two release together.

A tensor records how it is stored

Activations are held channels-last, but a model output written by a terminal Transpose keeps the axis order the model states, and nothing in the plan said which of the two a given boundary was. A caller could not recover it, and neither could the cross-repo contract harness, which had been guessing from the shape of the graph that produced the tensor. Schema 6 made the boundary dtype self-describing; TIGRIS_TENSOR_LINEAR does the same for axis order.

The flag is the load-time half of a compiler change that lets one graph hold both kinds of tensor, which is what makes a matrix product expressible at all.

Two activations can be multiplied

kern_matmul and kern_matmul_s8 compute a product whose second operand is an activation rather than a weight. The fully-connected kernel cannot express that, because it reads its weight from the weight table. The loader states the contract: both operands the same rank, batch extents agreeing, inner dimensions meeting, and no broadcasting of batch dimensions, which is rejected rather than guessed at. The int8 kernel takes both zero points from the tensor table and requantizes with a multiplier the compiler derives from the product of the two input scales.

The float fully-connected kernel batches

The int8 kernel has always taken [N, IC] against [OC, IC] and written [N, OC]. The float one computed a single row, and the loader enforced that asymmetry for float only, so a linear layer applied per sequence position was executable in int8 and rejected in float.

Softmax runs in a tiled stage

Both kernels refused tiled execution outright. Normalization runs along the final stored dimension, which is the channel axis in both NLC and NHWC, while a tile cuts the axis ahead of it, so a tile always holds whole normalization rows and needs nothing from its neighbours. A tensor larger than the fast pool no longer has to fit whole.

Version sources are held together

idf_component.yml had sat a release behind on the integration branch, and the CMake project version had not moved since 0.6.0. scripts/check_version_sources.py compares the four places that state this library's version, and the publish workflow checks them against the tag before uploading, so a stale manifest cannot consume a registry version.