Repository navigation
Releases: raws-labs/tigris-runtime
Release list
v0.11.4
Patch release. Accepts plan schema 2 through 9 and pairs with compiler v0.11.4.
Fixed
- Plans with schema 2 or 3 skipped operator semantic validation, so a record
no kernel can execute, such as a Conv with an unsupported group, loaded
anyway. The loader now validates the built-in operators those schemas
already defined. Transpose, MatMul and opcodes outside that table stay open
to custom dispatch, as before.
v0.11.3
v0.11.2
Patch release. Accepts plan schema 2 through 9 and pairs with compiler v0.11.2.
Fixed
- v0.11.1 failed to build with GCC 11 at
-O3when warnings are errors,
because the fixed-point Softmax parameters were reported as maybe
uninitialized. They are now set on every path. Behavior is unchanged from
v0.11.1.
v0.11.1
Patch release. Accepts plan schema 2 through 9. Its source does not build with GCC 11 at -O3 when warnings are errors; use v0.11.2, which pairs with compiler v0.11.2.
Fixed
- The int8 Softmax kernel computed in float, so its output could differ by one
LSB from TFLite, TFLite Micro and CMSIS-NN. With TFLite's int8 Softmax output
quantization (scale 1/256, zero point -128) it now runs their fixed-point
algorithm and matches them bit for bit. Other output quantizations keep the
float computation. - An int8 global average pool that stands for a whole-map
AveragePoolcan
now round like TFLite'sAVERAGE_POOL_2D: the integer sum divided with ties
away from zero, instead of theMEANrequantization that could differ by one
LSB. Plans from compiler v0.11.1 select it with a new operator attribute.
Older runtimes refuse such plans instead of running them with the other
rounding. The CMSIS-NN backend runs plain global average pools on the
reference kernel, sincearm_avgpool_s8rounds likeAVERAGE_POOL_2D.
v0.11.0
First paired release: compiler and runtime now release together with the same version number. Accepts plan schema 2 through 9 and pairs with compiler v0.11.0.
Added
tigris_host, a shared library for desktop and server machines that runs
compiled plans on the portable float32 and int8 reference kernels through a
small allocating C API (host/tigris_host.h). It backstigris runin
tigris-ml. Build it with-DTIGRIS_BUILD_HOST=ON.- Release archives of the host library for Linux x86-64 and ARM64
(manylinux_2_28), macOS 11 or newer on x86-64 and ARM64, and Windows
x86-64, each with its header, license, a manifest and a SHA-256 file. - The runtime and the host library build with MSVC.
Fixed
- A tile's top or bottom padding outside 0 to 65535 is refused with
TIGRIS_EXEC_ERR_TILEinstead of being truncated when stored.
Changed
- The installed CMake package accepts requests within its own minor release
line only, so CMake consumers changefind_package(tigris_runtime 0.10)to
find_package(tigris_runtime 0.11).
v0.10.2
Patch release. Accepts plan schema 2 through 9 and pairs with compiler v0.9.0, as v0.10.1 does.
Fixed
- On ESP-IDF the executor called
vTaskDelay(1)after every stage and every
tile, so each inference waited up to one FreeRTOS tick per stage and tile.
The call is removed andtigris_run()now runs to completion on the calling
task. Plans, outputs and the plan format are unchanged.
Documentation
- The README describes how to service a task watchdog when an inference
outlasts its timeout.
v0.10.1
Patch release. Accepts plan schema 2 through 9 and pairs with compiler v0.9.0, as v0.10.0 does.
Fixed
- An
Add,MulorSubinside a tiled chain could combine different image
rows when one operand kept halo rows the other path had already consumed. Each
operand is now aligned to the operation's row range before dispatch, and row
bookkeeping resets for every chain. Plans and the plan format are unchanged.
Documentation
- The README generates test fixtures with the compiler repository's
scripts/gen_fixtures.py.
v0.10.0
Accepts plan schema 2 through 9 and pairs with compiler v0.9.0.
Added
- Schema 8 operator attributes and schema 9 row bands for matrix pipelines.
- Kernels:
LayerNormalization,Erf,HardSwish,Sub,Split, a maximum
pool to a global maximum, and bilinearResize, float and int8. - Tiled execution for global reductions along their input, layout conversions,
rank-2 matrix pipelines along their rows, rank-3 tensors in 2D tiles and
Resizestages. AddandMulwith a per-channel second operand.- uint8 model inputs and outputs, converted at the boundary.
tigris_esp_nn_enable_dual_core()gives the ESP-NN kernels the ESP32-S3's
second core with unchanged results.
Changed
- The ESP-IDF component requires esp-nn 1.4.0 or later. Releases below 1.3
miscompute the mult8 1x1 convolution on some channel counts. tigris_esp_nn_prepare()returns -2 when a Conv or depthwise weight or bias
is not 16-byte aligned. ESP-NN reads filters in place and computes wrong
results from misaligned ones, so the plan buffer has to start on a 16-byte
boundary. The getting-started example embeds its plan throughmain/plan.S.TIGRIS_TENSOR_ALIGNis 16 on Xtensa.
Fixed
- ESP-NN Conv calls passed a filter channel count of 0, which ESP-NN 1.3 and
later read as a grouped convolution and run on their ANSI C path. With
esp-nn 1.4 on ESP32-S3, MobileNetV1 at a 128K fast arena goes from 12.5 s to
1.19 s. - ESP-NN depthwise writes stay within the output buffer.
- A tiled stage writes each output at its own rows.
- Half-pixel bilinear bands start at the first input row they read.
- The two counts a plan gives for its quantization parameters are reconciled.
v0.9.1
Patch release. No runtime or plan-format change: accepts plan schema 2 through 7 and pairs with compiler v0.8.0, exactly as v0.9.0 does.
The release tag's CI job required a compiler tag of the same name. Compiler and runtime version independently and have never shared one, so a release tag could only ever fail that check, and did on v0.8.0 and v0.9.0. A tag now pairs against the released compiler, which is the tip of its main.
The same job resolved a ref by bare name, and ls-remote patterns match a ref's tail, so a tag could pair against an unrelated branch whose name merely ended in it. Both sides match a full refs/heads path now.
v0.9.0
Pairs with compiler v0.8.0, which emits plan schema 7. This runtime accepts 2 through 7; earlier runtimes reject a schema-7 plan at load, so the two release together.
A tensor records how it is stored
Activations are held channels-last, but a model output written by a terminal Transpose keeps the axis order the model states, and nothing in the plan said which of the two a given boundary was. A caller could not recover it, and neither could the cross-repo contract harness, which had been guessing from the shape of the graph that produced the tensor. Schema 6 made the boundary dtype self-describing; TIGRIS_TENSOR_LINEAR does the same for axis order.
The flag is the load-time half of a compiler change that lets one graph hold both kinds of tensor, which is what makes a matrix product expressible at all.
Two activations can be multiplied
kern_matmul and kern_matmul_s8 compute a product whose second operand is an activation rather than a weight. The fully-connected kernel cannot express that, because it reads its weight from the weight table. The loader states the contract: both operands the same rank, batch extents agreeing, inner dimensions meeting, and no broadcasting of batch dimensions, which is rejected rather than guessed at. The int8 kernel takes both zero points from the tensor table and requantizes with a multiplier the compiler derives from the product of the two input scales.
The float fully-connected kernel batches
The int8 kernel has always taken [N, IC] against [OC, IC] and written [N, OC]. The float one computed a single row, and the loader enforced that asymmetry for float only, so a linear layer applied per sequence position was executable in int8 and rejected in float.
Softmax runs in a tiled stage
Both kernels refused tiled execution outright. Normalization runs along the final stored dimension, which is the channel axis in both NLC and NHWC, while a tile cuts the axis ahead of it, so a tile always holds whole normalization rows and needs nothing from its neighbours. A tensor larger than the fast pool no longer has to fit whole.
Version sources are held together
idf_component.yml had sat a release behind on the integration branch, and the CMake project version had not moved since 0.6.0. scripts/check_version_sources.py compares the four places that state this library's version, and the publish workflow checks them against the tag before uploading, so a stale manifest cannot consume a registry version.