[RFC] Shared accelerator operation contract #1770
Kenneth-Javier
started this conversation in
Ideas
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Summary
This is meant as the compute-side counterpart to #1057. That RFC established
how Colibri becomes model-agnostic without a generic forward: specialized family
engines, shared runtime mechanisms, migrated one consumer at a time behind a
contract that is the only way. This RFC proposes the same treatment for the
device side.
The starting point is not a design. It is an operation that four backends
already implement with the same meaning and nearly the same arguments, the
quantized matmul, and that each engine currently reaches in its own way. The
proposal is to make that existing contract explicit, mandatory and tested, and
then to use it as the template for the next operation.
No generic forward, no new runtime dependency, and no change of placement or
numerics in the first step.
Evidence (dev
4e28e398, read-only)Four backends already agree on the operation. Same arguments, same order,
same "1 = handled, 0 = not mine" return:
coli_cuda_matmul(ColiCudaTensor **t, y, x, weights, scales, fmt, S, I, O, device, gs)coli_metal_matmul(ColiMetalTensor **t, y, x, weights, scales, fmt, S, I, O, gs)coli_vk_matmul(ColiVkTensor **t, y, x, weights, scales, fmt, S, I, O, gs)coli_xdna_try_matmul(family, ColiXdnaPrepared **slot, fmt, q4, scale, I, O, gs, planar, y, x, S)Two more entry points are the same operation in another form:
coli_metal_gemm(fmt 1/2/4, the same parameters, but no per-tensor handle: a synchronous GEMM
over unregistered weights) and
coli_vk_matmul_pair(two quantized matmuls thatshare one input, fused into one submit). The contract has to cover both: a
provider may keep no per-tensor state, and a call may carry two outputs over one
input.
The semantic common core is y, x, weights, scales, fmt, S, I, O, gs.
Provider-local prepared state (
ColiCudaTensor,ColiVkTensor,ColiMetalTensor,ColiXdnaPrepared) may exist, but it belongs to theprovider's resource plane, not to the operation. XDNA2 adds
the two things a GPU-only design never had to spell out: a semantic tag
(
family, which qualification is bound to, never inferred from shape) and theweight layout (
planar).The engines reach it in different ways. 18 direct calls in four engines:
colibri.c6,kimi_k3.c7,qwen36_tier.c3,glm53.c2. Insidecolibri.calone, Metal and CUDA/HIP are tried inside
matmul_qt_ex(:1236), while Vulkanand XDNA2 are called at individual sites.
Placement is decided three different ways:
QT.cuda_eligible, set at loadVK_FMT_OK, Vulkan's format rule copied into the engine in 7 placesThat is also why the obvious refactor is the wrong one. Folding Vulkan and XDNA2
into
matmul_qt_exwould change which tensors may try them. The request has tobe separated from placement before dispatch can be centralised.
What is not a compute contract:
ColiEdgeAdapterand the Segment runtime areengine-level ABIs for distributed inference. They are the idiom to follow
(struct size, ABI version, explicit registration, opaque implementation), but
Colibri has no device-level operation contract today.
Invariants
consume operations; they stop owning accelerator implementations.
what to compute. Policy says which implementations may try. A provider
says how.
cuda_device,XRT context or AIE tile in a request; those stay inside providers.
README.md:36forbids automatic BF16substitution, so a reduced-precision provider is eligible only on an explicit
request, and measured placement compares only providers in the same
exactness class.
need not have the same right to use the same accelerator artifact.
only through the shared contract, and a test fails if it does not. This is
the [RFC] Model-agnostic Colibrì without a generic forward: shared runtime, specialized family engines #1057 rule: a half-finished abstraction is worse than none. How the test
detects a bypass (today, a scan for backend entry points) is an
implementation detail; the invariant outlives any renaming of backend
functions.
caller does not mean reusing a symbol with a new signature across the
runtime-loaded DLL boundary (
coli_cuda.dll,coli_hip.dll). A host mustnever call an old backend with a new calling convention.
later step, not part of the first migration.
and XDNA2 against synthetic helpers; every phase keeps that property.
Design sketch
Split what lives as long as the weights from what changes per call:
Providers keep their existing per-tensor state (
ColiCudaTensor,ColiVkTensor,ColiXdnaPrepared), or none, as Metal's GEMM does today. The first migrationdoes not move tensor ownership.
XDNA2's
COLI_XDNA_FAMILY_*becomes a sharedColiSemanticOp, because it wasnever an XDNA concept: it is the qualification key for what an operation means.
Capability moves, placement does not. Today the engine decides both whether
a tensor may go to Vulkan and whether Vulkan can handle its format. After the
first step the engine still makes the first decision; the provider answers the
second (
VK_FMT_OKbecomes a Vulkan-sidesupports(desc)).The DLL boundary. The loader already binds each
coli_cuda_*symbol by nameand treats a missing mandatory symbol as fatal for the backend, with a
diagnostic naming it (
backend_loader.c,RESOLVE). A new, versioned entrypoint, e.g.
coli_cuda_qmatmul_v1(ColiCudaTensor **, const ColiQMatmulDesc *, const ColiQMatmulCall *), resolved as mandatory, means an older DLL is refusedcleanly rather than called with the wrong arguments. The two mechanisms cover
different things: the versioned symbol governs the calling convention, and
struct_size/abi_versiongovern how the structures behind it evolve. Aget_api(requested_version)table can replace per-symbol evolution later, ifthe surface grows.
Phased migration
Each phase follows the #1057 discipline: coverage lands first, one mechanism
per PR, and nothing optional is left behind.
B2: the quantized-matmul contract
providers CI can execute, currently Vulkan/Lavapipe. For reduced-precision providers, explicit expected-result or
tolerance checks plus eligibility checks; the synthetic XDNA2 helpers validate
dispatch and failure semantics, including the existing guarantee that every
failure falls back to a bit-identical
matmul_qtresult. And compile coverage for everycaller being migrated, which is not there today: CI builds
colibriwithVK=1but neverglm53orkimi_k3, and itsMETAL=1loop coverscolibri,inklingandkimi_k3but notglm53. Five of the 18 calls(
glm532,kimi_k33) would change signature with nothing compiling them. That gap existsregardless of this RFC; the first commit closes it.
ColiQMatmulDesc/ColiQMatmulCall, including the pair form; every enginecall site in scope and the tests migrated in the same PR. The counts in this
RFC are baseline facts for
4e28e398: implementation starts by re-auditingcurrent
devand migrates every caller in scope at that point.backend entry points of this operation (
coli_cuda_matmul,coli_metal_matmul,coli_metal_gemm,coli_vk_matmul,coli_vk_matmul_pair,coli_xdna_try_matmul), with the provider filesexempt, derived from the sources rather than from a hand list. Each name was
checked to be this operation and nothing else, so the scan forbids no
unrelated backend code.
VK_FMT_OKleaves the engine.longer resolved.
ownership, dependencies.
B3: provider chain and policy. Placement becomes explicit alongside the
request, as policy / eligibility metadata rather than code location, never
inside the compute request itself. Dispatch becomes one provider chain, and measured selection within
an exactness class builds on the startup GEMV timing that v1.12.1 introduced for
Qwen3.6.
B4: a second operation, the expert group. CUDA and Vulkan already agree on
its signature (
coli_*_expert_group); Metal'smoe_blockneeds an adapter.This is the phase that shows whether the contract generalises or merely fitted
matmul.
Non-goals
are treated as artifact producers for an existing provider; runtime
compilation is outside this RFC.
ink_cuda_matmul_bf16): a different operation withits own backend, left for a later phase rather than half-migrated now.
Questions for the maintainer
ColiSemanticOp/COLI_SEM_*for what XDNA2 calls a family?coli_cuda.dll/coli_hip.dlloutright, with the existingmissing-symbol diagnostic: acceptable, given that host and DLL ship together?
test_makefile_deps.py, or in a newcontract test?
Related
#1057 (shared runtime, specialized engines), #1261 (the XDNA2 lane that surfaced
the semantic tag and the layout dimension), #1590 (Linux portability, which a
provider boundary makes a runtime question rather than an engine change), #959.
All reactions