Skip to content

v0.0.1 — A foundry for agents

Choose a tag to compare

@zhen8838 zhen8838 released this 09 Aug 05:09
· 94 commits to main since this release
403db1f

This is the first public release of TileFoundry, a tile-based, agentic platform for automatic high-performance program generation across hardware.

pip install tilefoundry   # Python ≥ 3.12, PyTorch ≥ 2.7

Why we built this

High-performance kernels have always been expert work. The optimization space is huge — tiling, layouts, pipelining, and instruction choice all interact — and every decision made for speed still has to be checked for correctness.

LLM agents can now do much of this work, but existing compiler stacks were not built for them. The intermediate states are opaque graphs, the knowledge an agent needs is scattered across the codebase, and "is this still correct?" has no mechanical answer.

TileFoundry starts from one question: what should a program-generation platform look like when an agent is the primary user?

The answer comes down to three things, one for each problem above:

  • Source to source — so agent edits stay visible and structurally safe
  • Context on demand — so the right information arrives at the right time
  • Fail closed, verify explicitly — so nothing passes silently

Source to source

The workflow is a loop:

  1. Describe. Write a published model in HIR, TileFoundry's high-level IR — a readable reference implementation, finished when it agrees with the original on real weights.
  2. Optimize. Write a runtime twin of one module, check it against the reference, and iterate while the numbers disagree.
  3. Repeat. Take the next module. Fuse neighbours. Check again.

Both sides of every comparison are readable Python source; nothing is hidden behind an opaque graph. Runtime twins preserve the reference's module and function structure, so omissions fail early instead of surviving to the final check.

Correctness is always one check away. That lets an agent optimize aggressively — and lets a human trust the result.

Context on demand

Question Command
What has been described? tilefoundry models
What does the contract say? tilefoundry spec
Does it agree with its reference? tilefoundry check
What does it cost? tilefoundry analyze

The design lives in a normative specification, installed with the package and queryable from the command line. Instead of front-loading that knowledge into a prompt, TileFoundry lets an agent ask the question it has now. Every command level describes itself, so context arrives at the moment it becomes useful. For an agent driving the CLI in a loop, this is the difference between converging and thrashing.

Fail closed, verify explicitly

check has no default tolerance. Every output is judged by a bound the caller states, because a silent PASS teaches nothing.

Every shipped model states how far it has been verified, with the test cited as evidence. The yardstick is the model's official implementation — the Hugging Face Transformers code people already run and trust:

  • L1 — each layer matches the official implementation, on random weights
  • L2 — the whole forward chain matches it, end to end
  • L3 — decoding from the released checkpoint reproduces its output, token for token

How much a model is trusted is a lookup, not a judgment call.

In this release

The heart of the release is the loop itself — the language for describing models, and the models already carried through it:

  • A typed, shard-aware HIR — a Python DSL for describing models as nested Modules of @func kernels, covering current inference dtypes (bf16, fp8e4m3, f8e8m0 block scales).
  • Seven real models already described with it — DeepSeek V4 Flash, Gemma2 2B, MiniCPM3 4B, Qwen2.5 1.5B, Qwen3 1.7B (L2); Kimi Linear 48B A3B, Qwen3.5 35B A3B (L1) — installed with the package as references to copy from.

Around them, the machinery and the on-ramp:

  • Scheduling built for agents — polyhedral modeling, storage-driven instruction selection, and generated kernel skeletons with holes left for the agent to fill in.
  • Three end-to-end examples of optimized decode — hand-written CUDA, CuTe DSL, and TileLang — and a tutorial that teaches the loop: what to do, in what order, at what granularity.

What's next

This release establishes the foundation: the DSL for describing models, the check loop for runtime implementations, and the CLI an agent works through.

The next phase is to make the platform more useful inside the optimization loop. TileFoundry should not only report a cost or a failure; it should use the current IR, target facts, and measurements to give the agent precise, stage-appropriate guidance about what to try next. In parallel, we are extending the same describe–check–optimize workflow to more hardware targets and more inference scenarios.

Follow along

TileFoundry is early — alpha, and the APIs are still moving. Which means design feedback matters more now than it ever will again.

  • Star and watch the repo to follow releases; each release note will explain the design as it grows.
  • Open an issue with questions or objections — the spec is meant to be argued with.
  • Try the loop: install TileFoundry 0.0.1 from PyPI, then run tilefoundry models and tilefoundry tutorial.

The foundry is open. Come build it with us.