Repository navigation
Releases: tile-ai/TileFoundry
Release list
v0.0.2 — Compiler feedback in the loop
TileFoundry 0.0.2 turns compiler analysis into feedback that an agent can act on while optimizing real programs.
pip install tilefoundry==0.0.2Highlights
New blog: AI Compilers in the Agentic Era
AI Compilers in the Agentic Era explains the boundary TileFoundry draws between an agent and a compiler.
The agent proposes strategies, rewrites programs, and implements kernels. TileFoundry keeps the authored HIR and runtime twin connected through three forms of compiler feedback:
analyzeattributes performance costs before a kernel is written.checkprotects correctness across large structural rewrites.- HIR types and placement expose optimization opportunities to the agent.
New example: a whole LLM decode step in one kernel
The new Nemotron-3.5-Lightning-30B-A3B example describes all 52 layers of a Mamba2, attention, and MoE model in authored HIR, then implements the decode step as one persistent TileLang cooperative kernel. (#143)
On the measured H200 setup, it reaches:
- 287.4 tok/s at context 32, or 97.5% of SGLang
- 231.6 tok/s at context 262080, or 83.2% of SGLang
- one kernel launch per decode step, compared with 3212 launches in the operator-by-operator path
The full agent run, HIR, runtime twin, checks, measurements, and optimization history are included in the example.
Features
- Added topology-scoped compute costs, authored-loop access footprints, execution timelines, and analytical performance bounds. (#80, #94, #105, #110)
- Made the HIR execution domain a first-class
MeshScoperegion. (#151) - Added symbolic placement and expanded HIR operation capabilities. (#88, #93)
- Added same-kernel module calls and unified function-call typing and call edges. (#86, #129)
- Added a B200 SXM target. (#76)
- Added external target services and custom target modes. (#70, #81)
- Added an executable GQA tutorial and improved the CLI surfaces through which agents discover tutorials and specifications. (#62, #115, #155)
- Added broader parser, installed-package, and CI coverage. (#58, #98, #100, #101)
Bug fixes and refactoring
- Made
analyzereports reproducible, readable, and locatable in authored source. (#150) - Unified placement syntax and semantics across parser and analysis paths. (#150)
- Preserved layouts and symbolic slice dimensions through IR transformations. (#102)
- Retained authored failure context in diagnostics. (#141)
- Fixed runtime callee dispatch and optional-weight handling. (#145)
- Added validation for HIR return annotations and fused boundaries. (#130, #142)
- Split
checkinto explicit selection, input preparation, execution, and comparison stages. (#144) - Unified module authoring and function parsing through the decorator API. (#79, #111)
- Fixed model analysis, parser surfaces, and authored expression gates. (#123, #124)
- Made analysis visitors linear and consolidated the analysis implementation. (#126, #134)
Thanks
Thanks to @bigSheep123 for contributing the B200 SXM target in #76.
v0.0.1 — A foundry for agents
This is the first public release of TileFoundry, a tile-based, agentic platform for automatic high-performance program generation across hardware.
pip install tilefoundry # Python ≥ 3.12, PyTorch ≥ 2.7Why we built this
High-performance kernels have always been expert work. The optimization space is huge — tiling, layouts, pipelining, and instruction choice all interact — and every decision made for speed still has to be checked for correctness.
LLM agents can now do much of this work, but existing compiler stacks were not built for them. The intermediate states are opaque graphs, the knowledge an agent needs is scattered across the codebase, and "is this still correct?" has no mechanical answer.
TileFoundry starts from one question: what should a program-generation platform look like when an agent is the primary user?
The answer comes down to three things, one for each problem above:
- Source to source — so agent edits stay visible and structurally safe
- Context on demand — so the right information arrives at the right time
- Fail closed, verify explicitly — so nothing passes silently
Source to source
The workflow is a loop:
- Describe. Write a published model in HIR, TileFoundry's high-level IR — a readable reference implementation, finished when it agrees with the original on real weights.
- Optimize. Write a runtime twin of one module, check it against the reference, and iterate while the numbers disagree.
- Repeat. Take the next module. Fuse neighbours. Check again.
Both sides of every comparison are readable Python source; nothing is hidden behind an opaque graph. Runtime twins preserve the reference's module and function structure, so omissions fail early instead of surviving to the final check.
Correctness is always one check away. That lets an agent optimize aggressively — and lets a human trust the result.
Context on demand
| Question | Command |
|---|---|
| What has been described? | tilefoundry models |
| What does the contract say? | tilefoundry spec |
| Does it agree with its reference? | tilefoundry check |
| What does it cost? | tilefoundry analyze |
The design lives in a normative specification, installed with the package and queryable from the command line. Instead of front-loading that knowledge into a prompt, TileFoundry lets an agent ask the question it has now. Every command level describes itself, so context arrives at the moment it becomes useful. For an agent driving the CLI in a loop, this is the difference between converging and thrashing.
Fail closed, verify explicitly
check has no default tolerance. Every output is judged by a bound the caller states, because a silent PASS teaches nothing.
Every shipped model states how far it has been verified, with the test cited as evidence. The yardstick is the model's official implementation — the Hugging Face Transformers code people already run and trust:
- L1 — each layer matches the official implementation, on random weights
- L2 — the whole forward chain matches it, end to end
- L3 — decoding from the released checkpoint reproduces its output, token for token
How much a model is trusted is a lookup, not a judgment call.
In this release
The heart of the release is the loop itself — the language for describing models, and the models already carried through it:
- A typed, shard-aware HIR — a Python DSL for describing models as nested
Modules of@funckernels, covering current inference dtypes (bf16,fp8e4m3,f8e8m0block scales). - Seven real models already described with it — DeepSeek V4 Flash, Gemma2 2B, MiniCPM3 4B, Qwen2.5 1.5B, Qwen3 1.7B (L2); Kimi Linear 48B A3B, Qwen3.5 35B A3B (L1) — installed with the package as references to copy from.
Around them, the machinery and the on-ramp:
- Scheduling built for agents — polyhedral modeling, storage-driven instruction selection, and generated kernel skeletons with holes left for the agent to fill in.
- Three end-to-end examples of optimized decode — hand-written CUDA, CuTe DSL, and TileLang — and a tutorial that teaches the loop: what to do, in what order, at what granularity.
What's next
This release establishes the foundation: the DSL for describing models, the check loop for runtime implementations, and the CLI an agent works through.
The next phase is to make the platform more useful inside the optimization loop. TileFoundry should not only report a cost or a failure; it should use the current IR, target facts, and measurements to give the agent precise, stage-appropriate guidance about what to try next. In parallel, we are extending the same describe–check–optimize workflow to more hardware targets and more inference scenarios.
Follow along
TileFoundry is early — alpha, and the APIs are still moving. Which means design feedback matters more now than it ever will again.
- Star and watch the repo to follow releases; each release note will explain the design as it grows.
- Open an issue with questions or objections — the spec is meant to be argued with.
- Try the loop: install TileFoundry 0.0.1 from PyPI, then run
tilefoundry modelsandtilefoundry tutorial.
The foundry is open. Come build it with us.