Skip to content

v0.0.2 — Compiler feedback in the loop

Latest

Choose a tag to compare

@zhen8838 zhen8838 released this 10 Sep 05:06
· 30 commits to main since this release
4690c58

TileFoundry 0.0.2 turns compiler analysis into feedback that an agent can act on while optimizing real programs.

pip install tilefoundry==0.0.2

Highlights

New blog: AI Compilers in the Agentic Era

AI Compilers in the Agentic Era explains the boundary TileFoundry draws between an agent and a compiler.

The agent proposes strategies, rewrites programs, and implements kernels. TileFoundry keeps the authored HIR and runtime twin connected through three forms of compiler feedback:

  • analyze attributes performance costs before a kernel is written.
  • check protects correctness across large structural rewrites.
  • HIR types and placement expose optimization opportunities to the agent.

New example: a whole LLM decode step in one kernel

The new Nemotron-3.5-Lightning-30B-A3B example describes all 52 layers of a Mamba2, attention, and MoE model in authored HIR, then implements the decode step as one persistent TileLang cooperative kernel. (#143)

On the measured H200 setup, it reaches:

  • 287.4 tok/s at context 32, or 97.5% of SGLang
  • 231.6 tok/s at context 262080, or 83.2% of SGLang
  • one kernel launch per decode step, compared with 3212 launches in the operator-by-operator path

The full agent run, HIR, runtime twin, checks, measurements, and optimization history are included in the example.

Features

  • Added topology-scoped compute costs, authored-loop access footprints, execution timelines, and analytical performance bounds. (#80, #94, #105, #110)
  • Made the HIR execution domain a first-class MeshScope region. (#151)
  • Added symbolic placement and expanded HIR operation capabilities. (#88, #93)
  • Added same-kernel module calls and unified function-call typing and call edges. (#86, #129)
  • Added a B200 SXM target. (#76)
  • Added external target services and custom target modes. (#70, #81)
  • Added an executable GQA tutorial and improved the CLI surfaces through which agents discover tutorials and specifications. (#62, #115, #155)
  • Added broader parser, installed-package, and CI coverage. (#58, #98, #100, #101)

Bug fixes and refactoring

  • Made analyze reports reproducible, readable, and locatable in authored source. (#150)
  • Unified placement syntax and semantics across parser and analysis paths. (#150)
  • Preserved layouts and symbolic slice dimensions through IR transformations. (#102)
  • Retained authored failure context in diagnostics. (#141)
  • Fixed runtime callee dispatch and optional-weight handling. (#145)
  • Added validation for HIR return annotations and fused boundaries. (#130, #142)
  • Split check into explicit selection, input preparation, execution, and comparison stages. (#144)
  • Unified module authoring and function parsing through the decorator API. (#79, #111)
  • Fixed model analysis, parser surfaces, and authored expression gates. (#123, #124)
  • Made analysis visitors linear and consolidated the analysis implementation. (#126, #134)

Thanks

Thanks to @bigSheep123 for contributing the B200 SXM target in #76.