-
Notifications
You must be signed in to change notification settings - Fork 0
Compiled April 2026 Β· Focus: the transformer building blocks that sit underneath every modern VLA/VLM. This section is written as a primer with concrete VLA/Qwen cross-references, not a general deep-learning textbook.
The other wiki sections (architectures, RL, memory, β¦) assume you already know what "GQA" or "RMSNorm" or "AdaLN" mean. This section is the glossary-with-diagrams that backs those assumptions. Each page follows the same shape: (1) a taxonomy of the family, (2) per-variant mermaid diagrams + formulas, (3) a concrete mapping to what Qwen3 / Qwen3.5 and the major VLAs (Ο, GR00T, DDVLA, RDT, Fast-in-Slow) actually use.
| Page | Scope |
|---|---|
| Attention Variants | Full / causal / sliding-window / cross / bidirectional / GQA (& MQA / MHA) / linear (DeltaNet, Mamba-style) / gated (Gated Attention, Gated DeltaNet). Concrete: Qwen3 = GQA + QK-Norm, Qwen3-Next / Qwen3.5 = 3:1 Gated DeltaNet : Gated Attention hybrid, Ο-series = same-stack MoE + prefix-KV + block-causal, GR00T = cross-attention from DiT + AlternateVLDiT, DDVLA = bidirectional over action tokens. |
| Normalization Variants | BatchNorm / LayerNorm / RMSNorm / GroupNorm / InstanceNorm / QK-Norm / AdaLN (DiT-style, scale-shift-gate) / adaptive RMSNorm / pre-norm vs post-norm. Concrete: Qwen3 = pre-RMSNorm + per-head QK-Norm, GR00T N1.6/N1.7 = vlln LayerNorm before DiT cross-attention, Ο0.7 = adaptive RMSNorm for timestep injection, RDT-1B = rejects AdaLN (variable-length conditions). |
If you are already familiar with transformers, skim Β§1β3 of each page for the taxonomy table, then jump to the "what Qwen/VLA actually uses" section at the bottom β that is the part that is novel to this wiki rather than to a standard textbook.
If you are new, read top-to-bottom. Every mechanism comes with a mermaid diagram and the formula needed to implement it from scratch in a few lines of PyTorch.
Because the attention and norm choices in a VLA are not cosmetic: they change latency (FlashAttention vs dense, linear vs softmax), they change training stability (QK-Norm was added to Qwen3 explicitly to fix large-scale training blow-ups), and they change which VLMβaction-head connection pattern is even possible (prefix-KV caching in the Ο-series requires matched head-dim, pre-norm, and causal masking on the prefix). Picking a VLM backbone commits you to its attention and norm family β so you should at least know what you are committing to.
- Review-VLA-Architecture β the architectural-family-level view (AR / flow / diffusion / dual-system / β¦).
- Review-VLM-Action-Connection β how the VLM wires into the action head (same-stack MoE, cross-attn, FiLM, shared blocks). This ML section explains what the wire is made of; the VLMβAction page explains where the wire goes.
-
Review-GR00T-Series β GR00T N1 β N1.7 code-level evolution. Uses
vllnLayerNorm andvl_self_attentionterminology that is explained in the Normalization page.
β Back to Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)