Start here
Attention Variants & Norms, in plain language
Grouped-Query Attention (GQA), Multi-Query (MQA), RoPE embeddings, and RMSNorm topology. Attention and normalization variants determine context quality, KV-cache cost, training stability, and serving throughput in large Transformers.
For a small example, one sentence contains subject, action, and location relationships. Let separate heads emphasize each relation, then concatenate their outputs and inspect why identical heads add little. This is the mechanism to keep in view as the lesson becomes more technical. Before moving on, identify the input, transformation, output, and one observation that would falsify your conclusion.
Key points
- Head dimensions and Q/K/V projection shapes.
- KV-cache growth during autoregressive generation.
- LayerNorm, RMSNorm, residual paths, and positional encodings.