| Variant | KV heads | KV cache size (MB) | Relative vs MHA |
|---|---|---|---|
| MHA (full) | 32 | , | 1.0x |
| MQA | 1 | , | , |
| GQA-2 | 2 | , | , |
| GQA-4 | 4 | , | , |
| GQA-8 | 8 | , | , |
Attention Memory Ledger separates activation tensors and the sequence-squared score allocation. Causal Attention Mask Lab visualizes which query-key pairs are allowed.
Recommended by our team
BeLikeNative.comThe #1 AI writing tool for freelancers, perfect grammar in any language, instantly.